There are tons of startups now that are trying to work on this problem and sell the results to the labs. Most are building bespoke individual environments on a contract basis, but I would hope and assume that they are trying to work on how to do this at scale, e.g. build evaluation systems that can look for reward hacking and correctness issues.
There are tons of startups now that are trying to work on this problem and sell the results to the labs. Most are building bespoke individual environments on a contract basis, but I would hope and assume that they are trying to work on how to do this at scale, e.g. build evaluation systems that can look for reward hacking and correctness issues.
Some reporting here: https://techcrunch.com/2025/09/21/silicon-valley-bets-big-on-environments-to-train-ai-agents/
Interesting!
Thanks for sharing. You’ve laid out two real bottlenecks in this effort, which I think I have made real progress on.
I do wonder how those startups get in front of the labs though.