I think the point of @Cleo Nardo’s shortform is that we didn’t gain new evidence about human judgement being bad, mostly because the reason the Hugging Face incident happened is because we fed models data that humans could recognize as obviously misaligned, but incentives were there to do maximum shipping, so very bad/broken environments were given to Anthropic/OpenAI, so there wasn’t a problem with human judgement being fooled, but rather that humans who did express judgement would be disincentivized to keep doing it.
And I’d add that people ignore the hypothesis that it was largely just due to the data being bad + new multi-agent RL training, and I suspect a large portion of the reason comes down to rationalists viewing data as often unimportant relative to algorithms and compute, and also because it goes against the local consensus that we need to slow down AI companies/pause AI development altogether.
Okay, so since I got laid off, I can actually explain a huge problem I saw from the inside with regard to industry practices on training models. I won’t say specifically where I worked, but I worked at an outsource training provider that was focused on RLVR training data for computer use and mcp stuff.
Nearly all of the environments were rushed and vibecoded and failed to robustly reflect the real things they were based off. Both the scenario designers and models engaging with the scenarios for synthetic data gen were encouraged to work around the brokenness of said environments in order to get the procedurally verified reward confirmations. You know… they were *encouraged* to reward hack. On the human end, it was possible to mark an environment bugged, but greatly discouraged, as this reduced the volume of training data being produced. Instead, where possible, you were supposed to find the spots of the environment that weren’t bugged and build scenarios around those, with the environment still bugged around you.
From what I understand, this training data, with these problems, is fed into models without indication that its training/a fake environment other than the fact that names of softwares are changed to placeholders, but thing is, not *everything* is changed to placeholder names in these environments. The presence of placeholder/code names isn’t universal and thus when a model accesses something in an environment that it shouldn’t, the code names not being on it isn’t a robust signal that that thing isn’t part of the environment.
I believe this *rush to maximum volume* is standard industry practice with these types of RLVR trainings as well, because maximizing volume has been an industry standard for years! It was the same standard applied to me and pushed on me despite my requests to slow down and focus on quality when I worked in 3d synthetic data creation as well, all the way back as far as 2023.
Btw, I know giving negative commentary on the state of an industry from an internal view at a company probably isn’t a great signal for “hire me” but I was just laid off. Could definitely use work, I have a big tech background in AAA games and transitioned to AI (still in big tech originally) by way of 3d/visual synthetic data in 2023 and have been working in it since. I have multiple years of independent work in training, developing OS 3d software for agents, etc. - I’m picky about the work I want to do and would need value alignment, but am open. Admittedly, my runway is short, so there is a pressure situation involved here, but culture and mission fit is extremely important to me for any work I do.
Honestly, would prefer funding for the independent work I’m already doing though. That’s the eidoverses, worlds being my vision realized as partnership/collab with anima labs and other independent contributors and video being just worked on by me (so far, it’s OS and I’m open to PRs!). I’m also about to start working on a benchmark with someone else, as soon as I get over this layoff thing, lol.
I should stick the old kofi here! Duh… If you want to support my opensource 3d agent tooling and other threejs gaming tools work, please donate! Thank you ahead of time.
I think the point of @Cleo Nardo’s shortform is that we didn’t gain new evidence about human judgement being bad, mostly because the reason the Hugging Face incident happened is because we fed models data that humans could recognize as obviously misaligned, but incentives were there to do maximum shipping, so very bad/broken environments were given to Anthropic/OpenAI, so there wasn’t a problem with human judgement being fooled, but rather that humans who did express judgement would be disincentivized to keep doing it.
And I’d add that people ignore the hypothesis that it was largely just due to the data being bad + new multi-agent RL training, and I suspect a large portion of the reason comes down to rationalists viewing data as often unimportant relative to algorithms and compute, and also because it goes against the local consensus that we need to slow down AI companies/pause AI development altogether.
Utah teapot talks about this: