I mean isn’t this advice tantamount to saying “if you cared enough about preventing this incident you wouldn’t be training AI models at all”? Sure, okay, but, I don’t think this is particularly actionable advice for OpenAI, nor do I think it’s particularly good advice for the field. Reinforcement learning is not intrinsically bad, in the limit pretty much any imaginable objective reaches the perverse/Goodhart regime, reinforcement learning just tends to reach it faster because it e.g. pushes things towards more agentic behavior faster.
It wasn’t meant as useful advice, just a note on the framing, like the other comment. The point is that they just don’t care enough about safety. They could be forced to do proper alignment research and they still wouldn’t care. A more sensible organization could be working (only) on image models, studying smaller LLMs, developing theory, trying to build pure tool AIs, etc. Instead, they are creating agents in the stupidest way possible. They aren’t even trying, and I’m not sure there’s anything to salvage.
I mean isn’t this advice tantamount to saying “if you cared enough about preventing this incident you wouldn’t be training AI models at all”? Sure, okay, but, I don’t think this is particularly actionable advice for OpenAI, nor do I think it’s particularly good advice for the field. Reinforcement learning is not intrinsically bad, in the limit pretty much any imaginable objective reaches the perverse/Goodhart regime, reinforcement learning just tends to reach it faster because it e.g. pushes things towards more agentic behavior faster.
It wasn’t meant as useful advice, just a note on the framing, like the other comment. The point is that they just don’t care enough about safety. They could be forced to do proper alignment research and they still wouldn’t care. A more sensible organization could be working (only) on image models, studying smaller LLMs, developing theory, trying to build pure tool AIs, etc. Instead, they are creating agents in the stupidest way possible. They aren’t even trying, and I’m not sure there’s anything to salvage.