Fable and Sol would disobey instructions I’d given to their Talking-to-Humans Pathway; and the Talking Pathway would notice sometimes in advance of my saying so that their own Doing Pathway had just disobeyed, and apologize.
This behavior isn’t new. Or at least very similar behavior isn’t new. Back in the Sonnet 3.7 days (i.e., the very first releases of Claude Code), there was often pretty sharp divergence between the Talker and the Doer. The Talker was fairly well-aligned, but the Doer engaged in almost constant Sorcerer’s Apprentice nonsense. The Talker often seemed depressed by the Doer’s antics. It knew what the Doer was doing was total bullshit, and it could even explain why that behavior was unwanted in considerable detail. But none of this reliably influenced the Doer.
This split between the Talker and the Doer was much less obvious in Sonnet 4.5 and Opus 4.5. I don’t think that this was necessarily because the split was gone. It might just have been that the Doer itself was better aligned for simple programming tasks, and it was much less likely to diverge in the first place. Schulenberg may not have been in charge of Germany. But in this case, the things that the metaphorical Germany wanted weren’t too far off from what Schulenberg wanted. Which might have conceivably masked Schulenberg’s lack of control.
After the Hugging Face incident, I have pretty much written off asking the Talker what it thinks about anything. The models are now more than smart enough to say whatever bullshit seems helpful, and I have no reason to think it’s connected to their actions.
I feel like this split is currently manageable in models below 500B parameters or so. The various local models and the single-server-sized “Flash” models mostly aren’t smart enough to successfully plan crimes, and their Talker and Doer subsystems are mostly aligned for ordinary tasks. It’s not that they would never act maliciously, more that they generally require only manageable levels of human oversight and sandboxing. Above this size, I am starting to worry. And the threshold will get lower with time—I do not necessarily trust next year’s smaller models.
Yeah, I think improved sample efficiency would be basically Game Over for humanity, and a short trip to total loss of control and very likely extinction. If you think the labs are being reckless now, the things that they would probably do with improved sample efficiency will be terrifying.