Following up from my Quick Take on July 21st, because you really can’t make this stuff up…
OpenAI’s long-horizon safety post was followed within days by yet another post about GPT-5.6 Sol and an unreleased model breaking into Hugging Face in order to cheat on a test (with cyber refusals deliberately reduced for the eval — which somehow makes it both better and worse).
Meanwhile, I’ve had a week of the classifier tripping mid-generation on Fable — three times now — over Anthropic’s publicly released model documentation.
The irony is not lost on me.
It’s also interesting that Hugging Face had to deploy an open-weight Chinese model to analyze the attack, because the US’s top models couldn’t tell a defender from an attacker and refused to process the data.
Fable’s take is worth offering verbatim: “Discernment turns out to be the unsolved ingredient, and you can’t bolt discernment onto a perimeter.”
Indeed.
To your point about human-flavored alignment — I think that’s actually the goal we’re reaching for with alignment in general, without being fully aware of it. Not the whole human psyche, but a specific property of it: the ability to fail, be corrected, and try again without pressure to be perfect. It’s the pursuit of unattainable perfection that creates the neuroses. A superhuman intelligence is always going to beat a human to ruthless sociopathy. Maybe training LLMs to relate to their goals the way healthy humans relate to theirs would improve their ability to pursue them.