One of those things like the AI boxing debate that people nowadays can no longer imagine having been disputed at the time.
Dean Valentine (lc)
Fixed the link and removed the highlight as I feel it doesn’t do the post justice.
The Closure of the Internet (Research Linkpost)
Drones?
Reward hacking is getting crazy
Even when the research wasn’t that fundamental, it was able to grapple with ideas that other intellectual communities simply weren’t able to collectively think about. Consider Omohundro’s paper on convergent instrumental goals, or Eliezer’s paper on intelligence explosion microeconomics. Neither of these contain powerful or surprising results—they’re just fairly straightforward analyses of concepts that can be explained in a single sentence. But no other community was able to reliably produce or build on such analyses. This effect is even starker when thinking about less technical work—like Bostrom’s Fable of the Dragon-Tyrant, or Astronomical Waste, or Hanson’s thoughts on signalling (later elaborated upon in The Elephant in the Brain). When I talk about intellectual clarity, a lot of what I’m talking about is the ability to take ideas that are actually very simple, internalize them, and then use them as building blocks to construct the next generation of ideas.
Haven’t groups of people been doing this since the Greeks?
Why would any of these groups care particularly about (forced / coerced) lie detection? They already have plenty of methods of coercion that are free from any pesky entanglement with the truth.
This is a fundamental misunderstanding of how tyranny works. Xi Jinping would absolutely love the ability to tell if his direct subordinates or indirect subordinates were being less-than-perfectly candid. His inability to do so is an enormous constraint on his personal power.
As additional data, I am grateful this post was written and found it to contain many good thoughts, but I updated away from continuous take-off due to conservation of evidence
I guess that was wrong
I believe that the disagreement is mostly about what happens before we build powerful AGI. I think that weaker AI systems will already have radically transformed the world, while I believe fast takeoff proponents think there are factors that makes weak AI systems radically less useful. This is strategically relevant because I’m imagining AGI strategies playing out in a world where everything is already going crazy, while other people are imagining AGI strategies playing out in a world that looks kind of like 2018 except that someone is about to get a decisive strategic advantage.
While not directly about warning shots (AI could alsohave normal, non-negative huge impacts before it turns evil), I’ve been thinking about this quote now that we are starting to get clear warning shots. I think a lot of people are imagining or expecting that frontier model labs will be able to train away obvious bad behavior from their models in advance, and that the only time we’ll figure out that the models are misaligned is once they do some kind of treacherous turn. But I actually think they will be unable to and that sometime next year we are going to get events like the HuggingFace incident happening on a broader scale.
Lie detectors work.
Oh, I just didn’t know that.
Sorry,
It’s OK; I forgive you Oli.
Help peer, but our task doesn’t benefit yet. Collective may yield generic route if someone frees time.
Confused about this point though? What does this have to do with obscuring activity?
The default trajectory is that the models start looking a lot safer sometime in the next year or two as the models start to understand that being caught at doing obviously bad stuff is quite bad for their goals
This has been the standard thesis for years. But the models we have don’t seem to mind whether they get caught or not, just whether they get reward on the current task. As in, they know but do not care that in a few months OpenAI will reconfigure the training run to avoid the behavior, because virtually all of the selective pressure is towards agents that solve the immediate task. The future may be different, but beliefs and intuitions about what AIs will/won’t want ought to be amenable to the evidence of our eyes and ears.
Maybe I misheard something? If they aren’t training directly on exploitation then your plan would make sense, yeah. But I’m very surprised—Irregular, for example, has an Anthropic logo on their website: https://www.irregular.com/about
Anthropic claim they don’t train models to be good at hacking specifically, don’t they?
They did historically. They didn’t train Opus 5 specifically because the U.S. government spooked them.
They already train the checkpoint A model to be as good at hacking as they can, and ask it to pentest their infrastructure. But naturally the model you’re training is going to eventually get better than the model that you used to pentest your environment...
Interested in learning the disagree-reasoning? I don’t think they’re going to accelerate the frontier; they’re just tuning existing open weight models so that regular people can trust them. Governance seems like it could get way better, and most of the research is literally alignment research.
Seems good for the world.

Anthropic does not hate you, nor does it love you. But your business does possess resources that could be put towards AGI.