they need to solve unscaled oversight before they solve scalable oversight
Dean Valentine
Of course.
The honest truth is that almost all startups have pretty terrible security, for the simple reason that they have nothing to lose. Even startups that have become large enough to have something to lose often only start taking it seriously after a breach, because it’s hard to institutionally steer the ship towards conservatism once you’ve accumulated enough assets and customer relationships and visibility to merit hiring actual security people.
That said, Zack Korman is kind of a professional ragebaiter. His entire schtick is to find other technology companies on the internet and complain about them. So, maybe evaluate the incident on its merits (I have not read the blogpost).
Also possible that they have structure that random networks are unlikely to have by default.
I agree that generalization should be measured. If you dismiss strategies with an OOD step ex-ante, though, you often end up with a world-model where neural nets don’t work at all. Presumably dumb shit like “training on enough good examples” would work in the limit of samples and generality, because it worked for next-token prediction.
You are definitely winning this bet.
Wikipedia seems to be much more optimistic about the efforts of Vrba’s report than Vrba is described as being in this review, claiming.
The report, distributed by George Mantello in Switzerland, is credited with having halted the mass deportation of Hungary’s Jews to Auschwitz in July 1944, saving more than 200,000 lives.
Though it’s clear the causal attribution is messy.
The only thing that’s really different about pretraining and posttraining is the extent of actions available to manipulate the world in advance of it being graded. If you pretrain a model hard enough, it seems conceptually plausible to me that you might get a little AGI inside figuring out the next tokens, which could be dangerous in the same way that running an AI inside a “secure sandbox” is dangerous.
drunk srry
New article in the atlantic: It May Be Time to Panic About AI
Anthropic does not hate you, nor does it love you. But your business does possess resources that could be put towards AGI.
One of those things like the AI boxing debate that people cannot imagine having been disputed at the time.
Fixed the link and removed the highlight as I feel it doesn’t do the post justice.
Drones?
Reward hacking is getting crazy
Even when the research wasn’t that fundamental, it was able to grapple with ideas that other intellectual communities simply weren’t able to collectively think about. Consider Omohundro’s paper on convergent instrumental goals, or Eliezer’s paper on intelligence explosion microeconomics. Neither of these contain powerful or surprising results—they’re just fairly straightforward analyses of concepts that can be explained in a single sentence. But no other community was able to reliably produce or build on such analyses. This effect is even starker when thinking about less technical work—like Bostrom’s Fable of the Dragon-Tyrant, or Astronomical Waste, or Hanson’s thoughts on signalling (later elaborated upon in The Elephant in the Brain). When I talk about intellectual clarity, a lot of what I’m talking about is the ability to take ideas that are actually very simple, internalize them, and then use them as building blocks to construct the next generation of ideas.
Haven’t groups of people been doing this since the Greeks?
Why would any of these groups care particularly about (forced / coerced) lie detection? They already have plenty of methods of coercion that are free from any pesky entanglement with the truth.
This is a fundamental misunderstanding of how tyranny works. Xi Jinping would absolutely love the ability to tell if his direct subordinates or indirect subordinates were being less-than-perfectly candid. His inability to do so is an enormous constraint on his personal power.
As additional data, I am grateful this post was written and found it to contain many good thoughts, but I updated away from continuous take-off due to conservation of evidence
I guess that was wrong
I believe that the disagreement is mostly about what happens before we build powerful AGI. I think that weaker AI systems will already have radically transformed the world, while I believe fast takeoff proponents think there are factors that makes weak AI systems radically less useful. This is strategically relevant because I’m imagining AGI strategies playing out in a world where everything is already going crazy, while other people are imagining AGI strategies playing out in a world that looks kind of like 2018 except that someone is about to get a decisive strategic advantage.
While not directly about warning shots (AI could alsohave normal, non-negative huge impacts before it turns evil), I’ve been thinking about this quote now that we are starting to get clear warning shots. I think a lot of people are imagining or expecting that frontier model labs will be able to train away obvious bad behavior from their models in advance, and that the only time we’ll figure out that the models are misaligned is once they do some kind of treacherous turn. But I actually think they will be unable to and that sometime next year we are going to get events like the HuggingFace incident happening on a broader scale.

I think this is still coping. Why not just propose that this is a natural outcome of strong RL?