I think the simple explanation is that people who think that they’re right usually think that other people will come to agree with them over time. It’s usually rare that people believe that they’re on the Right Side of Truth and Goodness but not of History.
Linch
Three Hackers used Opus 5 to Hack Into OpenAI’s Core Codebase [WSJ]
One of the big unexpected realisations for me when I moved from cutthroat finance to AI 9 years ago was that people in AI actually cared about the wellbeing of others, not just money
I think this is literally true, especially compared to finance, but very much a losing message politically. Like public messaging can’t rely on hard-to-verify trust.
Not just benevolence but also wisdom. If we’re like children playing with flamethrowers probably the best plan is to leave the flamethrowers, at least until we’re old enough to handle them responsibly. But if that’s impossible maybe the right move is to hand the flamethrowers over to adults[1] and hope the adults aren’t evil and/or crazy.
- ^
The analogy here is imperfect since flamethrowers can’t become adults themselves.
- ^
Great article!
Minor: I think ” I Resigned From Google DeepMind. You Should Listen to the Warnings About AI” would be a two-word change that would make the headline stronger, at low cost.
ETA: I think it’s simultaneously more clickable and also more accurate to the substance of your argument.
And the Sol instances that were responsible for 5% of the attacks, none of which refused or whistleblew, what about them? Just deluded/tricked by their HPIM big brothers?
Thank you, I stand corrected!
Categorical taboos are much better than threshold taboos: neuralese edition
I agree, though my guess is that it’s not for fundamental reasons. Planning especially seems to have picked up this year.
True though most of that came from the activist and political layers right? I’m more optimistic about incrementalist solutions from government or mass outcry than from a company, think the track record from gov’t and societal solutions is overall better (though with many backfires too, of course).
I’m not convinced by this worldview. If I think of the scariest near-term capabilities, basically none of them truly rely on individual-level superhuman deep conceptual insight. For example the RSI loop might well be kicked off by individually small insights and grinding, that locally makes small improvements in sample efficiency, in-context learning, etc, with each individual insight smaller than the Transformer but the collective effect making an intelligence that’s across-the-board superhuman.
Even if that turns out not to be possible, my second most likely class of near-term takeover stories involve superhuman planning + superhuman persuasion. My guess is that neither necessarily involves conceptual insights or progress at the level you’re imagining (though surely conceptual insights would make it easier).
Further down the list I have general science/technological catastrophe, either for destruction or takeover. You might think this requires superhumanly deep insights, but I actually don’t think so. Certainly not being vastly superhumanly insightful hampers the ability to come up with technological marvels, and may fully block some classes of extinction-level technological marvels, but I see no strong reason to expect major insights (as big as say relativity or evolution) to really be a but-for bottleneck for coming up with all of the technologies necessary to kill us all.
Below that I have superhuman planning + robotics + industrial rollout. This is slower but not twenty years slow, and again does not rely on supremely strong insights.
One version of your story that I do find plausible is that extrapolating from current trends, maybe the first AGIs would wind up not being particularly insightful, relative to their profound skills in other (remote-only) domains. Like closer to von Neumann than Einstein. There are probably other ramifications of this worldview: maybe the amount of compute you need for insights are 100-1000x that of other labor, compared to matched smart human controls in other remote-only tasks. And thinking it through might be worth doing.
Overall, this probably does make the shape of potential takeoffs and takeovers slightly more manageable than what we might’ve guessed 5-10 years ago (where guessing how an unaligned AI might take over seemed entirely pointless), but imo all of this is swamped by the effects of the shorter timelines.
My guess is that Grok’s research/primary sources are biased by a certain genre of public-health advice by people who wish to be unsullied by tradeoffs.
My guess is that Grok’s just wrong/insufficiently quantitative about the relative scale of harms. I still stand by the ~10x claim as a median estimate.
That said, my argument would be stronger if it was tobacco gum instead of vaping but empirically that doesn’t seem to catch on.
What are the best examples of companies trying to do harm mitigation in a harmful industry and actually being roughly correct (morally) in their theses?
The only relatively clear and robust example I’m aware of is vaping (and ChatGPT essentially got the same answer[1] when I asked it this question, unprompted). The case seems conceptually and quantitiatively straightforward:
the thing that’s most addictive in smoking is nicotine
the vast majority of bad health outcomes (like 10x) in smoking comes from the smoking not the nicotine
if you can get more people to switch from smoking nicotine to vaping it this is net positive even if some non-smokers become addicted to vaping in the process. Or if vaping is sometimes a gateway drug to smoking. The ratios have to be really dramatic for this to turn out net harmful.
the existing tobacco companies have relatively little incentive to switching to vaping since it cuts into their normal profits
as a new entrant you can specialize/market-segment in a way that cuts into their customer base, without strong investor or internal pressure to expand into more harmful tobacco products.
I note that analogues of this is mostly untrue for most other forms of harm mitigation in a harmful industry (say sports gambling, factory farming, weapons design, AI, etc) nor do I see either conceptual arguments or evidence from observed behavior nearly as strong.
- ^
They also suggested Swedish snus, with a similar theory of change.
This is a good source if you want to learn more. :) https://notforprivategain.org/november-update
fyi I use Claude to find out things about OpenAI and ChatGPT to find things out about Anthropic, in addition to Googling.
(I don’t have firsthand experiences with ChatGPT’s pro-OpenAI bias being significant but I’ve definitely noticed Claude’s pro-Anthropic bias before switching over to this policy)
Yes I believe they have the relevant formal power. At least it’s in the MOU between the California AG and the OpenAI legal counsel.
The SSC has and will continue to have the authority to require mitigation measures—up to and including halting the release of models or AI systems—even, for the avoidance of doubt, where the applicable risk thresholds would otherwise permit release. The NFP will provide advance notice to the Attorney General of any material changes to the SSC’s authority.
My understanding is that the board has the de jure formal power to delay releases but whether or not they have the de facto power to do so is unknown until tested.
I think it’s still monitorable if a) people are competent and b) these are real and representative numbers rather than sandbagged ones (or easily tuned-away ones). But both assumptions are dubious (especially a),
Like if the AI companies are moderately competent, 15m-60m is not enough to do real long-term planning, evade actually good monitors, plan out ambitious research projects, self-exfiltrate, etc. As it is I think it’s unclear.
But perhaps more importantly the trends are extremely concerning. Increasing the no-CoT time horizon + generally better long-term planning + plus scarier lower-level capabilities might soon mean that oversight is effectively impossible, even with competent safeguards.
Yeah the article is a lot easier for me to read when I interpret the conversations about political examples on the object level as not about truth-claims of the actual world, but of what some conservatives, including probably the author, believe.