But where is the part where they actually slow down? The part that a foreign government can’t provide, I think, is allowing Dario to call up Sam and Elon and agree to actually stop doing some commerce activity (to avoid antitrust). I guess it’s different if another government is the one making the calls and telling providers to slow down
MichaelLowe
The 10% figure by Evan has been frequently cited by the media, and probably lead to increased attention as a result.
There might be a small minority of online commenters who were already negatively predisposed attacking that figure, but I don’t think that extends to how most people perceive it.
A huge part will probably be the combination of clarity and the credible signal of earnestness by resigning.
1)
He made very forthright and clear statements such as “The people building AI earnestly believe that it could kill us all by the end of the decade.”; “They are racing straight to self-improving superintelligence and gambling with our lives.”
These are just very intelligible to broad audiences. Previous statements where less clear to non technical audiences. Compare a statement such as Hinton’s estimate of 10% chance of human extinction, which might generate less virality because many people could be reading 10% as “won’t happen”. Also compare with Miles Brundage’s “Neither OpenAI nor any other frontier lab is ready, and the world is also not ready.” when he left OpenAI: we read “not ready” and think “not ready to prevent human extinction”, but most people would just not read it with the same frame of mind, and I guess might think something mild such as “not ready to handle the economic transition painlessly”.
2) Resigning is just a very clear signal that he is serious about this. It’s really hard for lay people to evaluate arguments about safety risks on their merits, so it makes sense to instead defer to such signals. You can’t explain this away with being motivated by PR, and common sense dictates that a person foregoing a big fat pay check would do so only for very good reasons.
Contrast this with Sam Altman, who has also been relatively clear in the past, but it would take a couple of inferential steps to take this seriously, if the first thought is “then why the hell are you building it?”.
This does not seem convincing enough to be the primary reason for going viral. There are tons of news articles out there usually, and the WSJ is not extremely widely read. But maybe I am missing something.
Isn’t the implication of him being the person to write it up that the NYU mathematician would get the (?) millennium prize? This sure sounds generous to me, I am sure OpenAI has no shortage of mathematicians on staff that would like to do this.
That you can destroy a country by implementing disastrous policies, seems very realistic. The part that would have made me indeed skeptical, is that Gorbachev could make it to the top without strengthening the regime in bad ways during that process.
> Though on my model of the situation, if it broke into the host machine,
Sorry, is it known that the model(s) did find info such as the flag or the API key for the judge?
>my impression often is that these views … are the most common
I find it hard to be well calibrated myself, but I have the intuition that looking at the attitude expressed in mainstream media (i.e. news papers or TV, not necessarily the comments section) is, ironically, more representative of the view of the general population than social media like reddit, youtube comments or x on this topic. Writing comments on these platforms is a very strong selection effect (I recall a statistic that about 1% of users write 90% of comments or something) that imho does not select for sanity, plus terminally online people are known to interpret everything through their own politics.
There is a separate but partially related effect where people in the tech community have pre-existing instincts (i.e. favoring open-source, being wary of luddites, “enlightened skepticism”) that make them distrust anything coming out of these companies.
Would be very interested in more of these, incl. other sources!
“There’s a loop that consequentialists can get into that’s something like: … “Now that I’ve gained power, ”
But I am not at Anthropic, and neither are any of my friends. So the dynamic does not apply to me.And at any rate, the argument is not that ppl at Anthropic or OpenAI are not deluding themselves. It’s rather whether their effect on the world is better/worse than the counterfactual AI company. You correctly point towards the need for the “safety-consciousness” to play out in anything meaningful to count as evidence, but then you only list as a criterion that their current alignment techniques should actually generalize.
To me it is pretty clear that we have solid evidence: with their taking a stand on autonomous weapons Antropic have behaved quite clearly differently than the counterfactual AI company would have. In general, comparing frontier AI companies in general to tobacco or oil companies, the differences in ethical behavior seems very clear to me. You might of course still argue that this is in fact worse for the world, because chumps take the safety postures of frontier companies seriously, thus preventing actually meaningful political actions.
I greatly enjoy feeding ducks in my local park, so I am biased, but to me the reasons for not feeding them have never quite held up. (although of course they should be fed food appropriate for their digestive system). Specifically on the Malthusian trap, it seems that I either don’t feed the wild animal and it dies from starvation, or I feed it, it reproduces and then its offspring dies from starvation, which is neutral in terms of animals dying from starvation at least within that species.
We reviewed a demonstration of this specific technique being used to identify a small number of previously known, minor vulnerabilities. These vulnerabilities all appear relatively simple, and we have found that other publicly-available models are able to discover them as well without requiring a bypass.
To me, this seems like potentially pretty misleading communication from Anthropic and somewhat silly.
Amazon claims that they used a technique that successfully bypasses Fable’s safeguards to get an response from Mythos, which IIUC Anthropic does not dispute. To the question about whether the safeguards work, it is probably immaterial whether the response is something you could have gotten from other models. The value of this demonstration is that it shows that the safeguards did not work, not that in this specific case any harm could/was done. Unless Anthropic makes the imho very implausible claim that the safeguards are developed in a way that they detect whether a particular vulnerability is previously known (or minor), this technique is evidence that an attacker totally could have used this for vulnerabilities that other models could not find.
Anthropic has not demonstrated that their safeguards work reliably enough in practice, and thus I am not very sympathetic to their perspective on the merits of the case (although procedurally this is obviously bad from the Trump admin). It seems good if the result of this is that the burden of proof of safety shifts towards the companies.
Yes, external testing has not found an universal jailbreak according to the model report, but it’s pretty unclear that the focus on universal jailbreaks is sufficient. No external party has tested Anthropic’s models (or other labs’ for that matter) AND their defense in depth approach to explicitly certify that they work sufficiently well, instead it’s isolated fishing for universal jailbreaks, and otherwise taking Anthropic at their word. I don’t think a convincing case has been made that universal jailbreaks are necessary to cause harm with a model, in fact it seems very possible that an attacker could use a custom jailbreak that works for a given environment or/and pivot to different jailbreaks when one stops working.
Ally quickly and perhaps you preserve more lives, as the Tlaxcala (another group in the region that were historical enemies of the Mexica) may have done?
In hindsight, that seems like a pretty robust strategy to me? I understand that you mention possible zero-sum dynamics, but I don’t think this negates the relatively straightforward case for it.
The Spanish did not come in with genocidal intentions and were a pretty legalistic culture anyway. I also guess that they were not interested in total dominiation, they wanted trade and probably the recognition of Spain’s king as the official ruler.
Sure they tried, but the question is whether they succeeded in keeping key decision makers in the dark, right? There was a large contingent of media that showed frequent mishaps of Biden and seemed convinced of his cognitive decline, so the info was out there and it is plausible that Biden’s inner circle failed to deceive key decision makers ( I have not read the book)
>Either way, having a people’s assembly selected from protesters is pretty silly, since it defeats the point of a citizen’s assembly as a representative sample of the population.
I am not sure. There could be various benefits to using citizen assemblies in the described manner that go beyond narrow PR purposes, even if they are not perfectly representative. My guess is that representativeness is only one of the benefits many non-dumb people ascribe to citizen assemblies.
Just a short list of the possible benefits: Engaging in a discussion about the topics in a protest can lead to better retention and understanding on the side of participants; People having more political discussions in meat space rather than online is actually good; Allows making some social connections among the participants which increases the likelihood they come back; Practicing such ad-hoc citizen assemblies can increase support for them on a national level (though you could call that PR).
Sorry, that was a bit unclear.
What I have in mind as a central example is a chain of thought that says something like,”the system prompt I received looks like it’s from a chat interface, and there are enough details in the request that look realistic, so this request is probably real”. Since models sometimes do verbalise eval awareness rather clumsily ( for example Gemini 3 stating that it’s probably in an evaluation because the date is too advanced), I’d expect some examples of such reasoning if situational awareness is very high.
FWIW, I’d update strongly if I see examples demonstrating that it is relatively common for LLMs to reason about the system prompts they receive and what that entails for how they should reply. So far, I am not aware of many such examples.
Very prosaic, but taking a strong stance is good for PR if you think there is value to be perceived externally and internally as an ethical company. If you decide to forego the military contract, might as well reap those benefits, and that quote in particular is probably uplifting to many people.
Also, them signaling that they are happy to support the DoW with offboarding their models does not sound like they want to force an escalation.
>Am I correct here that METR’s embedded asessor for Anthropic will be… ex-Anthropic employee?
No, you are not; you are merely speculating.