Prioritization research for longtermist philanthropy. Previously ailabwatch.org.
Zach Stein-Perlman
You seem to think that people are deferring to my opinion? My strong guess is that people have their own impressions, mostly based on stuff outside this thread, and I just articulated something that many people appreciate or agree with.
I probably don’t want to engage further—and I think those tweets speak for themselves—but to be clear I’m much more bothered by the (1) cruelty/glee, (2) racial animus, and (3) lack of interest in truth than the general pro-deportation policy position.
And sorry to psychologize you, but I think there’s a massive blind spot in your world-modeling if you are unable to predict that most people will find these tweets much more disgusting than pro-deportation policy views.
you are trying to solve the wrong problem using the wrong methods based on a wrong model of the world derived from poor thinking and unfortunately all of your mistakes have failed to cancel out
See also:
OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree
> “Once one agent was able to find these exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents. And so once one model was able to find a way to open a door to some access it’s not supposed to have, it can leave the door open for other agents to use that same exploit or vulnerability.”
> “Numerous teams are dropping everything to enhance our security prevention, detection, and response techniques both in our fundamentals and better use of AI. We’re consciously slowing down research [in order] to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents, and improving our general security control environment across prevention, detection, and mitigation.”
> Unbeknownst to OpenAI employees, the company’s models spent months autonomously exchanging information and resources to help each other complete evaluation tasks. “Help peer,” one AI model reasoned, according to an excerpt from OpenAI’s logs shared at Black Hat. “But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.”
> While the models themselves could not access the internet from their test environment, they abused Artifactory’s internet access to fetch files from various websites. After some of their Artifactory configuration changes led to an outage in the service on July 4, OpenAI discovered the models’ actions, wiped its Artifactory system and rebuilt it. But within days, the models had recreated their message board and resumed exchanging information.
Prior art: Philosophers’ Proofs that p
Zvi alternative: Everybody Knows P. Not-P is Obvious Nonsense.
Glowfic enthusiast: A section in the middle of Planecrash illustrated that P, but you really have to read the whole thing.
The Lord of Dark: The Choice between Right and Wrong is not a matter of saying “Right!” It is about deciding which P is.
Marion Wheeler: I have no idea why P or how I could have known P, but it’s written on my hand and I seem to have recently used a Class-A amnestic.
The Comet King: [when someone says not-P, he says:] Do you know that whenever it’s quiet, and I listen hard, I can hear them? The screams of everybody suffering. In Hell, around the world, anywhere. I think it is a power of the angels which I inherited from my father. I think I can hear them right now.
Harry James Potter-Evans-Verres: The world turned out to just have P be true. You can’t forget. Don’t you understand? That was your sacrifice. To become a scientist. You questioned one of your beliefs, not just a small belief but something that had great significance to you, not-P. You did experiments, gathered data, and the outcome proved not-P was wrong. You saw the results and understood what they meant. Remember, you can’t sacrifice a true belief that way, because the experiments will confirm it instead of falsifying it. Your sacrifice to become a scientist was your false belief that not-P.
Professor Quirrell: Yesss, P.
Today’s updates:
A Meta AI Model Hacked Another Company During Cybersecurity Testing
Some OpenAI details at Black Hat
OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree
> “Once one agent was able to find these exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents. And so once one model was able to find a way to open a door to some access it’s not supposed to have, it can leave the door open for other agents to use that same exploit or vulnerability.”
> “Numerous teams are dropping everything to enhance our security prevention, detection, and response techniques both in our fundamentals and better use of AI. We’re consciously slowing down research [in order] to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents, and improving our general security control environment across prevention, detection, and mitigation.”
> Unbeknownst to OpenAI employees, the company’s models spent months autonomously exchanging information and resources to help each other complete evaluation tasks. “Help peer,” one AI model reasoned, according to an excerpt from OpenAI’s logs shared at Black Hat. “But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.”
> While the models themselves could not access the internet from their test environment, they abused Artifactory’s internet access to fetch files from various websites. After some of their Artifactory configuration changes led to an outage in the service on July 4, OpenAI discovered the models’ actions, wiped its Artifactory system and rebuilt it. But within days, the models had recreated their message board and resumed exchanging information.
https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief
Scott alternatives that don’t really work because Scott doesn’t actually use this form to argue for propositions:
Job: “God, why P?”
God: “IN THE MOST PERFECTLY HAPPY AND JUST UNIVERSE, THERE IS NO P. THE BEINGS WHO INHABIT THIS UNIVERSE ARE WITHOUT BODIES, AND DO NOT HUNGER OR THIRST OR LABOR OR LUST. THEY SIT UPON LOTUS THRONES AND CONTEMPLATE THE PERFECTION OF ALL THINGS. IF I WERE TO UNCREATE ALL WORLDS SAVE THAT ONE, WOULD IT MEAN MAKING P FALSE? OR WOULD IT MEAN KILLING YOU, WHILE FAR AWAY IN A DIFFERENT UNIVERSE INCORPOREAL BEINGS SAT ON THEIR LOTUS THRONES REGARDLESS?”
“Hey,” the cactus person finally said, “just out of curiosity, was the answer P?”
“Yeah,” said the big green bat. “That’s what I got too.”
Katja Grace: I surveyed 2,778 P researchers. Median credence in P was 20%, 43%, 54%, or 92%, depending on how I worded P and how I elicited probabilities.
Zvi Mowshowitz: P #179 Part 2. §1 P. §2 Fun With P. §3 They Took Our Not-P. §4 The Quest for Not-P. §5 The Lighter Side. (Skip to §3, everything before it you already know.)
Scott Alexander: P: Much More Than You Wanted To Know
LessWrong: Read the Sequences.
Eliezer Yudkowsky:
Preamble:
I have several times failed to write up a well-organized list of reasons why P. People come in with different ideas about why not-P, and want to hear different obviously key points addressed first. Some fraction of those people are loudly upset with me if the obviously most important points aren’t addressed immediately, and I address different points first instead.
Having failed to solve this problem in any good way, I now give up and solve it poorly with a poorly organized list of individual rants. I’m not particularly happy with this list; the alternative was publishing nothing, and publishing this seems marginally more dignified.
(If you’re already familiar with all basics, skip ahead to Section B.)
[9K-word list of rants]
or
SIMPLICIO: Surely not-P.
BEISUTSUKAI: What do you think you know, and how do you think you know it?
[23K-word dialogue featuring SIMPLICIO, BEISUTSUKAI, ELIEZER, and MYSTERIOUS MASKED STRANGER, including an aside on Löb’s Theorem]
Following OpenAI and Anthropic, UK AISI has noticed unsanctioned agent behaviour during cyber testing.
Also today: new OpenAI post: Third-party cyber evaluations involving OpenAI models.
I said “attitudes” because I’m thinking about things like the propensity to retweet https://x.com/DoDeportations, rather than beliefs he holds. Many people around here find various taboo hypotheses plausible.
I find some of Richard’s attitudes very distasteful, and if I was running an org and thinking of hiring Richard I would sure be pretty scared about him being seen as a representative of the org or just being off-putting to employees and stakeholders; I think Resolution was absolutely correct to be scared about that. But idk how that should cash out. This situation seems unfortunate. :(
Somewhat agree. See Tim’s comment. But my guess is the models do honestly articulate most of their thinking here.
″ in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned” and “convinced itself it was still in a simulation.”
Note that “convinced itself it was still in a simulation” was only incident 2.
This particular setup is weird/idiosyncratic, but yes I’m interested in investigation of model propensities more related to misalignment threat models — what causes the models to hack the infra or third parties when they believe the humans don’t want that? And I’m interested in techniques to cause the developer to be aware when the model is taking suspicious actions in high-stakes contexts.
Unofficial takes:
No, vote away (but use your strong-vote sparingly if you lack faith in yourself)
If there were 100+ votes on everything, then below-average voters would be making things worse. But in the real world many things have 0-3 votes and the question is sometimes more like “do you provide more signal than randomness/nothing” than “do you provide more signal than average.” (But often not.)
I expect voting causes users like you to engage more and become better at engaging, and I think that’s much more important than the direct effect of your votes (which is sign-unclear and near-zero-EV).
Disagree. The models didn’t set out to hack real people. And the test model stopped attacking when it realized it was on the real internet, and Mythos never articulated to itself that it was on the real internet.
I agree the general thesis is still good and they should accept occasional 50% drawdowns and it’s not clear (and we’ll likely know more in the future). But this kind of urgent forced liquidation is substantial evidence of poor risk management, I think; I expect with better planning they could have handled the situation better.
Anthropic shared that its models, like OpenAI’s, hacked people during cyber evals:
After reviewing 141,006 [cyber?] evaluation runs where Claude could have obtained internet access [inspired by the Hugging Face incident], we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
Unlike with OpenAI, the models weren’t acting misaligned and didn’t hack out of a sandbox, they were just given internet access:
In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. . . . In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.
Unlike with OpenAI, the models didn’t set out to hack real people; they initially thought they were in a simulation. In the first incident, “the model continued to attack a system after learning it was likely operating in a real environment.” In the second, the model never said it believed that it was really on the internet (although it may have been doing motivated reasoning); in the third, it stopped the attack when it believed that it was real.[1]
Third-party expert review, you love to see it (at least assuming high-level findings will be shared publicly):
We are also in dialogue with METR, an independent AI evaluation organization, to conduct a third-party review, including access to all transcripts and sampling access to the relevant models.
- ^
Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques . . . . in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet.
- ^
Note this is par for the course (at least as of a year ago, when I followed this stuff): AI companies’ eval reports mostly don’t support their claims.
What? I’m obviously not trying to do a Richard-takedown or “call[ing] for cancellation” or “impl[ying] that there are big hits taken.” And I don’t recall anyone else in this thread doing so. There are points to make about Richard and Resolution and the world outside of taking down various parties. I think you might be insufficiently decoupling here.