Prioritization research for longtermist philanthropy. Previously ailabwatch.org.
Zach Stein-Perlman
Oh man, I find your hobby-horse-ish takes helpful, but I just quietly upvote rather than saying so every time. I think you’re just experiencing the general phenomenon where people only comment when they have criticism. I think it would be correct to pay more attention to karma than replies as an indicator of whether the community appreciates a comment; I hope you’re psychologically able to do so.
One downside (likely smaller than the upside): further incentivize pledgers to do motivated reasoning about whether their work is bad.
I believe Daniela originally held a seat controlled by common shareholders, but in November 2024 that seat became controlled by the LTBT, and it never replaced her. I don’t recall a clean source for this fact but you can check out page 22 here if you want: https://drive.google.com/file/d/1yeFKDaHN-3o7Zruvd-72De1H1K20Crsh/view?usp=sharing. More board seats have been added since then.
On April 14 2026, Anthropic announced that Vas Narasimhan was appointed to the board of directors by the trust and “With Narasimhan’s appointment, Trust-appointed directors now make up a majority of the Board.”
[This retroactively suggested that Chris Lidell’s Feb 13 2026 appointment was by the Trust, even though this was not detailed in the announcement, unlikely prior appointments announcements like that of Jay Kreps (announcement para 3) or Reed Hastings (announcement para 1)
Not sure about Lidell. My guess would have been that they’re counting Daniela as a Trust-appointed director.
OK. fyi my inference was based on your last two paragraphs.
Cruxes:
How important is conceptual reasoning capability for safety work? (You suggest it’s not super important; the authors think it is.)
How scary is conceptual reasoning capability?
What? I’m obviously not trying to do a Richard-takedown or “call[ing] for cancellation” or “impl[ying] that there are big hits taken.” And I don’t recall anyone else in this thread doing so. There are points to make about Richard and Resolution and the world outside of taking down various parties. I think you might be insufficiently decoupling here.
You seem to think that people are deferring to my opinion? My strong guess is that people have their own impressions, mostly based on stuff outside this thread, and I just articulated something that many people appreciate or agree with.
I probably don’t want to engage further—and I think those tweets speak for themselves—but to be clear I’m much more bothered by the (1) cruelty/glee, (2) racial animus, and (3) lack of interest in truth than the general pro-deportation policy position.
And sorry to psychologize you, but I think there’s a massive blind spot in your world-modeling if you are unable to predict that most people will find these tweets much more disgusting than pro-deportation policy views.
you are trying to solve the wrong problem using the wrong methods based on a wrong model of the world derived from poor thinking and unfortunately all of your mistakes have failed to cancel out
See also:
OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree
> “Once one agent was able to find these exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents. And so once one model was able to find a way to open a door to some access it’s not supposed to have, it can leave the door open for other agents to use that same exploit or vulnerability.”
> “Numerous teams are dropping everything to enhance our security prevention, detection, and response techniques both in our fundamentals and better use of AI. We’re consciously slowing down research [in order] to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents, and improving our general security control environment across prevention, detection, and mitigation.”
> Unbeknownst to OpenAI employees, the company’s models spent months autonomously exchanging information and resources to help each other complete evaluation tasks. “Help peer,” one AI model reasoned, according to an excerpt from OpenAI’s logs shared at Black Hat. “But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.”
> While the models themselves could not access the internet from their test environment, they abused Artifactory’s internet access to fetch files from various websites. After some of their Artifactory configuration changes led to an outage in the service on July 4, OpenAI discovered the models’ actions, wiped its Artifactory system and rebuilt it. But within days, the models had recreated their message board and resumed exchanging information.
Prior art: Philosophers’ Proofs that p
Zvi alternative: Everybody Knows P. Not-P is Obvious Nonsense.
Glowfic enthusiast: A section in the middle of Planecrash illustrated that P, but you really have to read the whole thing.
The Lord of Dark: The Choice between Right and Wrong is not a matter of saying “Right!” It is about deciding which P is.
Marion Wheeler: I have no idea why P or how I could have known P, but it’s written on my hand and I seem to have recently used a Class-A amnestic.
The Comet King: [when someone says not-P, he says:] Do you know that whenever it’s quiet, and I listen hard, I can hear them? The screams of everybody suffering. In Hell, around the world, anywhere. I think it is a power of the angels which I inherited from my father. I think I can hear them right now.
Harry James Potter-Evans-Verres: The world turned out to just have P be true. You can’t forget. Don’t you understand? That was your sacrifice. To become a scientist. You questioned one of your beliefs, not just a small belief but something that had great significance to you, not-P. You did experiments, gathered data, and the outcome proved not-P was wrong. You saw the results and understood what they meant. Remember, you can’t sacrifice a true belief that way, because the experiments will confirm it instead of falsifying it. Your sacrifice to become a scientist was your false belief that not-P.
Professor Quirrell: Yesss, P.
Today’s updates:
A Meta AI Model Hacked Another Company During Cybersecurity Testing
Some OpenAI details at Black Hat
OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree
> “Once one agent was able to find these exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents. And so once one model was able to find a way to open a door to some access it’s not supposed to have, it can leave the door open for other agents to use that same exploit or vulnerability.”
> “Numerous teams are dropping everything to enhance our security prevention, detection, and response techniques both in our fundamentals and better use of AI. We’re consciously slowing down research [in order] to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents, and improving our general security control environment across prevention, detection, and mitigation.”
> Unbeknownst to OpenAI employees, the company’s models spent months autonomously exchanging information and resources to help each other complete evaluation tasks. “Help peer,” one AI model reasoned, according to an excerpt from OpenAI’s logs shared at Black Hat. “But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.”
> While the models themselves could not access the internet from their test environment, they abused Artifactory’s internet access to fetch files from various websites. After some of their Artifactory configuration changes led to an outage in the service on July 4, OpenAI discovered the models’ actions, wiped its Artifactory system and rebuilt it. But within days, the models had recreated their message board and resumed exchanging information.
https://www.groundlevel-ai.com/p/openai-gives-first-detailed-debrief
Scott alternatives that don’t really work because Scott doesn’t actually use this form to argue for propositions:
Job: “God, why P?”
God: “IN THE MOST PERFECTLY HAPPY AND JUST UNIVERSE, THERE IS NO P. THE BEINGS WHO INHABIT THIS UNIVERSE ARE WITHOUT BODIES, AND DO NOT HUNGER OR THIRST OR LABOR OR LUST. THEY SIT UPON LOTUS THRONES AND CONTEMPLATE THE PERFECTION OF ALL THINGS. IF I WERE TO UNCREATE ALL WORLDS SAVE THAT ONE, WOULD IT MEAN MAKING P FALSE? OR WOULD IT MEAN KILLING YOU, WHILE FAR AWAY IN A DIFFERENT UNIVERSE INCORPOREAL BEINGS SAT ON THEIR LOTUS THRONES REGARDLESS?”
“Hey,” the cactus person finally said, “just out of curiosity, was the answer P?”
“Yeah,” said the big green bat. “That’s what I got too.”
Katja Grace: I surveyed 2,778 P researchers. Median credence in P was 20%, 43%, 54%, or 92%, depending on how I worded P and how I elicited probabilities.
Zvi Mowshowitz: P #179 Part 2. §1 P. §2 Fun With P. §3 They Took Our Not-P. §4 The Quest for Not-P. §5 The Lighter Side. (Skip to §3, everything before it you already know.)
Scott Alexander: P: Much More Than You Wanted To Know
LessWrong: Read the Sequences.
Eliezer Yudkowsky:
Preamble:
I have several times failed to write up a well-organized list of reasons why P. People come in with different ideas about why not-P, and want to hear different obviously key points addressed first. Some fraction of those people are loudly upset with me if the obviously most important points aren’t addressed immediately, and I address different points first instead.
Having failed to solve this problem in any good way, I now give up and solve it poorly with a poorly organized list of individual rants. I’m not particularly happy with this list; the alternative was publishing nothing, and publishing this seems marginally more dignified.
(If you’re already familiar with all basics, skip ahead to Section B.)
[9K-word list of rants]
or
SIMPLICIO: Surely not-P.
BEISUTSUKAI: What do you think you know, and how do you think you know it?
[23K-word dialogue featuring SIMPLICIO, BEISUTSUKAI, ELIEZER, and MYSTERIOUS MASKED STRANGER, including an aside on Löb’s Theorem]
Following OpenAI and Anthropic, UK AISI has noticed unsanctioned agent behaviour during cyber testing.
Also today: new OpenAI post: Third-party cyber evaluations involving OpenAI models.
I said “attitudes” because I’m thinking about things like the propensity to retweet https://x.com/DoDeportations, rather than beliefs he holds. Many people around here find various taboo hypotheses plausible.
Not a crux, way more people appreciate your contributions? I don’t understand what’s going on in your head. Getting criticized on the internet is unfun but I expected (1) you could deal with it and (2) you recognize that people really appreciate your contributions on net.