The AI safety ecosystem is so well resourced that it has been correctly identified by many as one of the best paths into high prestige AI research jobs.
This person on twitter has written a popular article about getting into frontier ai labs and a Field Guide to AI Fellowships. The “AI Fellowships” are mostly AI safety programs funded by CG/OpenPhil. I have also noticed that people in ML research are quite likely to have heard of MATS and be interested in participating, even when they have very little interest in AI safety.
Idk how good or bad this is. It definitely causes a lot of ML researchers to engage with the AI safety literature, where they otherwise would not have. But it’s worth noting that while the primary driver of applications to programs like MATS used be concern about AI safety, now it is increasingly a desire to work at a frontier lab.
I used to select participants for LASR labs, one of programs listed, and we actively tried to choose people who cared about AI safety, but I think we often did not succeed and indeed some people on the program now work at frontier labs in roles that I think have little to do with safety.
Safety-washing in practice mostly looks like people rationalizing research or jobs that they would have done anyway as necessary for safety. It’s easy to do because it is in fact genuinely unclear in most cases what is helpful and harmful for safety. Only by looking at the larger pattern can we notice the suspicious abundance of conveniently overlapping opportunities that further both safety and a person’s short-term interests. However, I think it will be less common in people whose primary motivation is to prevent AI catastrophe.
Maybe unpopular but I think that if you’re concerned about defection after programs like MATS then even more than “caring about AI safety” you specifically want to select for effective altruists, who are much likelier to care about AI actually going well than people who specifically care about AI safety. For the latter group, the motivation is partly not wanting everyone to die but also partly being AGI-pilled, being interested in the technical problem of alignment, and following prestige gradients, none of which are going to robustly keep participants on track.
maybe i’m overconfident but i generally find that it’s somewhat easy to tell if someone genuinely cares. you can’t really fake genuinely caring (or at least as genuinely as one can about anything—perhaps everything bottoms out in some other drives, but that’s good enough for me). there are some false negatives—people who genuinely care who have taken on the affectations of faking it, because they mistakenly think this helps their chances at success. there are also some people who care so much that they go crazy and start being counterproductive. but i don’t think i’ve ever heavily overestimated how much someone cares about AI safety.
This also seems right to me, but I don’t think MATS or the labs (or the funders) care if someone actually cares. They seem happy to tell themselves stories that they are making things better by getting people into the labs (or lab-adjacent orgs) even if they really don’t seem very safety motivated.
Maybe you could make an argument that they don’t care enough, but I’m pretty confident that MATS does care about this. When I did it, they had a whole reading group program whose primary aim, as I understand it, was to get people to understand and care more about the fundamental safety issues. And I also believe they try to select for people who care about safety, at least to some extent.
I predict that a randomly selected MATS participant from cohorts in the past two years would more likely than not fail to pass my ITT about why AI poses an existential risk.
Very interested in running the experiment, if anyone has a way to source random MATS scholars.
I think it’s gotten worse over time. I should have said “I don’t think MATS or the labs (or the funders) care much if someone actually cares”. I agree the caring isn’t zero.
i want to spend money on making there be more good alignment researchers and care about them actually caring about alignment. what would you recommend i do?
I think up until very recently funding Lightcone would have been a good bet at your scale (though you should of course be appropriately skeptical of me saying this). Funding projects doing good object-level work seems good.
Before I think about concrete recommendations, by “alignment researchers” do you mean people doing technical alignment work, or eval work, or do you include technical governance work like MIRI’s technical governance team? Or are you just using it as a proxy for work trying to make alignment go well in some form or another?
I think this is very hard thing to get right, and I don’t think there’s a scaleable way or org. that you can spend money on.
I think the current best bet would be to find existing people that are technical, driven, safety-pilled and already done some independent research, and fund them to continue doing that?
If true, I think part of the change can be attributed to a more dramatic increase in those who are technically competent and have various accomplishments yet don’t care so much for safety (though would love a role at Anthropic). If the sample of those folks has increased, it becomes more challenging to turn down those who are highly competent but maybe care a bit less about safety. Especially if you are a mentor and you just want someone to execute exceptionally well on your project.
Last time I mentored (for SPAR), I had to turn down a lot of folks who were more accomplished and likely to succeed (e.g. professors) so that I could accept those I felt would have more long-term potential to contribute to safety.
Last time I mentored (for SPAR), I had to turn down a lot of folks who were more accomplished and likely to succeed (e.g. professors) so that I could accept those I felt would have more long-term potential to contribute to safety.
Wouldn’t professors be more likely to only apply if they genuinely care about safety, while younger people (overall) might be still figuring things out and happy for any opportunity to advance their career?
In the specific situation I’m referencing, no. My impression is that they were just curious about automated research scaffolds, agent frameworks and applying interpretability for capabilities. They did not seem all that interested in alignment.
I could imagine that in practice many professors end up optimizing for prestige and more paper publications (that have nothing to do with superintelligence), whereas a younger person may not be locked into that mindset yet.
I think it’s important to note that the reading list was new in MATS 5, revised for MATS 6, and I’m unsure if it persisted to MATS 7, since attendance beyond (iirc) ~week 2 or so cratered in both cohorts. From the inside, it felt like a thing a few staff members desperately felt scholars needed, rather than an effective or popular educational initiative.
I think that individual people at MATS (including both directors) care a great deal. I think that MATS, as an institution, habitually under invests in value alignment in order to better serve the immediate desires of its users (mentors who want the most competent applicants; applicants who want the most lucrative job).
In the three cohorts I spent at MATS, there was never an explicit org-wide filter for value alignment. Mentors picked their scholars. Some seemed to filter on value alignment, and others did not seem to.
It depends a lot on the type and quantity of interactions you have. I think it can be pretty hard to assess this from just a CV, short answer questions and a 20 minute interview.
I mentor some people for AI fellowships/safety programs and worry about this a fair amount. How do you select people who care about safety (specifically: what criteria/vibes do use) and overall do you think it’s still positive in expectation to do this type of mentorship for the next 12 months?
Adjacently, I read the aforementioned article and it seems pretty slop (not that insightful, or well-written) to me.
I think successful prior career in industry could be a good sign. Or any other career, or generally being older, so that it’s less likely you thought 2 years ago that MATS would look great in your CV.
Dave Banerjee’s test is a decent first pass filter, given his observation that “a large fraction of researchers in AI safety/governance fellowships cannot do any of these things”. Like any interactive interview, you’d want to probe their mental models with a few follow-up questions.
I think this makes the right kind of mentorship and exercising the right kind of judgment more valuable. FWIW the article’s author got outed as a fabulist of some form in the last few hours; he admitted to lying about MATS and hasn’t furnished any proof about his other resume lines.
I think this is a legibility problem; the whole rationalist corpus is available for easy consumption, Bayesian reasoning is not that hard to figure out, and there’s a pretty universal presumption of good faith (the author still has a lot of defenders!), so it’s really hard to sit two people down, one of whom is a mercenary there to use the fellowship as a stepping stone to wealth, power, etc, the other there out of genuine belief, and detect which is which reliably off of conversation alone.
At least at the selection layer, it seems like the most important thing is finding small honesty and motivation tells. I’m a recent college grad and in college I helped found a club which had pretty explosive and rapid success which we believed was thanks to our distinct internal culture and norms which put us at odds with our otherwise famously-mercenary college.
In the end we found two filters to be reliable and workable. The first was a very brief screening application which explicitly prohibited AI-generated responses. Whoever read the application plugged results into Pangram if they were suspicious. Our belief was that if the applicant couldn’t do a ~15 minute screener with their own writing, then we had no reliable positive indications of their motivation or work quality, and a pretty strong negative signal about their honesty.
The next was a curveball question in which we explicitly encouraged “I don’t know, here’s what I do know and here’s how I’d start solving it” and clarifying-follow ups as answers. The questions always varied and were always premised around very specific information about our field; having an answer from the hip was very high-signal for motivation and work quality, asking the right questions was higher signal for work quality but lower for motivation.
These two filters were the only absolutes in the process, and any other information we asked (resume, specific question content, etc) was purely about placement or deciding between marginal applicants. When I graduated, the club had basically kept its original culture while being successful. Maybe this approach will be helpful for thinking through the problems?
My guess at one of the best metics / interview questions here is “tell me about a lesswrong post you really liked” or something similar. Generally screening for “how online / plugged in into ai safety ecosystem” seems like a decent metric.
This is not exactly what I want, though, since I think anyone seriously applying to these programs will have done some reading and would be able to answer about their favorite LW post competently.
I do, or at least there’s a small cluster I keep returning to, and the one I’d name first is Eliezer’s “Local Validity as a Key to Sanity and Civilization.”
The core move is simple and I find it keeps paying out: you should be able to evaluate whether a single step of reasoning is valid independently of whether you like where the argument lands. And the post’s real claim isn’t just epistemic hygiene for individuals — it’s that this habit is load-bearing for civilization. A society where people can agree “that inference is invalid” even when they disagree about the conclusion has a working immune system; one where validity-judgments get pulled toward tribal allegiance loses the ability to error-correct at all. That second-order framing — local validity as the thing that lets pluralistic systems stay sane — is what elevates it above a standard “commit no fallacies” essay.
What I like is that it’s a genuinely useful idea rather than a clever one. It gives you a concrete thing to watch for in yourself: the moment you notice you’re scrutinizing an argument harder because you dislike its conclusion, you’ve caught yourself doing the bad thing.
The honest caveat is that “favorite” for me is closer to “highest hit-rate on rereads” than to nostalgia. By that standard, the runners-up are Scott’s “The Tails Coming Apart as Metaphor for Life” and Garrabrant’s “Goodhart Taxonomy” — both for similar reasons, that they name a structural failure mode crisply enough to actually use.
Do you have one? I’d be curious whether yours skews toward the reasoning-tools cluster or somewhere else entirely.
Remove some of the obvious LLM tells, and I’d have a lot of trouble telling it apart from a fluent LW’er “genuinely” interested in civilizational issues, discourse norms, and AI safety.
Agree with Leo that this is not a hard thing to distinguish, provided that value alignment matters for these fellowships.
When we do interviews for CMU’s AI Safety org, it seems like open-ended questions about viewpoints (e.g. “what about the current pace of AI keeps you up at night?” or “if you weren’t doing AI Safety research, what would the alternative be?”) enable us not just to distinguish between people who “speak the language”. Another shibboleth is actually under-awareness of the community—someone who is quick to recite a bunch of names may be less concerned with the issues at hand. Whether we want to do this is another question.
FWIW, MATS clearly does still source people who are excited about safety; most other fellows in my cohort act as if they are fighting for their future! Still others are intrigued by the more challenging theoretical and empirical questions. I trust the staff in their experience vetting out people who do not wish to engage the space genuinely.
The thing I am more worried about are AI fellowships taking money from participants (or fellowships created to reach personal agendas of the founders).
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible—but I hear the same people express fear privately. No other human activity poses this level of danger.
If you work at an AI lab, possibly the most impactful you can do now is resign and try to start a social cascade that stigmatises participating in the AI race.
But please do it right. Most AI employee resignations so far have been unstrategic, seemingly spur of the moment decisions with little consideration about making the most of the opportunity. You should have a media strategy lined up before you leave, journalists briefed, statements written and edited.
Or even better, you should not directly resign but instead use your newfound freedom to send the costliest possible signal to your company and to the world that you are actually scared about AI—to the point where you are less worried about keeping your job than doing something to slow AI progress.
If you are in this position, DM me if you want help executing this as effectively as possible.
I witness that this resignation and the warning going with it was (briefly but faithfully) mentionned in the morning news on the leading french news radio. That’s a good update that such an alert can even reach such a mainstream foreign media usually focused on politics and economics.
I live far from the Bay Area and have many friends outside of this world. I think people here underestimate how much young “normies” genuinely hate this technology and the companies building it. It generally comes out as pushback against datacenters and environmental concerns but I think the younger generation is fairly open to the idea at least that they will be replaced soon and can imagine this tech fully disempowering them, setting up a surveillance state etc. I don’t think you need to convince them that it will be the AI killing them to convince them to support a pause. They already think CEOs are maximally evil and want everyone else dead. Even if there’s issues with the specifics, I hope more people here come to this conclusion and realize that a political solution is likely the easiest path to prevent catastrophic scenarios soon. I’ve been convincing people to call their senators/house reps in support of a pause. Outside of the tech world there is huge support for regulating AI.
In many ways I am finding technical people outside of the safety community to be shockingly irrational about this while people who don’t understand the tech have a strong sense that something will likely go wrong along the way. There’s something to be said for the past sins of the tech industry setting up the right mental model among the general populace to distrust the people developing capabilities.
This does not seem convincing enough to be the primary reason for going viral. There are tons of news articles out there usually, and the WSJ is not extremely widely read. But maybe I am missing something.
People are really getting concerned about alignment. Just this morning I saw several liberal commentators on UK bluesky say that senior politicians need to start speaking up about the risk. Also I think mainstream commentators who’ve been quietly worried about it for a while (like Yglesias) are feeling emboldened to speak up about it more.
A huge part will probably be the combination of clarity and the credible signal of earnestness by resigning.
1)
He made very forthright and clear statements such as “The people building AI earnestly believe that it could kill us all by the end of the decade.”; “They are racing straight to self-improving superintelligence and gambling with our lives.”
These are just very intelligible to broad audiences. Previous statements where less clear to non technical audiences. Compare a statement such as Hinton’s estimate of 10% chance of human extinction, which might generate less virality because many people could be reading 10% as “won’t happen”. Also compare with Miles Brundage’s “Neither OpenAI nor any other frontier lab is ready, and the world is also not ready.” when he left OpenAI: we read “not ready” and think “not ready to prevent human extinction”, but most people would just not read it with the same frame of mind, and I guess might think something mild such as “not ready to handle the economic transition painlessly”.
2) Resigning is just a very clear signal that he is serious about this. It’s really hard for lay people to evaluate arguments about safety risks on their merits, so it makes sense to instead defer to such signals. You can’t explain this away with being motivated by PR, and common sense dictates that a person foregoing a big fat pay check would do so only for very good reasons.
Contrast this with Sam Altman, who has also been relatively clear in the past, but it would take a couple of inferential steps to take this seriously, if the first thought is “then why the hell are you building it?”.
I’d make an easier request of developer employees: just state your opinion on AI risks publicly, and include your perception of the range of opinions inside your org.
You can get more media boost if you quit, but each of these statements helps shift the Overton window, and is a powerful reference for those of us trying to communicate AI risks.
We may have an opportunity right now to create the cascade that shifts the Overton window and public opinion dramatically in a short time, much like how opinions on COVID shifted in early 2020.
This might be stupid, but I have a proposal: If you are resigning, consider calling the local police, the FBI, or whoever else might help, inform them of the situation, and try to convince them that they should raid the offices of your former employer to stop them from doing something that is a potentially imminent danger. After that fails (presumably), go public and say what you attempted. Twitter seems full of people who think this is just a publicity stunt. Most of them probably can’t affect much, but some might, and actually calling the police (or another law enforcement agency), especially if a raid actually happens and disrupts their productivity, might help convince them otherwise.
It’s bad they think Jacob Coxon’s resignation is a publicity stunt because it would be to somehow help the AI industry. If whoever resigns next puts enough effort into making it clear they want to take down their former employer, or at least stop attempts at recursive self-improvement, it doesn’t matter much if it is meant as publicity, as people would understand it is publicity meant to help achieve those goals.
I don’t necessarily agree with this proposal on the object level, but the insights are interesting. If we want this to stay in the news, ideally instead of a boring drip-drip-drip of employee resignations, we get a boom-bang-crash of drama. If there’s a way to resign in a different, flashy way which somehow “advances the story”, adds color, etc. that’s liable to generate continuing coverage.
Furthermore, the ideal resignation would somehow address this suspicion that it’s all about creating hype for investors. For example, imagine if a resigning employee called on politicians to tax away all the profits from the AI industry, to reduce the incentive to invest in doom. Headline: “Capitalism will kill us all, resigning AI employee says”. At that point, it becomes difficult to claim that the resignation is some sort of corporate gambit, since what corporation would beg the government to punitively tax them?
Another idea is to resign from OpenAI and then sue OpenAI. There must be something you can sue them for, right? Or how about calling on securities regulators to block the upcoming IPOs? Is there any way they could do that? Essentially call for some sort of policy which will hurt the profits of these AI companies so people don’t think it’s all about investor hype. “Ethical investors should boycott upcoming IPO, says resigning AI engineer.” Simply linking to PauseAI could also be a credible signal that you actually want a pause and you’re not just marketing.
Another idea is to resign from OpenAI and then sue OpenAI. There must be something you can sue them for, right?
You could rightly say that they are endangering your life, which is grounds for a lawsuit. Whether a judge would go for that is another question. But anyone could file that lawsuit, not just an ex-employee. Maybe there’s a better angle an ex-employee would have.
A mental model for media. There are various podcasts and broadcast news channels. Some of them might be interested in doing an interview with someone on the topic of AI safety. We should aim to saturate these interview channels to the greatest degree possible.
For example, if Coxon has more media requests than he has time to field, it might be possible to coordinate and have him say something like: “Sorry I don’t have time to chat—but now there is news that ANOTHER employee has resigned from Anthropic, for similar reasons. Would you like to talk to them?”
Or, it’s possible that certain channels are open to covering this topic, but don’t want to talk to Coxon because they feel the Coxon “scoop” has already been covered to death. So when formulating media strategy for another resignation, you could specifically target channels which did not cover Coxon, but are still “newsy”/”techy”/etc. enough that such a resignation would be within the channel’s topical scope. Then you contact them, with the subtext that: “You missed the Coxon scoop, but you could be among the first channels to cover this new scoop!”
The desired end result is that no matter where a person gets their news, they’ve seen coverage related to the big AI company resignations.
I imagine working with a publicist could be really helpful since they’re professionals at coordinating stuff like this.
This is an attempt to compile all publicly available primary evidence relating to the recent death of Suchir Balaji, an OpenAI whistleblower.
This is a tragic loss and I feel very sorry for the parents. The rest of this piece will be unemotive as it is important to establish the nature of this death as objectively as possible.
I was prompted to look at this by a surprising conversation I had IRL suggesting credible evidence that it was not suicide. The undisputed facts of the case are that he died of a gunshot wound in his bathroom sometime around November 26 2024. The police say it was a suicide with no evidence of foul play.
Most of the evidence we have comes from the parents and George Webb. Webb describes himself as an investigative journalist, but I would classify him as more of a conspiracy theorist, based on a quick scan of some of his older videos. I think many of the specific factual claims he has made about this case are true, though I generally doubt his interpretations.
Webb seems to have made contact with the parents early on and went with them when they first visited Balaji’s apartment. He has since published videos from the scene of the death, against the wishes of the parents[1] and as a result the parents have now unendorsed Webb.[2]
The cause of death was decided by the authorities in 14 (or 40, unclear) minutes.[4]
The parents arranged a private autopsy which “made their suspicions stronger”.[5]
The parents say “there are a lot of facts that are very disturbing for us and we cannot share at the moment but when we do a PR all of that will come out.”[6]
The parents say “his computer has been deleted, his desktop has been messed up”.[7]
Although the parents also said that their son’s phone and laptop are not lost and are in escrow.[8][9] I think the claim of the computer being deleted is more up-to-date, but I’m not sure as that video was posted earlier.
It was his birthday and he bought a bike on the week of his death.[10]
He said he didn’t want to work and he was going to take a gap year, “leaving the AI industry and getting into machine learning and neuroscience” but also he was planning to start his own company and was reaching out to VCs for seed funding.[11]
He had just interviewed with the New York Times and he was supposed to do further interviews in the days after his death.[12]
According to the parents and Webb, there are signs of foul play at the scene of death:
There are several areas with blood, [Confirmed from pictures] suggesting to Webb and the parents he was trying to crawl out of the bathroom.[13][14]
Webb says the body had bleeding from the genitals.[15] I’m not aware of a better source for this claim, so right now I think it is probably false.
The trash can in the bathroom was knocked over.[13][16] [Confirmed from pictures].
A floss pick is on the floor.[13][17] [Confirmed from pictures]. Webb interprets this as being dropped at the time of death, suggesting that Balaji was caught by surprise.
The path of the bullet through the head missed the brain. I’m not sure what the primary source for this is, but I’m not sure why Webb would invent this, so I think it’s true. Webb takes this as evidence that it was shot during a struggle rather than at the considered pace of a suicide.[18]
The bullet did not go all the way through the head, suggesting a lower caliber, quiet gun.[19]
According to Webb and the parents, the drawers of the apartment were ransacked, the cupboards were thrown open.[8][20] From the pictures this looks false, although the apartment is very messy and his hiking backpacks are strewn around with much of their contents on the table (he had recently returned from a hiking trip).
The blood on the sink looks different, suggesting to Webb that it came from a different part of the body.[21] This is not obvious to me from the pictures but not implausible and the main pool of blood looks surprisingly dark.
There is a half-eaten meal at the desk in the apartment. [Confirmed from pictures].
There is a tuft of Balaji’s hair, soaked in blood, under the bathroom door. [Confirmed that’s what it looks like in the pictures], again suggesting to Webb a violent struggle.
According to the parents, he had a USB thumb drive which is now missing, containing important evidence for an upcoming court case about OpenAI’s use of copyrighted data.[8][22]
People that spoke to him around the time of his death report him to have been in high spirits and making plans for the near future.[23]
George Webb claims there were security cameras working all on floors except the floor which he lived on.[24] This appears to conflict with the parents’ claim that the police said no one came in or out (see below), but may be referring to different cameras, as the parents also mention that the murderer could have come through a different entrance to the main one.[25]
The parents say “OpenAI has deleted the copyright data that was evidence that was given to the discovery for the [New York Times] lawsuit. They deleted the data and now my son is also gone, so now they’re all set for winning the lawsuit… It’s also said that my son had the documentation to prove the copyright violation. His statement, his testimony would have turned the AI industry upside down...”[26]
Looking into the details of this, OpenAI did delete some data but this wasn’t a permanent deletion of any of the primary sources and I think it was probably an accident and not significant to the outcome of the case.
I don’t see any strong reason to believe Balaji had secret evidence that would have been critical to the outcome of the case.
Ilya Sutskever had two armed bodyguards with him at NeurIPS.
Evidence against:
One reason the authorities gave for declaring it a suicide was that CCTV footage showed that no one else came in or out of the apartment.[27]
In high school Balaji won a $100,000 prize for a computer science competition. His parents didn’t find out until they saw the news online, suggesting he may not have been very open with them.[28]
My interpretations:
If we interpret the apartment as simply messy (as it looks to me), rather than ransacked, then we can probably discount the knocked-over trash can, the floss pick on the floor and the half-eaten meal. We can also probably discard the hypothesis of someone trying to locate a USB drive with secret information, which raises more questions than it answers (why didn’t he reveal this information before? why didn’t he back up this crucial data anywhere?).
In my uninformed view, it doesn’t look like the pictures of the scene of death strongly suggest a struggle between murderer and victim, although I don’t know how to explain the tuft of hair.
The motivations of OpenAI or some other actor to murder a whistleblower are unlikely. The most plausible to me is that they want to send a warning to other potential whistleblowers, but this isn’t very compelling.
There’s no smoking gun and the parents (understandably) do not look like they are thinking very systematically to establish a case for foul-play. This is notable because their claim of foul-play is the main factor that privileged this hypothesis to credible people.
Balaji appeared from the outside to be a happy and highly successful person with important plans in the next few days. It is surprising that someone like that would commit suicide.
Overall my conclusion is that this was a suicide with roughly 96% confidence. This is a slight update downwards from 98% when I first heard about it and overall quite concerning.
I encourage people to trade on this related prediction market and report further evidence.
I’m not linking to this evidence here, in the spirit of respecting the wishes of the parents, but this is an important source that informed my understanding of the situation.
The undisputed facts of the case are that he died of a gunshot wound in his bathroom sometime around November 26 2024. The police ruled it as a suicide with no evidence of foul play.
As in, this is also what the police say?
Did the police find a gun in the apartment? Was it a gun Suchir had previously purchased himself according to records? Seems like relevant info.
Yes, edited to clarify. The police say there was no evidence of foul play. All parties agree he died in his bathroom of a gunshot wound.
Did the police find a gun in the apartment? Was it a gun Suchir had previously purchased himself according to records? Seems like relevant info.
The only source I can find on this is Webb, so take with a grain of salt. But yes, they found a gun in the apartment. According to Webb, the DROS registration information was on top of the gun case[1] in the apartment, so presumably there was a record of him purchasing the gun (Webb conjectures that this was staged). We don’t know what type of gun it was[2] and Webb claims it’s unusual for police not to release this info in a suicide case.
Well, it seems quite important whether the DROS registration could possibly have been staged. If e.g. there is footage of Suchir buying a gun 6+ months prior, using his ID, etc. then the assassins would have had to sneak in and grab his own gun from him etc. which seems unlikely.
Is the interview with the NYT going to be published?
Is any of the police behavior actually out of the ordinary?
Well, it seems quite important whether the DROS registration could possibly have been staged.
That would be difficult. To purchase a gun in California you have to provide photo ID[1], proof of address[2] and a thumbprint[3]. Also it looks like the payment must be trackable[4] and gun stores have to maintain video surveillance footage for up to year.[5]
My guess is that the police haven’t actually invested this as a potential homicide, but if they did, there should be very strong evidence that Balaji bought a gun. Potentially a very sophisticated actor could fake this evidence but it seems challenging (I can’t find any historical examples of this happening). It would probably be easier to corrupt the investigation. Or the perpetrators might just hope that there would be no investigation.
There is a 10-day waiting period to purchase guns in California[5], so Balaji would probably have started planning his suicide before his hiking trip (I doubt someone like him would own a gun for recreational purposes?).
Is the interview with the NYT going to be published?
I think it’s this piece that was published before his death.
Is any of the police behavior actually out of the ordinary?
Epistemic status: highly uncertain: my impressions from searching with LLMs for a few minutes.
It’s fairly common for victim’s families to contest official suicide rulings. In cases with lots of public attention police generally try to justify their conclusions. So we might expect the police to publicly state if there is footage of Balaji purchasing the gun shortly before his death. It could be that this will still happen with more time or public pressure.
Ilya Sutskever had two armed bodyguards with him at NeurIPS.
Some people are asking for a source on this. I’m pretty sure I’ve heard it from multiple people who were there in person but I can’t find a written source. Can anyone confirm or deny?
Ilya Sutskever had two armed bodyguards with him at NeurIPS
I don’t understand how Ilya hiring personal security counts as evidence, especially at large events like a conference. Famous people often attract unwelcome attention, and having professional protection close by can help deescalate or deter random acts of violence, it is a worthwhile investment in safety if you can afford it. I see it as a very normal thing to do. Ilya would be vulnerable to potential assassination attempts even during his tenure at OpenAI.
Thank you, this is very interesting and it seems like you did a valuable public service in compiling it
The motivations of OpenAI or some other actor to murder a whistleblower are unlikely. The most plausible to me is that they want to send a warning to other potential whistleblowers, but this isn’t very compelling
What do you think of the motive that he was counterfactually going to testify in a very damaging way, or that he had very damaging evidecne/data that was deleted?
The U.K. will seek to use its upcoming G20 presidency to forge a global agreement on AI, Prime Minister Andy Burnham said ahead of a visit to the U.N. General Assembly in New York.
This comes as discussions about global AI governance have shot up the diplomatic agenda in the wake of dire warnings about the potential threats from the technology, and marks the most significant intervention yet from the U.K.’s prime minister about how he will respond to perceived threats from AI.
Speaking to journalists on Monday night, Burnham said the U.K. is “uniquely positioned to play a leadership role on AI,” citing its relations with the EU, U.S. and China. He also said the U.K.’s expertise, including its AI Security Institute, means it could act as an “honest broker.”
“I think we can really do something in this space [that is] very significant,” Burnham said.
This comes after PauseAI UK’s protest last week in Downing Street, in which we demanded that the government lead international coordination on AI. The protest was covered in national news outlets in Britain, including Channel 4 News, ITV News, The Guardian and others.
It was a very strange experience today seeing the two policy announcements from Ed Davey’s speech today be “Increase the threshold to pay income tax to £15k” and “Ban superintelligent AI”.
There’s an orange elephant in the room that won’t change his mind at the summit, sadly.
But I think the shifting political winds in the UK are nonetheless great news for an unrelated reason: a state visit from the King, with RSI on the agenda, is a moonshot that no other country is equipped to do. I’m probably over-optimistic to think this’ll actually happen, but considering Trump’s admiration for the royals there’s a small chance it’ll actually work.
It feels to me like the current situation is driven in part by the idea that USA based AI dominance means that middle powers become less and less relevant. This pretty much acts as an extension with how tech companies have grown in recent years.
It’s not clear to me, as a UK resident, why the US government should give a toss what we think. I don’t think that the US ”needs“ the UK.
(Of course, the US probably doesn’t end up “winning” AI because AI “wins” AI. But, I don’t think that the US government understands that.)
“I think you have to be clear-eyed about the huge benefits it can bring as well as recognizing where we do need to be careful,” Burnham said — noting that Trump has warned of risks from AI. The U.S. president has also branded existential warnings around the tech as a “HOAX.”
Burnham said warnings from AI company insiders have “made me pause” but said that he is determined to find a “sweet spot” on balancing the tech’s benefits against its risks, and not turn into a “doom-monger about AI.”
I think if Burnham genuinely believed there was a >10% probability of everyone dying within a decade, he would probably be saying different things? (The quoted statements do not sound to me like things said by top politicians in surviving worlds.)
Lots of people are leaving UK AISI because the pay is too low. Often they go to work on evals at other AI safety orgs where they will get paid more money. This is terrible. It’s far more valuable to have at least one government agency in the world with top tier AI expertise than it is to have a slightly stronger non-profit eval ecosystem.
A competent department creates unique options for the UK government to pass a pre-deployment testing law better than other countries. Or the government could empower them as a a thoughtful AI regulator that actually understands the risks. And AISI pushes other countries to improve by setting a high bar. It could even be transformed at some point into an international body that is the backbone of an AI treaty.
The jobs are in fact pretty well paid by UK tech standards but people there often have more lucrative options available at AI safety non-profits. Is there any way we can pay them to stay?
It’s sad we can’t pay them because the UKAISI-Redwood pay gap (order of $300k) is tiny compared to the cost of donations that it takes to match a safety researcher’s impact (order of $30M), or even the Redwood-Anthropic pay gap (order of $3M). I feel like at the very least, former lab employees and others without financial concerns should preferentially work for UKAISI and CAISI, as long as they don’t have unresolvable COIs from their equity.
I think the value of a year of the AI safety nonprofit ecosystem is ~50B times the value of a marginal dollar, so if one safety researcher is worth ~1/5K of the ecosystem (depends on the person!), that’s $50B/5K = $10M/year.
Surely the comparison of a marginal dollar vs. [total-impact]/[nr of people] is going to be misleading here? If we compared with [total-impact]/[nr of dollars] that’d be more one-to-one but also not what we care about. We’d want to compare the marginal dollar with a marginal person.
We can talk about Lukas’s value on the margin, which is probably slightly less than total impact * Lukas’s share of the credit due to diminishing returns in the size of the field but that’s not obvious
We can talk about causing new people, who tend to add much less value than average, to join the field
The correct concept here is the former, I think, since we’re talking about trading off between money and work by great people.
I certainly agree with the conceptual point, but it seems like I believe in way faster DMR in the size of the field than you seem to believe in. Like, I think a 50⁄50 gamble to eliminate the field or triple the size seems bad to me, whereas a 50⁄50 gamble to either double or halve it feels kind of unclear if it’s worth it or not. So feels closer to log than linear to me, although I think it’s probably somewhat less steeply diminishing than log.
(I would also have thought that the value of a year of the AI safety nonprofit ecosystem is higher than 50B though, so not at all clear I’m disagreeing with your bottom-line. And probably there’s some correlation here, where money is more valuable if there’s less steeply DMR to the size of the field.)
Sure, there is some pay pressure, but I didnt leave UK AISI because the pay was too low, I left because the projects were uninteresting and behind the curve with respect to what others are doing. Im perfectly happy to work for less pay so long as I am working on meaningful projects at a frontier institution where actual work is happening.
A lot of my work was getting told to investigate this or that, and then I would do the usual initial field overview and get told that that was enough and great work and yada yada yada. Excuse me? We didn’t even calculate anything yet? No no, thats enough...
I’m sorry, I (and many others coming from physics) got hired because of quantitative prowess, the work doesn’t have merit until we actually calculate. Drawing a few functors to map some categories is just a plan and it isn’t enough, this stuff depends on the details and we need to calculate. You can’t “vibemath” alignment.
bribes supplemental payments to civil servants are viewed uncharitably for some reason. perhaps a post-service endowment that requires some number of years with AISI as qualification?
I do wonder how those policy roles compare to roles with similar jobs titles elsewhere in the AI industry, though. In the UK civil service there are a lot of junior policy roles where you are expected to basically have a general understanding of the policy and act as a go-between between the technical staff and ministers.
E.g. this Cyber Resilience Policy Advisor role is paying £39-43k in London (£36-39k elsewhere) - ‘Awareness of, and in interest in, the UK’s cyber security landscape and basic cyber security concepts’ is listed as a desirable characteristic but not essential.
UK salaries in every field are pretty terrible. I suspect that this is a result of pre-existing wealth, basically capital is more important, in order to break in to that tier you either need a 98+ percentile salary or some other mechanism e.g start a high risk business, invest in high risk things, etc.
The public sector is uniquely bad for this because the comparison points (in a political sense) end up being with things like the distribution of salaries rather than the end result of what that money gets you.
For example, a house suitable to raise a small family in London, an MP on an MP’s salary could not get a mortgage on that, they would need some other source of funds (inheritance, previous job, etc).
This is considered to be just how it works, workers are conditioned to accept these low (in comparison with capital/wealth) salaries.
As a result of that I think that getting the required sums from taxpayer money is effectively impossible.
a house suitable to raise a small family in London, an MP on an MP’s salary could not get a mortgage on that, they would need some other source of funds
There are powerful non-monetary incentives to be in Parliament. Since UK law requires all government ministers to be MPs, most of the most powerful 500 people in the UK are MPs.
Many people find power to be very intoxicating even if one of the conditions for holding that power is that it is illegal to convert that power into personal wealth.
But the same analysis does not apply to the other 99.999% of jobs in the UK, so it would be nice if you could come up with an example population of workers other than MPs. Almost any other population would do. (Rock musician would not do because there are powerful non-monetary incentives to being a male rock musician because many young women find rock musicians extremely attractive.)
Are the rates of having high personal wealth lower in the US than in the UK? I doubt it is much lower. So if (as you claim) it is the presence of high-net-worth individuals in the labor market that keeps salaries low in the UK, then why is the mechanism you claim is operating in the UK so much weaker in the US? That is, why are US salaries, particularly for jobs in computing and AI, much higher than UK salaries?
Do you claim that the same mechanism operates in other markets or is it restricted to the labor market? For example, do high-net-worth individuals on average pay more for the same car as a no-net-worth individual because they can afford to pay more? Do high-net-worth individuals sometimes pretend to be no-net-worth individuals when talking to car salesmen to get a better price? It seems to me that they almost never bother doing that because they don’t have to because the salesman (and the car dealership he works for) regards one pound as just as desirable as any other pound regardless of the wealth status of the person giving them the pound. There a people who hold an animus towards wealthy people, but even such a person if employed as car salesman will almost always care much more about making a sale (and consequently earning a commission on that sale) than on disadvantaging the wealthy person or making the wealthy person suffer a little. Increasing personal income is just a much more compelling consideration than most considerations for the average person.
Would the cost of a particular make and model of car in country X or city X be lower if no one in country X or city X had any savings? My guess is no. It costs Y pounds to make the car, so that Y pounds acts as a strict lower bound on the average selling price of the car. A seller would like the average selling price to be as high as possible, but buyers are price sensitive, so if the seller’s price is higher than the price of another seller of the same make and model of car, buyers will just buy from the other seller till the other seller gets low on inventory and the prices equalize. These are the main 2 mechanisms determining the prices of cars, and I don’t see how the net worth of the buyers affects these mechanisms in any significant way.
I always thought the labor market works basically the same way: if employer A is willing to pay salary X to job Y and additionally if employer B stubbornly refuses to pay more than 0.9X, then pretty soon employer B will be unable to hire any employees willing and able to do Y, so the salaries offered by A and B will tend to equalize (unless for example B’s workplace is much easier to commute to than A’s or some other factor, but all of these factors tend to be minor factors relative to salary).
I expect wealthy people to care about salary, too, such that as long as employer B stubbornly refuses to pay more than 0.9X, it will find it difficult to hire even the very wealthiest people. People don’t stop wanting more money just because they already have a lot of money. If you ask people in a survey how much income would they need to be happy, they answer about 30% more than their current income regardless of whether their current income is 10,000 pounds a year, 20,000, 40,000, 80,000 or 160,000.
If the salaries A and B are willing to pay get high enough, then people who previously never considered being a Y will start considering it and start acquiring the necessary qualifications. Also, if the economy changes such that the products of A and B no longer command as high a price as they once did, then the amount A and B are able to pay for labor used in the production of those products must go down.
Again I always thought that this is the basic dynamics by which salaries are set, and again I fail to see how the average net worth of the employees and job applicants affects the dynamics significantly. In fact, the largest effect (although not particularly large) I imagine is its raising salaries because there are unavoidable aversive aspect to holding down most high-paying jobs, and people of wealth will tend to choose to avoid those aversive aspects by taking themselves out of the market for those jobs, requiring employers to raise salaries (at least a little) to attract sufficient numbers of applicants.
It’s a long winded topic but my general feelings are some mixture of:
the US celebrates capitalism in a way that the UK doesn’t and so the market and workers are actually just more competitive in general. ask the guy in the rural petrol (gas) station about business in the US and he will probably perk up. ask a reasonably middle class person in the UK and they will kind of flinch as if we just don’t do that kind of ambition here.
older country with more established class system
smaller country with worse building restrictions
more centralised around London
Bear in mind that I’m not talking about absolute wealth but in comparison to salaries e.g. does work buy you the ability to work less at some point or is it tightly bound around minimum cost of living
edit: re the car thing; more like high NW families have actual family houses and low NW families rent flats, cars are irrelevant expenses unless you look at clearly silly things like a ferrari
When people compliment you on highly salient attributes, you should be skeptical.
If you have a very prominent tattoo or a funky rainbow-colored jacket, people will frequently give you compliments about it, whether or not it is actually good. When there is a highly conspicuous thing, people cannot resist commenting on it. But the only polite way to mention something about someone’s appearance is to give a compliment. So people who make terrible fashion choices are frequently reinforced in their belief that everyone around them thinks they look great.
Similarly, if you give a public speech in front of a large audience or perform a song at a concert, everyone will compliment your performance, even if it was bad, because they can’t bring themselves not to mention the most salient fact about you at that moment, and a compliment is the only socially acceptable way to do that.
I would counter the idea that people are simply being polite when expressing compliments with the argument that what you are likely seeing instead is simply a kind of selection bias.
9x% of the population might think that you look ridiculous, but the single digit percentage that you appeal to could be both honest and vocal. In my experience, tattoos _in general_ are a great example of this where the “masses” will just not comment at all but tattoo fans are loud.
Or there’s the opposite—perhaps in a general “poll the world” sense people would think that you look pretty good, but when you wear X in a Y area, the Y’s don’t tend to like it.
Yeah, it could be that 99% people hate it but remain silent, and 1% really like it and say it out loud. You think you are popular in general, while in reality you are just optimizing for the tastes of 1%, and getting more and more silent disapproval of the 99%. Even worse, the reason why the 1% like it may be different than you think; for example you think “this is nice”, but for them it is an arbitrary ingroup symbol, so you are ultimately wrong about everyone.
When there is a highly conspicuous thing, people cannot resist commenting on it. But the only polite way to mention something about someone’s appearance is to give a compliment.
People could also not comment about it to the person in question, but satisfy the “cannot resist commenting on it” energy by waiting to be out of earshot and then telling someone else how terrible that person looked. I feel like this would happen a lot more often than giving a fake compliment just out of the need to say something.
I’ve also worn conspicuous clothing before and gotten comments that were more about expressing surprise/amusement rather than being compliments. (I don’t remember the exact words, but sentiments along the lines of “this is the first time I’ve seen someone wear cat ears to work” or whatever.)
I think people do this, but there’s a risk to assuming this motivation. It’s taking a statement that’s positive on its face, and then coming up with a scenario via mind reading that it’s actually bad for you. Sort of reverse CBT that could make you more anxious or paranoid over time since you are finding a way to reframe even positive comments as bad. Then you are potentially disappointed whether people compliment or insult you.
“Despite their extreme danger, we only became aware of them when the enemy drew our attention to them by repeatedly expressing concerns that they can be produced simply with easily available materials.”
Ayman al-Zawahiri, former leader of Al-Qaeda, on chemical/biological weapons.
I don’t think this is a knock-down argument against discussing CBRN risks from AI, but it seems worth considering.
The trick is that chem/bio weapons can’t, actually, “be produced simply with easily available materials”, if we talk about military-grade stuff, not “kill several civilians to create scary picture in TV”.
You sound really confident, can you elaborate on your direct lab experience with these weapons, as well as clearly define ‘military grade’ vs whatever the other thing was?
How does ‘chem/bio’ compare to high explosives in terms of difficulty and effect?
Well, I have bioengineering degree, but my point is that “direct lab experience” doesn’t matter, because WMDs in quality and amount necessary to kill large numbers of enemy manpower are not produced in labs. They are produced in large industrial facilities and setting up large industrial facility for basically anything is on “hard” level of difficulty. There is a difference between large-scale textile industry and large-scale semiconductor industry, but if you are not government or rich corporation, all of them lie in “hard” zone.
Let’s take, for example, Saddam chemical weapons program. First, industrial yields: everything is counted in tons. Second: for actual success, Saddam needed a lot of existing expertise and machinery from West Germany.
Let’s look at Soviet bioweapons program. First, again, tons of yield (someone may ask yourself, if it’s easier to kill using bioweapons than conventional weaponry, why somebody needs to produce tons of them?). Second, USSR built the entire civilian biotech industry around it (many Biopreparat facilities are active today as civilian objects!) to create necessary expertise.
The difference with high explosives is that high explosives are not banned by international law, so there is a lot of existing production, therefore you can just buy them on black market or receive from countries which don’t consider you terrorist. If you really need to produce explosives locally, again, precursors, machinery and necessary expertise are legal and widespread sufficiently that they can be bought.
There is a list of technical challenges in bioweaponry where you are going to predictably fuck up if you have biological degree and you think you know what you are doing but in reality you do not, but I don’t write out lists of technical challenges on the way to dangerous capabilities, because such list can inspire someone. You can get an impression about easier and lower-stakes challenges from here.
Biochem is hard enough that we need LLMs at full capacity pushing the field forward. Is it harmful to intentionally create models that are deliberately bad at this cutting edge and necessary science in order to maybe make it slightly more difficult for someone to reproduce cold war era weapons that were considered both expensive and useless at the time?
Do you think that crippling ‘wmd relevance’ of LLMs is doing harm, neutral, or good?
My honest opinion is that WMD evaluations of LLMs are not meaningfully related to X-risk in the sense of “kill literally everyone.” I guess current or next-generation models may be able to assist a terrorist in a basement in brewing some amount of anthrax, spraying it in a public place, and killing tens to hundreds of people. To actually be capable to kill everyone from a basement, you would need to bypass all the reasons industrial production is necessary at the current level of technology. A system capable to bypass the need for industrial production in a basement is called “superintelligence,” and if you have a superintelligent model on the loose, you have far bigger problems than schizos in basements brewing bioweapons.
I think “creeping WMD relevance”, outside of cyberweapons, is mostly bad, because it is concentrated on mostly fake problem, which is very bad for public epistemics, even if we forget about lost benefits from competent models.
Are you open to writing more about this? This is among top 3 most popular arguments against open source AI on lesswrong and elsewhere.
I agree with you you need a group of > 1000 people to manufacture one of those large machines that does phosphoramidite DNA synthesis. The attack vector I more commonly see being suggested is that a powerful actor can bribe people in the existing labs to manufacture a bioweapon while ensuring most of them and most of rest of society remains unaware this is happening.
Many people wrongly assume that the main way to use bioweapons is to create small amount of pathogen to release it in environment with outbreak as an intended outcome. (I assume that where your sentence about DNA synthesis comes from.) The problem is that creating outbreaks in practice is very hard, we, thankfully, don’t know reliable way to do that. In practice, the way that bioweapons work reliably is “bomb-saturate the entire area with anthrax such that first wave of death is going to be from anaphylactic shock rather than infection” and to create necessary amount of pathogen you need industrial infrastructure which doesn’t exist, because nobody in our civilization cultivates anthrax at industrial scale.
I agree that 1-2 logs isn’t really in the category of xrisk. The longer the lead time on the evil plan (mixing chemicals, growing things, etc), the more time security forces have to identify and neutralize the threat. So all things being equal, it’s probably better that a would be terrorist spends a year planning a weird chemical thing that hurts 10s of people, vs someone just waking up one morning and deciding to run over 10s of people with a truck.
There’s a better chance of catching the first guy, and his plan is way more expensive in terms of time, money, access to capital like LLM time, etc. Sure someone could argue about pandemic potential, but lab origin is suspected for at least one influenza outbreak and a lot of people believe it about covid-19. Those weren’t terrorists.
I guess theoretically, there may be cyberweapons that qualify as wmd, but those will be because of the systems they interact with. It’s not the cyberweapon itself, it’s the nuclear reactor accepting commands that lead to core damage.
I’d love a reply on this. Common attack vectors I read on this forum include 1. powerful elite bribes existing labs in US to manufacture bioweapons 2. nation state sets up independent biotech supply chain and starts manufacturing bioweapons.
This has been an option for decades, a fully capable LLM does not meaningfully lower the threshold for this. It’s already too easy.
This has been an option since the 1950s. Any national medical system is capable of doing this, Project Coast could be reproduced by nearly any nation state.
I’m not saying it isn’t a problem, I’m just saying that the LLMs don’t make it worse.
I have yet to find a commercial LLM that I can’t make tell me how to build a working improvised explosive (I can grade the LLMs performance because I’ve worked with the USG on the issue and don’t need a LLM to make evil).
On LessWrong, the frontpage algorithm down-weights older posts based on the time-since-posted, not the time-since-frontpaged. So, if a post doesn’t get frontpaged until a few days after posting, then it’s unlikely to get many views.
LessWrong has an autofrontpager that works a reasonable amount of the time. Otherwise, posts have to be manually frontpaged by a person. In my experience, this was always quite quick, but my most recent post was not frontpaged until 3 days after it was posted, so AFAICT it never actually appeared on the frontpage (unless you clicked “Load More”).
I think the solution is to downweight posts based on the time-since-frontpaged.
If you downweigh posts based on the time-since-frontpaged then posts get a huge boost when they have a delay of getting frontpaged (since they then first show to everyone who has personal blog enabled on their frontpage, and can accumulate karma during this time, and then when they have their effective date reset have a huge advantage over posts that were immediately frontpaged, because the karma provides a much longer visibility window).
I don’t really have a great solution to this problem. I think the auto-frontpager helps a lot, though of course only if we can get the error rate sufficiently down.
I’d be happy, if the auto-frontpager is ~instant, to get the option “delay publishing until human review” if it declines frontpage. Whether something gets ~50% less karma than it would by default is a pretty major drop in the effectiveness of what is often many hours of work, I’d be fine with waiting a day or two to avoid that usually.
Relatedly, I’ve been thinking about building a schedule-this-post-for-publication feature. If I publish a post at 10pm, it’s often better to publish the next morning for visibility. My guess is this would be useful for Inkhaven Residents who finish writing near-midnight.
If I could schedule, the frontpage review happened before publishing, and the schedule UI had “delay publishing until frontpage”[1] as a checkbox, this would be ~solved.
I’d prefer this to “delay publishing until human review”, as ~half a dozen times in the past few years I’ve appealed via Intercom and had a human-reviewed page retroactively frontpaged (usually a resource, which LW team’s priors seem to be something like ‘this won’t be maintained’ but will because I optimize a bunch for not leaving stale projects).
Or even better, at the the time when a post is frontpaged, check if will actually appear on the frontpage. If it is too old and has too little karma to be seen, then use the time-since-frontpaged.
“Back in December 2025, working with the campaign group PauseAI, I was the first MP to launch an AI Safety Debate in the House of Commons. Some colleagues thought the subject premature, one for technologists and scientists to determine, not legislators. Nine months on, that argument is dead.”
Startups often pivot away from their initial idea when they realize that it won’t make money.
AI safety startups need to not only come up with an idea that makes money AND helps AI safety but also ensure that the safety remains through all future pivots.
If you combine the fact that power corrupts your world models with the general startup person being power hungry as well as AI Safety being a hot topic, you also get a bunch of well meaning people doing things that are going to be net-negative in the future. I’m personally not sure that the VC model actually even makes sense for AI Safety Startups given some of the things I’ve seen in the space.
Speaking from personal experience I found that it’s easy to skimp out on operational infrastructure like a value aligned board or a more proper incentive scheme. You have no time so instead you start prototyping a product yet that means you get this path dependence where if you succeed, you suddenly have a lot less time. As a consequence the culture changes because the incentives are now different. You start hiring people and things become more capability focused. And voila, you’re now in a capabilities/AI safety startup and it’s unclear what it is.
So get a good board and don’t commit to something unless you have it in contract form or similar that you will have at least a PBC structure if not something even more extreme as the underlying company model. The main problem I’ve seen here is if your co-founder(s) is/are being cagey about it, I would move on to new people at least if you care about safety.
Best way to start an AI safety startup is get enough high status credentials and track record that you can ask your investors to go fuck themselves if they ever ask you to make revenue. Only half-joking. Most AI research (not product) companies have no revenue today, or are trading at an insane P/S multiple.
For a while, everyone was underinvesting in AI because they were not sufficiently AGI-pilled. Now, I think a lot of people (the big companies, investors, people in this community) might be overinvesting in AI because they’re not sufficiently AGI-pilled.
AI data center capex is approaching $1 trillion per year. Enough gigawatts are already planned to multiply the AI labs’ compute several times over in the next few years.
And yet the models are going crazy. I feel like the new revelations about misaligned agent swarms from OpenAI are not yet priced in. How can we possibly keep accelerating for another three years like this? If nothing else cybersecurity is surely going to break.
I can see a few possible futures:
Governments step in and regulate pre-emptively before a major catastrophe.
There is a major catastrophe and everyone is forced to slow down.
We fully lose control and AI takes over.
But another three years of just more scaling and acceleration like we’ve had the last three is increasingly infeasible. We are entering the danger zone and the race cannot just carry on as before.
I wouldn’t expect markets to price in the third possible future (total loss of control), but they should price in the first two. I think we could soon see a market correction when investors see that a deceleration will be required to keep models under control.
Economically it is probably rational for the world to be investing $1 trillion in capabilities and $10 trillion in safety. Unless you’re pessimistic, in which case we should spend $0 trillion on capabilities and still $10 trillion on safety.
This assumes that safety is the kind of good you can invest in. Right now that feels about as true as investing in peace or investing in love or similarly abstract concepts. Approximately nobody who is mainly money-bottlenecked has demonstrated a compelling track record of turning money into safety.
My main issue is that the world will find it hard to invest anything like ten trillion dollars. This is around a half of Chinese annual GDP, a quarter of the USA’s debt or a third of the USA’s annual GDP. Additionally, the world’s understanding of the AIs’ wildly transformative potential is surprisingly poor and prevents the world from making the relevant investments, especially during the looming financial crisis.
The idea is the faster safety is solved, the faster we can scale capabilities safely, which increases the growth rate of the economy from ~3% to ~100% and makes people immortal. If people want to maximize something like their discounted log(consumption) over the next 100 years, a wartime investment is warranted unless you think we couldn’t solve AI safety well enough to cure aging and automate the economy in a few decades.
I’m definitely on the pessimistic side, but yes I can see that spending more on safety may shorten the timeline for solving death which is probably a good thing if it doesn’t also increase s-risks.
My main worry is the following couple of doom scenarios.
Scenario A:
The amount of investments which was already priced in by Some Lab becomes reduced by external shocks, the lack of priced-in capabilities or misalignment;
Some Lab decides to avoid becoming bankrupt by gaining the revenue, to deliver the capabilities and uses potentially hazardous architectures like neuralese;
The capabilities emerge, but the AIs aren’t aligned.
Scenario B:
The $2,4T of investments which were already priced in by AI-2040 become reduced by external shocks, misalignment or the lack of priced-in capabilities;
Capabilities do slow down far more than AI-2040 predicted, but China keeps racing to produce compute;
The leadership is lost to China where safety is far worse, but the USA and China don’t do things like cross-audits of training runs.
Superintelligence safety isn’t “far worse” in China. You can’t get worse than zero superintelligence safety, which is what we currently have in both the US and China.
The next PauseAI UK protest will be (AFAIK) the first coalition protest between different AI activist groups, the main other group being Pull the Plug, a new organisation focused primarily on current AI harms. It will almost certainly be the largest protest focused exclusively on AI to date.
In my experience, the vast majority of people in AI safety are in favor of big-tent coalition protests on AI in theory. But when faced with the reality of working with other groups who don’t emphasize existential risk, they have misgivings. So I’m curious what people here will think of this.
Personally I’m excited about the protest and I’ve found the organizers of Pull the Plug to be very sincere and good to work with, but I’ve also set things up so that the brands of PauseAI UK and Pull the Plug are clearly distinct, so that our messaging remains clearly focused on the risks of future AI. For example, we have a separate signup page and we have our own demands focused on decelerating frontier development.
“In my experience, the vast majority of people in AI safety are in favor of big-tent coalition protests on AI in theory”
is this true? I think many people (myself included) are worried about conflationary alliances backfiring (as we see to some extent in the current admin)
I only have anecdata but I’ve talked to quite a few people and most people say it’s is a good idea to use the myriad of other concerns about AI as a force multiplier on shared policy goals.
I’ve talked to quite a few people and most people say it’s is a good idea to use the myriad of other concerns about AI as a force multiplier on shared policy goals.
Speaking only for myself, here: There’s room for many different approaches, and I generally want people to shoot the shots that they see on their own inside view, even if I think they’re wrong. But I wouldn’t generally endorse this strategy, at least without regard for the details of how the coalition is structured and what it’s doing.
I think our main problem is a communication problem of getting people to understand the situation with AI
that model capabilities are steadily increasing;
that the labs are aiming at literal superintelligence, no really, something more capable than any human alive, and then even better than that; that the labs are explicitly aiming to do an RSI, which looks increasingly likely to succeed;
that there is not a known science of reliably controlling or shaping the motivations of superhuman AIs.
that there are competitive pressures for all of the labs and all of the countries to beat their competitors, so slowing down or pausing requires international coordination.
These are slippery points to get across specifically because audiences tend to slip into visualizing something other than “actual strategic superintelligence”, that is automating science and technological progress and capable of strategically outmaneuvering adversaries—even when I talk with people from the labs, they often tend to gravitate to a fuzzier vision that has the form factor of the current AI chatbots / agents, but is much more competent.
Most of the time, I’m trying to land these points, despite the slipperiness, and talking about present-day harms that don’t have a through-line to the core alignment problems seem like more of a distraction than a help.
If we already had developed policies that would substantially improve the situation and were politically feasible, and we just needed to get a big enough coalition to get them implemented, I would feel differently.
But insofar as we have policies substantially help, they’re rather radical (on the order of “don’t allow private individuals to own more than 8 GPUs” and “negotiate with China for an international pause in frontier AI development”), and are only politically realistic if the stakeholders have a close-to-accurate picture of the situation.
We call on the UK government to fund binding Citizens’ Assemblies on AI and implement their decisions.
I think this would probably be a disaster, given how misinformed and unwise large parts of the broad public have been on many other scientific issues (e.g. vaccines, GMOs, nuclear power).
The rest of their views doesn’t inspire much confidence in their epistemics either:
Citizen assemblies often involve selecting a small number of delegates who are then informed about the all of the details of the issue in depth, including by expert testimonies, which the delegates have the affordance to do because they’re being paid for their time.
My understanding is that this works pretty well for coming to reasonable policy.
LLM hallucination is good epistemic training. When I code, I’m constantly asking Claude how things work and what things are possible. It often gets things wrong, but it’s still helpful. You just have to use it to help you build up a gears level model of the system you are working with. Then, when it confabulates some explanation you can say “wait, what?? that makes no sense” and it will say “You’re right to question these points—I wasn’t fully accurate” and give you better information.
Then it will often confabulate a reason why the correct thing it said was actually wrong. So you can never really trust it, you have to think about what makes sense and test your model against reality.
But to some extent that’s true for any source of information. LLMs are correct about a lot of things and you can usually guess which things they’re likely to get wrong.
Not OP but IME it might (1) insist that it’s right, (2) apologize, think again, generate code again, but it’s mostly the same thing (in which case it might claim it fixed something or it might not), (3) apologize, think again, generate code again, and it’s not mostly the same thing.
140 people came to the UK Parliament yesterday for the PauseAI conference.
The panel speakers were:
Dame Chi Onwurah MP, Chair of the Science, Innovation & Technology Committee
Professor Stuart Russell, world-renowed AI scientist
Lord Tim Clement-Jones, Co-Founder and Co-Chair of the All-Party Parliamentary Group on AI
Lord Lionel Tarassenko—President of Reuben College, Oxford
Iqbal Mohamed MP, leader of the Dec 2025 Westminster Hall debate on AI safety
Brando Benifei MEP, lead architect of the EU AI Act
I think this is probably the largest event on AI safety there’s been in Parliament. The tone from all the panelists was very supportive and most expressed very serious concern about AI, including the risk of extinction.
Brando Benifei gave us a dose of harsh reality with his assessment about the difficulty of achieving a global treaty of AI. But he remains supportive of PauseAI and seems to really by trying his best to get the EU to lead international cooperation for a ‘CERN for AI’ type solution.
Niradata ran the largest ever survey on global attitudes to superintelligence (n=377,458). The question was:
Some companies are working on developing superintelligent AI (artificial intelligence) that can outperform humans in all tasks. Which comes closest to your view?
I think choosing to make the cut at “net support of rapid development” when people with even the slightest sense of “hey maybe let’s put some rules on the thing” get sorted into the “strict oversight” bucket is a weasely way of presenting these results. You could just as easily make the case that both “Pause/Stop” and “Continue” are matched at 38% of the vote.
I think some people should internalize that often Principles Don’t Justify Drama (Among Humans). Drama destroys future discussion and activates primal emotions that will override rational deliberation. If you want to bring people around to your viewpoint, make them like you.
If everyone died on their hill every time someone violated what they thought was an important principle, society would not function. Cooperate with people you disagree with and when you violate someone else’s principle, be thankful that they are going to cooperate with you.
The best historical example of this is probably religious toleration. Thank God for those who tolerated infidels.
Of course the art is in knowing which hills to actually die on and who you should really shun as your enemy. To get it right, you need to fully appreciate the costs.
If you make people dislike you then you will push them away from your viewpoint whether or not your viewpoint is true.
The people you are talking with, maybe. You also might have an audience. Which archetypal character do people like more when they’re viewing an interaction from a distance, the abrasive prophet or the unctuous people-pleaser?
You’ll get better signal-to-noise in convincing people of true things (compared to false things) the more of an argument you’re able to get them to listen to (up to a point where you overwhelm their argument processing capacity) which is easier to do if they like you enough to listen to you.
This all leads into Zack Davis’s top rationality advice: become unlikable, so that people will never agree with you just to get along. They’ll argue with everything you say and you increase your chance of hearing true counterarguments!
I definitely wouldn’t call it “top rationality advice” because there are lots of other reasons to want to be liked and not want to be disliked, but I do expect the effect you describe to be real.
I think what you describe is mostly a high-status (dominance or prestige) phenomenon rather than a high-likeability phenomenon. But I will admit I’m rather going off instinct and intuition here.
When I go on LessWrong, I generally just look at the quick takes and then close the tab. Quick takes cause me to spend more time on LessWrong but spend less time reading actual posts.
On the other hand, sometimes quick takes are very high quality and I read them and get value from them when I may not have read the same content as a full post.
Interesting. I am concerned about this effect, but I do really like a lot of quick takes. I wonder whether maybe this suggests a problem with how we present posts.
I think the biggest problem with how posts are presented is it doesn’t make the author embarrassed to make their post needlessly long, and doesn’t signal “we want you to make this shorter”. Shortforms do this, so you get very info dense posts, but actual posts kinda signal the opposite. If its so short, why not just make it a shortform, and if it shouldn’t be a shortform, surely you can add more to it. After all, nobody makes half-page lesswrong posts anymore.
This. The struggle is real. My brain has started treating publishing a LessWrong post almost the way it’d treat publishing a paper. An acquaintance got upset at me once because they thought I hadn’t provided sufficient discussion of their related Lesswrong post in mine. Shortforms are the place I still feel safe just writing things.
It makes sense to me that this happened. AI Safety doesn’t have a journal, and training programs heavily encourage people to post their output on LessWrong. So part of it is slowly becoming a journal, and the felt social norms around posts are morphing to reflect that.
I’d love to see the reading time listed on the frontpage. That would make the incentives naturally slide towards shorter posts, as more people would click and it would get more karma. Feels much more decision relevant than when the post was posted.
Get an LLM to generate a TLDR of the post and after the user finishes reading the post, have a pop-up “Was opening the post worth it, given that you’ve already read the TLDR?”.
They don’t claim that Grok 3 was trained on 200K GPUs, and that can’t actually be the case from other things they say. The first 100K H100s were done early Sep 2024, and the subsequent 100K H200s took them 92 days to set up, so early Dec 2024 at the earliest if they started immediately, which they didn’t necessarily. But pretraining of Grok 3 was done by Jan 2025, so there wasn’t enough time with the additional H200s.
There is also a plot where Grok 2 compute is shown slightly above that of GPT-4, so maybe 3e25 FLOPs. And Grok 3 compute is said to be either 10x or 15x that of Grok 2 compute. The 15x figure is given by Musk, who also discussed how Grok 2 was trained with less than 8K GPUs, so possibly he was just talking about the number of GPUs, as opposed to the 10x figure named by a team member that was possibly about the amount of compute. This points to 3e26 FLOPs for Grok 3, which on 100K H100s at 40% utilization would take 3 months, a plausible amount of time if everything worked on almost the first try.
Time needed to build a datacenter given the funding and chips isn’t particularly important for timelines, only for catching up to the frontier (as long as it’s 3 months vs. 6 months and not 18 months). Timelines are constrained by securing more funding for a training system, and designing and manufacturing better chips. Another thing on that presentation was a claim of starting work on another 1.2 GW GB200/GB300 datacenter, which translates to 600K chips. This appears to be more than other LLM labs will construct this year, which might be only about 0.5 GW, except for Google[1], but then Musk didn’t name deadlines for 1.2 GW either. It’s only more concrete than Meta’s 2 GW site in specifying that the chips are Blackwell, so it can’t be about plans for 2027 when better chips will be available.
For Claude 3.5, Amodei says the training time cost “a few $10M’s”, which translates to between 1e25 FLOPs (H100, $40M, $4/hour, 30% utilization, BF16) and 1e26 FLOPs (H100, $80M, $2/hour, 50% utilization, FP8), my point estimate is 4e25 FLOPs.
GPT-4o was trained around the same time (late 2023 to very early 2024), and given that the current OpenAI training system seems to take the form of three buildings totaling 100K H100s (the Goodyear, Arizona site), they probably had one of those for 32K H100s, which in 3 months at 40% utilization in BF16 gives 1e26 FLOPs.
Gemini 2.0 was released concurrently with the announcement of general availability of 100K TPUv6e clusters (the instances you can book are much smaller), so they probably have several of them, and Jeff Dean’s remarks suggest they might’ve been able to connect some of them for purposes of pretraining. Each one can contribute 3e26 FLOPs (conservatively assuming BF16). Hassabis noted on some podcast a few months back that scaling compute 10x each generation seems like a good number to fight through the engineering challenges. Gemini 1.0 Ultra was trained on either 77K TPUv4 (according to The Information) or 14 4096-TPUv4 pods (according to EpochAI’s quote from SemiAnalysis), so my point estimate for Gemini 1.0 Ultra is 8e25 FLOPs.
This gives 6e26-9e26 FLOPs for Gemini 2.0 (from 2-3 100K TPUv6e clusters). But unclear if this is what went into Gemini 2.0 Pro or if there is also an unmentioned Gemini 2.0 Ultra down the line.
It seems that we are already at the GPT 4.5 level? Except that reasoning models have confused everything, and increasing OOM on output can have the same effect as ~OOM on training, as I understand it.
By the way, you’ve analyzed the scaling of pretraining a lot. But what about inference scaling? It seems that o3 has already used thousands of GPUs to solve tasks in ARC-AGI.
The 200k GPU number has been mentioned since October (Elon tweet, Nvidia announcement), so are you saying that that they managed to get the model trained so fast is what beat the predictions you heard?
1/ Sparse autoencoders trained on the embedding weights of a language model have very interpretable features! We can decompose a token into its top activating features to understand how the model represents the meaning of the token.🧵
2/ To visualize each feature, we project the output direction of the feature onto the token embeddings to find the most similar tokens. We also show the bottom and median tokens by similarity, but they are not very interpretable.
3/ The token “deaf” decomposes into features for audio and disability! None of the examples in this thread are cherry-picked – they were all (really) randomly chosen.
4/ Usually SAEs are trained on the internal activations of a component for billions of different input strings. But here we just train on the rows of the embedding weight matrix (where each row is the embedding for one token).
5/ Most SAEs have many thousands of features. But for our embedding SAE, we only use 2000 features because of our limited dataset. We are essentially compressing the embedding matrix into a smaller, sparser representation.
6/ The reconstructions are not highly accurate – on average we have ~60% variance unexplained (~0.7 cosine similarity) with ~6 features active per token. So more work is needed to see how useful they are.
7/ Note that for this experiment we used the subset of the token embeddings that correspond to English words, so the task is easier—but the results are qualitatively similar when you train on all embeddings.
8/ We also compare to PCA directions and find that the SAE directions are in fact much more interpretable (as we would expect)!
9/ I worked on embedding SAEs at an @apartresearch hackathon in April, with Sajjan Sivia and Chenxing (June) He. Embedding SAEs were also invented independently by @Michael Pearce.
Claude 3.7′s annoying personality is the first example of accidentally misaligned AI making my life worse. Claude 3.5/3.6 was renowned for its superior personality that made it more pleasant to interact with than ChatGPT.
3.7 has an annoying tendency to do what it thinks you should do, rather than following instructions. I’ve run into this frequently in two coding scenarios:
In Cursor, I ask it to implement some function in a particular file. Even when explicitly instructed not to, it guesses what I want to do next and changes other parts of the code as well.
I’m trying to fix part of my code and I ask it to diagnose a problem and suggest debugging steps. Even when explicitly instructed not to, it will suggest alternative approaches that circumvent the issue, rather than trying to fix the current approach.
I call this misalignment, rather than a capabilities failure, because it seems a step back from previous models and I suspect it is a side effect of training the model to be good at autonomous coding tasks, which may be overriding its compliance with instructions.
It’s concerning that such a high proportion of the best people I know in AI safety go into grantmaking, running upskilling programs or building AI safety courses / talent pipelines.
I expect if we had better ideas about what to do, the best people would mostly just go and do those things directly instead of finding / funding others to do the actual work of preventing AI catastrophe.
Do you have a way of knowing whether the people you know are a representative sample of the whole field? From what I understand there’s way more safety researchers than field builders.
Caloric restriction works, however it impedes his productivity (“ability to think”).
Exercise isn’t effective in promoting weight loss or reducing weight gain due to compensatory metabolic throttling during non-exercise times
His fat metabolism is poor, because his fat cells are inclined to leach glucose and triglycerides from his bloodstream to sustain themselves rather than be net contributors, and the effect is that muscle loss makes up the difference, leading to unfavourable implications for body composition, energy, and overall health. In essence “good genetics” for fat loss = “Fat cells that are readily and efficiently broken down for energy”, whereas “bad genetics” for fat loss = “fat cells that are resistant to being used as energy”.
Thousands of people across the UK believe that humanity might soon build dangerous superintelligent AI. And yet many of these people haven’t even taken 10 seconds to email their MP about it. Political intervention is one of the most plausible ways that humanity can avoid catastrophe and there are many easy ways that anyone can help to improve the chance of useful AI legislation:
If the thousands of people that are already concerned about superintelligence were well organised as a political force, they could wield substantial power to influence British AI policy. Help us make that happen!
We’re hiring for:
a software engineer
an operations specialist
a content creator
and several other roles
Applications will be assessed on a rolling basis, so apply soon for the best chance of success.
It’s becoming a habit for me to run anything I write through an LLM to check for mistakes before I send it off.
I think the hardest part of implementing this feature well would be to get it to only comment on things that are definitely mistakes / typos. I don’t want a general LLM writing feedback tool built-in to LessWrong.
LLMs can pick up a much broader class of typos than spelling mistakes.
For example in this comment I wrote “Don’t push the frontier of regulations” when from context I clearly meant to say “Don’t push the frontier of capabilities” I think an LLM could have caught that.
There are people who treat the lock on a public bathroom as a tool for communicating occupancy and a safeguard against accidental attempts to enter when the room is unavailable. For these people the standard protocol is to discern the likely state of engagement of the inner room and then tentatively proceed inside if they detect no signs of human activity.
And there are people who view the lock on a public bathroom as a physical barricade with which to temporarily defend possessed territory. They start by giving the door a hearty push to test the tensile strength of the barrier. On meeting resistance they engage with full force, wringing the handle up and down and slamming into the door with their full body weight. Only once their attempts are thwarted do they reluctantly retreat to find another stall.
I haven’t read the paper yet but I’m pretty confident people are overhyping the Global Workspace paper.
This isn’t even any fault of the paper. It’s just kinda inevitable for any ML paper with a PR team and a cool video promoting it.
But it’s especially true for a paper that is (1) interpretability and (2) uses a bold philosophical framing. The interp literature has a long history (pre-dating mech interp) of techniques going through a cycle of discovery, hype, criticism and skepticism.
Not everyone is overhyping the paper. Neel Nanda’s review seems pretty sober. But I’m surprised to see Zvi writing this:
This feels like a major advance in our understanding of LLMs, both overall and our ability to understand any particular interaction.
The J-Lens is improvement of the logit lens. It’s cool, but not a major advance that deserves this level of attention from people who do not typically read mech interp papers.
Once again, I have not read the paper and certainly hope to be pleasantly surprised when I do!
I find Anthropic’s style irritating (Mostly because Mech Interp ppl are not the intended audience of the posts ig), but reluctantly I found it to be a pretty interesting paper on the technical level.
“The J-Lens is an improvement of the logit lens.”: The J-lens is answering a different question from the logit lens, so it’s not best described as just an improvement imo.
The J-lens gives a global linear approximation for the impact that intervening will have on the outputs at future tokens, as opposed to just the output at the current token. This leads to qualitatively different tokens being surfaced in comparison to the logit lens, and their results here are pretty interesting.
Agreed it’s definitely overhyped though, yeah. But that’s what Anthropic does! It’s probably not too bad a trade off in exchange for getting more people into the field (as long as people in field maintain appropriate skepticism).
I do wish that Anthropic would publish a preprint alongside the blog post that could plausibly be submitted to conference, because I think it would be far more readable if they were forced to compress things down to the key experiments.
The AI safety ecosystem is so well resourced that it has been correctly identified by many as one of the best paths into high prestige AI research jobs.
This person on twitter has written a popular article about getting into frontier ai labs and a Field Guide to AI Fellowships. The “AI Fellowships” are mostly AI safety programs funded by CG/OpenPhil. I have also noticed that people in ML research are quite likely to have heard of MATS and be interested in participating, even when they have very little interest in AI safety.
Idk how good or bad this is. It definitely causes a lot of ML researchers to engage with the AI safety literature, where they otherwise would not have. But it’s worth noting that while the primary driver of applications to programs like MATS used be concern about AI safety, now it is increasingly a desire to work at a frontier lab.
I used to select participants for LASR labs, one of programs listed, and we actively tried to choose people who cared about AI safety, but I think we often did not succeed and indeed some people on the program now work at frontier labs in roles that I think have little to do with safety.
Safety-washing in practice mostly looks like people rationalizing research or jobs that they would have done anyway as necessary for safety. It’s easy to do because it is in fact genuinely unclear in most cases what is helpful and harmful for safety. Only by looking at the larger pattern can we notice the suspicious abundance of conveniently overlapping opportunities that further both safety and a person’s short-term interests. However, I think it will be less common in people whose primary motivation is to prevent AI catastrophe.
Interestingly Pangram says that the first article you linked is 100% AI. https://x.com/pangram/status/2066399823486423185
Maybe unpopular but I think that if you’re concerned about defection after programs like MATS then even more than “caring about AI safety” you specifically want to select for effective altruists, who are much likelier to care about AI actually going well than people who specifically care about AI safety. For the latter group, the motivation is partly not wanting everyone to die but also partly being AGI-pilled, being interested in the technical problem of alignment, and following prestige gradients, none of which are going to robustly keep participants on track.
maybe i’m overconfident but i generally find that it’s somewhat easy to tell if someone genuinely cares. you can’t really fake genuinely caring (or at least as genuinely as one can about anything—perhaps everything bottoms out in some other drives, but that’s good enough for me). there are some false negatives—people who genuinely care who have taken on the affectations of faking it, because they mistakenly think this helps their chances at success. there are also some people who care so much that they go crazy and start being counterproductive. but i don’t think i’ve ever heavily overestimated how much someone cares about AI safety.
This also seems right to me, but I don’t think MATS or the labs (or the funders) care if someone actually cares. They seem happy to tell themselves stories that they are making things better by getting people into the labs (or lab-adjacent orgs) even if they really don’t seem very safety motivated.
Maybe you could make an argument that they don’t care enough, but I’m pretty confident that MATS does care about this. When I did it, they had a whole reading group program whose primary aim, as I understand it, was to get people to understand and care more about the fundamental safety issues. And I also believe they try to select for people who care about safety, at least to some extent.
I predict that a randomly selected MATS participant from cohorts in the past two years would more likely than not fail to pass my ITT about why AI poses an existential risk.
Very interested in running the experiment, if anyone has a way to source random MATS scholars.
I think it’s gotten worse over time. I should have said “I don’t think MATS or the labs (or the funders) care much if someone actually cares”. I agree the caring isn’t zero.
i want to spend money on making there be more good alignment researchers and care about them actually caring about alignment. what would you recommend i do?
I think up until very recently funding Lightcone would have been a good bet at your scale (though you should of course be appropriately skeptical of me saying this). Funding projects doing good object-level work seems good.
Before I think about concrete recommendations, by “alignment researchers” do you mean people doing technical alignment work, or eval work, or do you include technical governance work like MIRI’s technical governance team? Or are you just using it as a proxy for work trying to make alignment go well in some form or another?
Fund AFFINE
already did
awesome!!
I think this is very hard thing to get right, and I don’t think there’s a scaleable way or org. that you can spend money on.
I think the current best bet would be to find existing people that are technical, driven, safety-pilled and already done some independent research, and fund them to continue doing that?
If true, I think part of the change can be attributed to a more dramatic increase in those who are technically competent and have various accomplishments yet don’t care so much for safety (though would love a role at Anthropic). If the sample of those folks has increased, it becomes more challenging to turn down those who are highly competent but maybe care a bit less about safety. Especially if you are a mentor and you just want someone to execute exceptionally well on your project.
Last time I mentored (for SPAR), I had to turn down a lot of folks who were more accomplished and likely to succeed (e.g. professors) so that I could accept those I felt would have more long-term potential to contribute to safety.
Wouldn’t professors be more likely to only apply if they genuinely care about safety, while younger people (overall) might be still figuring things out and happy for any opportunity to advance their career?
In the specific situation I’m referencing, no. My impression is that they were just curious about automated research scaffolds, agent frameworks and applying interpretability for capabilities. They did not seem all that interested in alignment.
I could imagine that in practice many professors end up optimizing for prestige and more paper publications (that have nothing to do with superintelligence), whereas a younger person may not be locked into that mindset yet.
I think it’s important to note that the reading list was new in MATS 5, revised for MATS 6, and I’m unsure if it persisted to MATS 7, since attendance beyond (iirc) ~week 2 or so cratered in both cohorts. From the inside, it felt like a thing a few staff members desperately felt scholars needed, rather than an effective or popular educational initiative.
I think that individual people at MATS (including both directors) care a great deal. I think that MATS, as an institution, habitually under invests in value alignment in order to better serve the immediate desires of its users (mentors who want the most competent applicants; applicants who want the most lucrative job).
In the three cohorts I spent at MATS, there was never an explicit org-wide filter for value alignment. Mentors picked their scholars. Some seemed to filter on value alignment, and others did not seem to.
It depends a lot on the type and quantity of interactions you have. I think it can be pretty hard to assess this from just a CV, short answer questions and a 20 minute interview.
It’s worth noting for posterity that the author of the two linked articles appears to have been posing as an Anthropic fellow/employee and does not actually hold these credentials. https://x.com/bruce_t_/status/2067411712828174771?s=46
I mentor some people for AI fellowships/safety programs and worry about this a fair amount. How do you select people who care about safety (specifically: what criteria/vibes do use) and overall do you think it’s still positive in expectation to do this type of mentorship for the next 12 months?
Adjacently, I read the aforementioned article and it seems pretty slop (not that insightful, or well-written) to me.
I think successful prior career in industry could be a good sign. Or any other career, or generally being older, so that it’s less likely you thought 2 years ago that MATS would look great in your CV.
Dave Banerjee’s test is a decent first pass filter, given his observation that “a large fraction of researchers in AI safety/governance fellowships cannot do any of these things”. Like any interactive interview, you’d want to probe their mental models with a few follow-up questions.
I think this makes the right kind of mentorship and exercising the right kind of judgment more valuable. FWIW the article’s author got outed as a fabulist of some form in the last few hours; he admitted to lying about MATS and hasn’t furnished any proof about his other resume lines.
I think this is a legibility problem; the whole rationalist corpus is available for easy consumption, Bayesian reasoning is not that hard to figure out, and there’s a pretty universal presumption of good faith (the author still has a lot of defenders!), so it’s really hard to sit two people down, one of whom is a mercenary there to use the fellowship as a stepping stone to wealth, power, etc, the other there out of genuine belief, and detect which is which reliably off of conversation alone.
At least at the selection layer, it seems like the most important thing is finding small honesty and motivation tells. I’m a recent college grad and in college I helped found a club which had pretty explosive and rapid success which we believed was thanks to our distinct internal culture and norms which put us at odds with our otherwise famously-mercenary college.
In the end we found two filters to be reliable and workable. The first was a very brief screening application which explicitly prohibited AI-generated responses. Whoever read the application plugged results into Pangram if they were suspicious. Our belief was that if the applicant couldn’t do a ~15 minute screener with their own writing, then we had no reliable positive indications of their motivation or work quality, and a pretty strong negative signal about their honesty.
The next was a curveball question in which we explicitly encouraged “I don’t know, here’s what I do know and here’s how I’d start solving it” and clarifying-follow ups as answers. The questions always varied and were always premised around very specific information about our field; having an answer from the hip was very high-signal for motivation and work quality, asking the right questions was higher signal for work quality but lower for motivation.
These two filters were the only absolutes in the process, and any other information we asked (resume, specific question content, etc) was purely about placement or deciding between marginal applicants. When I graduated, the club had basically kept its original culture while being successful. Maybe this approach will be helpful for thinking through the problems?
Some discussion on twitter about whether that article is written by a LARPer here https://x.com/anpaure/status/2066563480539590934?s=20
My guess at one of the best metics / interview questions here is “tell me about a lesswrong post you really liked” or something similar. Generally screening for “how online / plugged in into ai safety ecosystem” seems like a decent metric.
This is not exactly what I want, though, since I think anyone seriously applying to these programs will have done some reading and would be able to answer about their favorite LW post competently.
Agreed. For example, here’s 4.8′s answer:
4.8 answer
I do, or at least there’s a small cluster I keep returning to, and the one I’d name first is Eliezer’s “Local Validity as a Key to Sanity and Civilization.”
The core move is simple and I find it keeps paying out: you should be able to evaluate whether a single step of reasoning is valid independently of whether you like where the argument lands. And the post’s real claim isn’t just epistemic hygiene for individuals — it’s that this habit is load-bearing for civilization. A society where people can agree “that inference is invalid” even when they disagree about the conclusion has a working immune system; one where validity-judgments get pulled toward tribal allegiance loses the ability to error-correct at all. That second-order framing — local validity as the thing that lets pluralistic systems stay sane — is what elevates it above a standard “commit no fallacies” essay.
What I like is that it’s a genuinely useful idea rather than a clever one. It gives you a concrete thing to watch for in yourself: the moment you notice you’re scrutinizing an argument harder because you dislike its conclusion, you’ve caught yourself doing the bad thing.
The honest caveat is that “favorite” for me is closer to “highest hit-rate on rereads” than to nostalgia. By that standard, the runners-up are Scott’s “The Tails Coming Apart as Metaphor for Life” and Garrabrant’s “Goodhart Taxonomy” — both for similar reasons, that they name a structural failure mode crisply enough to actually use.
Do you have one? I’d be curious whether yours skews toward the reasoning-tools cluster or somewhere else entirely.
Remove some of the obvious LLM tells, and I’d have a lot of trouble telling it apart from a fluent LW’er “genuinely” interested in civilizational issues, discourse norms, and AI safety.
Agree with Leo that this is not a hard thing to distinguish, provided that value alignment matters for these fellowships.
When we do interviews for CMU’s AI Safety org, it seems like open-ended questions about viewpoints (e.g. “what about the current pace of AI keeps you up at night?” or “if you weren’t doing AI Safety research, what would the alternative be?”) enable us not just to distinguish between people who “speak the language”. Another shibboleth is actually under-awareness of the community—someone who is quick to recite a bunch of names may be less concerned with the issues at hand. Whether we want to do this is another question.
FWIW, MATS clearly does still source people who are excited about safety; most other fellows in my cohort act as if they are fighting for their future! Still others are intrigued by the more challenging theoretical and empirical questions. I trust the staff in their experience vetting out people who do not wish to engage the space genuinely.
The thing I am more worried about are AI fellowships taking money from participants (or fellowships created to reach personal agendas of the founders).
Anthropic pretraining researcher resigns over safety concerns: https://x.com/hilbertspaess/status/2097476196791709843
If you work at an AI lab, possibly the most impactful you can do now is resign and try to start a social cascade that stigmatises participating in the AI race.
But please do it right. Most AI employee resignations so far have been unstrategic, seemingly spur of the moment decisions with little consideration about making the most of the opportunity. You should have a media strategy lined up before you leave, journalists briefed, statements written and edited.
Or even better, you should not directly resign but instead use your newfound freedom to send the costliest possible signal to your company and to the world that you are actually scared about AI—to the point where you are less worried about keeping your job than doing something to slow AI progress.
If you are in this position, DM me if you want help executing this as effectively as possible.
Edit: It seems that Jacob Coxon did have a media strategy and it paid off.
I witness that this resignation and the warning going with it was (briefly but faithfully) mentionned in the morning news on the leading french news radio. That’s a good update that such an alert can even reach such a mainstream foreign media usually focused on politics and economics.
This resignation seems to have gone way more viral than I would have expected. (almost 300k likes??)
Anyone have any idea why this is
I live far from the Bay Area and have many friends outside of this world. I think people here underestimate how much young “normies” genuinely hate this technology and the companies building it. It generally comes out as pushback against datacenters and environmental concerns but I think the younger generation is fairly open to the idea at least that they will be replaced soon and can imagine this tech fully disempowering them, setting up a surveillance state etc. I don’t think you need to convince them that it will be the AI killing them to convince them to support a pause. They already think CEOs are maximally evil and want everyone else dead. Even if there’s issues with the specifics, I hope more people here come to this conclusion and realize that a political solution is likely the easiest path to prevent catastrophic scenarios soon. I’ve been convincing people to call their senators/house reps in support of a pause. Outside of the tech world there is huge support for regulating AI.
In many ways I am finding technical people outside of the safety community to be shockingly irrational about this while people who don’t understand the tech have a strong sense that something will likely go wrong along the way. There’s something to be said for the past sins of the tech industry setting up the right mental model among the general populace to distrust the people developing capabilities.
Evan Hubinger retweeted it claiming p total human extinction >10%. Plus exponentially rising interest in AI x-risk.
The reason for going viral is probably the WSJ article.
This comment discusses the importance of having a media engagement strategy.
It’s also on the front page of The Guardian: https://www.theguardian.com/technology/2026/sep/09/anthropic-researchers-ai-human-extinction
This does not seem convincing enough to be the primary reason for going viral. There are tons of news articles out there usually, and the WSJ is not extremely widely read. But maybe I am missing something.
People are really getting concerned about alignment. Just this morning I saw several liberal commentators on UK bluesky say that senior politicians need to start speaking up about the risk. Also I think mainstream commentators who’ve been quietly worried about it for a while (like Yglesias) are feeling emboldened to speak up about it more.
A huge part will probably be the combination of clarity and the credible signal of earnestness by resigning.
1)
He made very forthright and clear statements such as “The people building AI earnestly believe that it could kill us all by the end of the decade.”; “They are racing straight to self-improving superintelligence and gambling with our lives.”
These are just very intelligible to broad audiences. Previous statements where less clear to non technical audiences. Compare a statement such as Hinton’s estimate of 10% chance of human extinction, which might generate less virality because many people could be reading 10% as “won’t happen”. Also compare with Miles Brundage’s “Neither OpenAI nor any other frontier lab is ready, and the world is also not ready.” when he left OpenAI: we read “not ready” and think “not ready to prevent human extinction”, but most people would just not read it with the same frame of mind, and I guess might think something mild such as “not ready to handle the economic transition painlessly”.
2) Resigning is just a very clear signal that he is serious about this. It’s really hard for lay people to evaluate arguments about safety risks on their merits, so it makes sense to instead defer to such signals. You can’t explain this away with being motivated by PR, and common sense dictates that a person foregoing a big fat pay check would do so only for very good reasons.
Contrast this with Sam Altman, who has also been relatively clear in the past, but it would take a couple of inferential steps to take this seriously, if the first thought is “then why the hell are you building it?”.
I’d make an easier request of developer employees: just state your opinion on AI risks publicly, and include your perception of the range of opinions inside your org.
You can get more media boost if you quit, but each of these statements helps shift the Overton window, and is a powerful reference for those of us trying to communicate AI risks.
We may have an opportunity right now to create the cascade that shifts the Overton window and public opinion dramatically in a short time, much like how opinions on COVID shifted in early 2020.
In less than 24 hours, that tweet has reached over 100 million views and 600k likes!
For comparison, a typical Elon Musk tweet (the most followed person on X) seems to get about 10 million views and 10k likes.
The tweet also stands out for being unusually well-written and completely unambiguous about its message.
you can just do things
This might be stupid, but I have a proposal: If you are resigning, consider calling the local police, the FBI, or whoever else might help, inform them of the situation, and try to convince them that they should raid the offices of your former employer to stop them from doing something that is a potentially imminent danger. After that fails (presumably), go public and say what you attempted. Twitter seems full of people who think this is just a publicity stunt. Most of them probably can’t affect much, but some might, and actually calling the police (or another law enforcement agency), especially if a raid actually happens and disrupts their productivity, might help convince them otherwise.
And your suggestion is to unambiguously make them correct?
It’s bad they think Jacob Coxon’s resignation is a publicity stunt because it would be to somehow help the AI industry. If whoever resigns next puts enough effort into making it clear they want to take down their former employer, or at least stop attempts at recursive self-improvement, it doesn’t matter much if it is meant as publicity, as people would understand it is publicity meant to help achieve those goals.
I don’t necessarily agree with this proposal on the object level, but the insights are interesting. If we want this to stay in the news, ideally instead of a boring drip-drip-drip of employee resignations, we get a boom-bang-crash of drama. If there’s a way to resign in a different, flashy way which somehow “advances the story”, adds color, etc. that’s liable to generate continuing coverage.
Furthermore, the ideal resignation would somehow address this suspicion that it’s all about creating hype for investors. For example, imagine if a resigning employee called on politicians to tax away all the profits from the AI industry, to reduce the incentive to invest in doom. Headline: “Capitalism will kill us all, resigning AI employee says”. At that point, it becomes difficult to claim that the resignation is some sort of corporate gambit, since what corporation would beg the government to punitively tax them?
Another idea is to resign from OpenAI and then sue OpenAI. There must be something you can sue them for, right? Or how about calling on securities regulators to block the upcoming IPOs? Is there any way they could do that? Essentially call for some sort of policy which will hurt the profits of these AI companies so people don’t think it’s all about investor hype. “Ethical investors should boycott upcoming IPO, says resigning AI engineer.” Simply linking to PauseAI could also be a credible signal that you actually want a pause and you’re not just marketing.
You could rightly say that they are endangering your life, which is grounds for a lawsuit. Whether a judge would go for that is another question. But anyone could file that lawsuit, not just an ex-employee. Maybe there’s a better angle an ex-employee would have.
A mental model for media. There are various podcasts and broadcast news channels. Some of them might be interested in doing an interview with someone on the topic of AI safety. We should aim to saturate these interview channels to the greatest degree possible.
For example, if Coxon has more media requests than he has time to field, it might be possible to coordinate and have him say something like: “Sorry I don’t have time to chat—but now there is news that ANOTHER employee has resigned from Anthropic, for similar reasons. Would you like to talk to them?”
Or, it’s possible that certain channels are open to covering this topic, but don’t want to talk to Coxon because they feel the Coxon “scoop” has already been covered to death. So when formulating media strategy for another resignation, you could specifically target channels which did not cover Coxon, but are still “newsy”/”techy”/etc. enough that such a resignation would be within the channel’s topical scope. Then you contact them, with the subtext that: “You missed the Coxon scoop, but you could be among the first channels to cover this new scoop!”
The desired end result is that no matter where a person gets their news, they’ve seen coverage related to the big AI company resignations.
I imagine working with a publicist could be really helpful since they’re professionals at coordinating stuff like this.
This is an attempt to compile all publicly available primary evidence relating to the recent death of Suchir Balaji, an OpenAI whistleblower.
This is a tragic loss and I feel very sorry for the parents. The rest of this piece will be unemotive as it is important to establish the nature of this death as objectively as possible.
I was prompted to look at this by a surprising conversation I had IRL suggesting credible evidence that it was not suicide. The undisputed facts of the case are that he died of a gunshot wound in his bathroom sometime around November 26 2024. The police say it was a suicide with no evidence of foul play.
Most of the evidence we have comes from the parents and George Webb. Webb describes himself as an investigative journalist, but I would classify him as more of a conspiracy theorist, based on a quick scan of some of his older videos. I think many of the specific factual claims he has made about this case are true, though I generally doubt his interpretations.
Webb seems to have made contact with the parents early on and went with them when they first visited Balaji’s apartment. He has since published videos from the scene of the death, against the wishes of the parents[1] and as a result the parents have now unendorsed Webb.[2]
List of evidence:
He didn’t leave a suicide note.[3]
The cause of death was decided by the authorities in 14 (or 40, unclear) minutes.[4]
The parents arranged a private autopsy which “made their suspicions stronger”.[5]
The parents say “there are a lot of facts that are very disturbing for us and we cannot share at the moment but when we do a PR all of that will come out.”[6]
The parents say “his computer has been deleted, his desktop has been messed up”.[7]
Although the parents also said that their son’s phone and laptop are not lost and are in escrow.[8][9] I think the claim of the computer being deleted is more up-to-date, but I’m not sure as that video was posted earlier.
It was his birthday and he bought a bike on the week of his death.[10]
He said he didn’t want to work and he was going to take a gap year, “leaving the AI industry and getting into machine learning and neuroscience” but also he was planning to start his own company and was reaching out to VCs for seed funding.[11]
He had just interviewed with the New York Times and he was supposed to do further interviews in the days after his death.[12]
According to the parents and Webb, there are signs of foul play at the scene of death:
There are several areas with blood, [Confirmed from pictures] suggesting to Webb and the parents he was trying to crawl out of the bathroom.[13][14]
Webb says the body had bleeding from the genitals.[15] I’m not aware of a better source for this claim, so right now I think it is probably false.
The trash can in the bathroom was knocked over.[13][16] [Confirmed from pictures].
A floss pick is on the floor.[13][17] [Confirmed from pictures]. Webb interprets this as being dropped at the time of death, suggesting that Balaji was caught by surprise.
The path of the bullet through the head missed the brain. I’m not sure what the primary source for this is, but I’m not sure why Webb would invent this, so I think it’s true. Webb takes this as evidence that it was shot during a struggle rather than at the considered pace of a suicide.[18]
The bullet did not go all the way through the head, suggesting a lower caliber, quiet gun.[19]
According to Webb and the parents, the drawers of the apartment were ransacked, the cupboards were thrown open.[8][20] From the pictures this looks false, although the apartment is very messy and his hiking backpacks are strewn around with much of their contents on the table (he had recently returned from a hiking trip).
The blood on the sink looks different, suggesting to Webb that it came from a different part of the body.[21] This is not obvious to me from the pictures but not implausible and the main pool of blood looks surprisingly dark.
There is a half-eaten meal at the desk in the apartment. [Confirmed from pictures].
There is a tuft of Balaji’s hair, soaked in blood, under the bathroom door. [Confirmed that’s what it looks like in the pictures], again suggesting to Webb a violent struggle.
According to the parents, he had a USB thumb drive which is now missing, containing important evidence for an upcoming court case about OpenAI’s use of copyrighted data.[8][22]
People that spoke to him around the time of his death report him to have been in high spirits and making plans for the near future.[23]
George Webb claims there were security cameras working all on floors except the floor which he lived on.[24] This appears to conflict with the parents’ claim that the police said no one came in or out (see below), but may be referring to different cameras, as the parents also mention that the murderer could have come through a different entrance to the main one.[25]
The parents say “OpenAI has deleted the copyright data that was evidence that was given to the discovery for the [New York Times] lawsuit. They deleted the data and now my son is also gone, so now they’re all set for winning the lawsuit… It’s also said that my son had the documentation to prove the copyright violation. His statement, his testimony would have turned the AI industry upside down...”[26]
Looking into the details of this, OpenAI did delete some data but this wasn’t a permanent deletion of any of the primary sources and I think it was probably an accident and not significant to the outcome of the case.
I don’t see any strong reason to believe Balaji had secret evidence that would have been critical to the outcome of the case.
Ilya Sutskever had two armed bodyguards with him at NeurIPS.
Evidence against:
One reason the authorities gave for declaring it a suicide was that CCTV footage showed that no one else came in or out of the apartment.[27]
In high school Balaji won a $100,000 prize for a computer science competition. His parents didn’t find out until they saw the news online, suggesting he may not have been very open with them.[28]
My interpretations:
If we interpret the apartment as simply messy (as it looks to me), rather than ransacked, then we can probably discount the knocked-over trash can, the floss pick on the floor and the half-eaten meal. We can also probably discard the hypothesis of someone trying to locate a USB drive with secret information, which raises more questions than it answers (why didn’t he reveal this information before? why didn’t he back up this crucial data anywhere?).
In my uninformed view, it doesn’t look like the pictures of the scene of death strongly suggest a struggle between murderer and victim, although I don’t know how to explain the tuft of hair.
The motivations of OpenAI or some other actor to murder a whistleblower are unlikely. The most plausible to me is that they want to send a warning to other potential whistleblowers, but this isn’t very compelling.
There’s no smoking gun and the parents (understandably) do not look like they are thinking very systematically to establish a case for foul-play. This is notable because their claim of foul-play is the main factor that privileged this hypothesis to credible people.
Balaji appeared from the outside to be a happy and highly successful person with important plans in the next few days. It is surprising that someone like that would commit suicide.
Overall my conclusion is that this was a suicide with roughly 96% confidence. This is a slight update downwards from 98% when I first heard about it and overall quite concerning.
I encourage people to trade on this related prediction market and report further evidence.
Useful sources:
https://x.com/RealGeorgeWebb1/status/1874166318053941567
https://www.youtube.com/watch?v=zu1whk7XdCo
https://rumble.com/v6411hm-forensics-say-the-open-ai-whistleblower-was-executed.html?e9s=src_v1_upp
https://www.youtube.com/watch?v=4LkteX3o1Co
I’m not linking to this evidence here, in the spirit of respecting the wishes of the parents, but this is an important source that informed my understanding of the situation.
https://x.com/RaoPoornima/status/1876101052065558998
Source: Poornima Ramarao (11:22)
Source: Poornima Ramarao (12:38)
Source: Poornima Ramarao (13:02)
Source: Poornima Ramarao (15:47)
Source: Poornima Ramarao (16:36)
Source: George Webb + Poornima Ramarao (1:45)
Source: George Webb (9:56)
Source: (23:27)
Source: (8:02)
Source: (26:00)
Source: George Webb + Poornima Ramarao (0:35)
Source: George Webb (6:53)
Source: George Webb (3:38)
Source: George Webb (5:44)
Source: George Webb (5:46)
Source: George Webb (0:05)
Source: George Webb (6:23)
Source: George Webb (9:12)
Source: Poornima Ramarao (1:18)
Source: George Webb (9:45)
Source: Poornima Ramarao (4:30)
Source: George Webb (9:30)
Source: Poornima Ramarao (2:40)
Source: Poornima Ramarao (4:14)
Source: Poornima Ramarao (12:42)
Source: Ramamurthy (17:37)
Source: George Webb (13:29)
Source: George Webb (5:43)
As in, this is also what the police say?
Did the police find a gun in the apartment? Was it a gun Suchir had previously purchased himself according to records? Seems like relevant info.
Yes, edited to clarify. The police say there was no evidence of foul play. All parties agree he died in his bathroom of a gunshot wound.
The only source I can find on this is Webb, so take with a grain of salt. But yes, they found a gun in the apartment. According to Webb, the DROS registration information was on top of the gun case[1] in the apartment, so presumably there was a record of him purchasing the gun (Webb conjectures that this was staged). We don’t know what type of gun it was[2] and Webb claims it’s unusual for police not to release this info in a suicide case.
Source: George Webb (10:10)
Source: George Webb (2:15)
Well, it seems quite important whether the DROS registration could possibly have been staged. If e.g. there is footage of Suchir buying a gun 6+ months prior, using his ID, etc. then the assassins would have had to sneak in and grab his own gun from him etc. which seems unlikely.
Is the interview with the NYT going to be published?
Is any of the police behavior actually out of the ordinary?
That would be difficult. To purchase a gun in California you have to provide photo ID[1], proof of address[2] and a thumbprint[3]. Also it looks like the payment must be trackable[4] and gun stores have to maintain video surveillance footage for up to year.[5]
My guess is that the police haven’t actually invested this as a potential homicide, but if they did, there should be very strong evidence that Balaji bought a gun. Potentially a very sophisticated actor could fake this evidence but it seems challenging (I can’t find any historical examples of this happening). It would probably be easier to corrupt the investigation. Or the perpetrators might just hope that there would be no investigation.
There is a 10-day waiting period to purchase guns in California[5], so Balaji would probably have started planning his suicide before his hiking trip (I doubt someone like him would own a gun for recreational purposes?).
I think it’s this piece that was published before his death.
Epistemic status: highly uncertain: my impressions from searching with LLMs for a few minutes.
It’s fairly common for victim’s families to contest official suicide rulings. In cases with lots of public attention police generally try to justify their conclusions. So we might expect the police to publicly state if there is footage of Balaji purchasing the gun shortly before his death. It could be that this will still happen with more time or public pressure.
https://www.fastbound.com/ffl-bound-book-software-features/dros/
https://giffords.org/lawcenter/state-laws/background-check-procedures-in-california/
https://www.gunpolicy.org/firearms/citation/quotes/7066
https://apnews.com/article/gun-stores-firearms-mass-shootings-credit-cards-abe3a28ea7117340d9a4a8bcde3693fe
https://giffords.org/lawcenter/state-laws/gun-dealers-in-california/
Some people are asking for a source on this. I’m pretty sure I’ve heard it from multiple people who were there in person but I can’t find a written source. Can anyone confirm or deny?
I don’t understand how Ilya hiring personal security counts as evidence, especially at large events like a conference. Famous people often attract unwelcome attention, and having professional protection close by can help deescalate or deter random acts of violence, it is a worthwhile investment in safety if you can afford it. I see it as a very normal thing to do. Ilya would be vulnerable to potential assassination attempts even during his tenure at OpenAI.
Thank you, this is very interesting and it seems like you did a valuable public service in compiling it
What do you think of the motive that he was counterfactually going to testify in a very damaging way, or that he had very damaging evidecne/data that was deleted?
UK Prime Minister Andy Burnham is aiming for a global agreement on AI, and Sir Ed Davey, leader of the Liberal Democrats (third-biggest party by MPs), has called for an AI pause.
This comes after PauseAI UK’s protest last week in Downing Street, in which we demanded that the government lead international coordination on AI. The protest was covered in national news outlets in Britain, including Channel 4 News, ITV News, The Guardian and others.
It was a very strange experience today seeing the two policy announcements from Ed Davey’s speech today be “Increase the threshold to pay income tax to £15k” and “Ban superintelligent AI”.
There’s an orange elephant in the room that won’t change his mind at the summit, sadly.
But I think the shifting political winds in the UK are nonetheless great news for an unrelated reason: a state visit from the King, with RSI on the agenda, is a moonshot that no other country is equipped to do. I’m probably over-optimistic to think this’ll actually happen, but considering Trump’s admiration for the royals there’s a small chance it’ll actually work.
It feels to me like the current situation is driven in part by the idea that USA based AI dominance means that middle powers become less and less relevant. This pretty much acts as an extension with how tech companies have grown in recent years.
It’s not clear to me, as a UK resident, why the US government should give a toss what we think. I don’t think that the US ”needs“ the UK.
(Of course, the US probably doesn’t end up “winning” AI because AI “wins” AI. But, I don’t think that the US government understands that.)
I think if Burnham genuinely believed there was a >10% probability of everyone dying within a decade, he would probably be saying different things? (The quoted statements do not sound to me like things said by top politicians in surviving worlds.)
Lots of people are leaving UK AISI because the pay is too low. Often they go to work on evals at other AI safety orgs where they will get paid more money. This is terrible. It’s far more valuable to have at least one government agency in the world with top tier AI expertise than it is to have a slightly stronger non-profit eval ecosystem.
A competent department creates unique options for the UK government to pass a pre-deployment testing law better than other countries. Or the government could empower them as a a thoughtful AI regulator that actually understands the risks. And AISI pushes other countries to improve by setting a high bar. It could even be transformed at some point into an international body that is the backbone of an AI treaty.
The jobs are in fact pretty well paid by UK tech standards but people there often have more lucrative options available at AI safety non-profits. Is there any way we can pay them to stay?
It’s sad we can’t pay them because the UKAISI-Redwood pay gap (order of $300k) is tiny compared to the cost of donations that it takes to match a safety researcher’s impact (order of $30M), or even the Redwood-Anthropic pay gap (order of $3M). I feel like at the very least, former lab employees and others without financial concerns should preferentially work for UKAISI and CAISI, as long as they don’t have unresolvable COIs from their equity.
Can you explain the 30M figure here?
I think the value of a year of the AI safety nonprofit ecosystem is ~50B times the value of a marginal dollar, so if one safety researcher is worth ~1/5K of the ecosystem (depends on the person!), that’s $50B/5K = $10M/year.
Surely the comparison of a marginal dollar vs. [total-impact]/[nr of people] is going to be misleading here? If we compared with [total-impact]/[nr of dollars] that’d be more one-to-one but also not what we care about. We’d want to compare the marginal dollar with a marginal person.
Yes but be careful about “marginal person.”
We can talk about Lukas’s value on the margin, which is probably slightly less than total impact * Lukas’s share of the credit due to diminishing returns in the size of the field but that’s not obvious
We can talk about causing new people, who tend to add much less value than average, to join the field
The correct concept here is the former, I think, since we’re talking about trading off between money and work by great people.
I certainly agree with the conceptual point, but it seems like I believe in way faster DMR in the size of the field than you seem to believe in. Like, I think a 50⁄50 gamble to eliminate the field or triple the size seems bad to me, whereas a 50⁄50 gamble to either double or halve it feels kind of unclear if it’s worth it or not. So feels closer to log than linear to me, although I think it’s probably somewhat less steeply diminishing than log.
(I would also have thought that the value of a year of the AI safety nonprofit ecosystem is higher than 50B though, so not at all clear I’m disagreeing with your bottom-line. And probably there’s some correlation here, where money is more valuable if there’s less steeply DMR to the size of the field.)
Sure, there is some pay pressure, but I didnt leave UK AISI because the pay was too low, I left because the projects were uninteresting and behind the curve with respect to what others are doing. Im perfectly happy to work for less pay so long as I am working on meaningful projects at a frontier institution where actual work is happening.
A lot of my work was getting told to investigate this or that, and then I would do the usual initial field overview and get told that that was enough and great work and yada yada yada. Excuse me? We didn’t even calculate anything yet? No no, thats enough...
I’m sorry, I (and many others coming from physics) got hired because of quantitative prowess, the work doesn’t have merit until we actually calculate. Drawing a few functors to map some categories is just a plan and it isn’t enough, this stuff depends on the details and we need to calculate. You can’t “vibemath” alignment.
What are you doing now?
bribessupplemental payments to civil servants are viewed uncharitably for some reason. perhaps a post-service endowment that requires some number of years with AISI as qualification?This is misleading. The L5 comp at UK AISI is about the overall median comp for software engineers in London, but 2x lower than the L5 comp at Meta London and many many multiples lower than HFT /frontier lab comp in London. Presumably the latter is the relevant reference class for the talent that AISI wants to hire.
AISI got an especial carve-out for those technical roles, which is why they are somewhat competitive, but the pure policy roles are not as high. I can’t find any current AISI policy job openings but in the past a Transformative AI Policy Analyst had a salary range of £44,195–£48,620, and a Strategy Delivery Manager, a range of £44-48k
I do wonder how those policy roles compare to roles with similar jobs titles elsewhere in the AI industry, though. In the UK civil service there are a lot of junior policy roles where you are expected to basically have a general understanding of the policy and act as a go-between between the technical staff and ministers.
E.g. this Cyber Resilience Policy Advisor role is paying £39-43k in London (£36-39k elsewhere) - ‘Awareness of, and in interest in, the UK’s cyber security landscape and basic cyber security concepts’ is listed as a desirable characteristic but not essential.
UK salaries in every field are pretty terrible. I suspect that this is a result of pre-existing wealth, basically capital is more important, in order to break in to that tier you either need a 98+ percentile salary or some other mechanism e.g start a high risk business, invest in high risk things, etc.
The public sector is uniquely bad for this because the comparison points (in a political sense) end up being with things like the distribution of salaries rather than the end result of what that money gets you.
For example, a house suitable to raise a small family in London, an MP on an MP’s salary could not get a mortgage on that, they would need some other source of funds (inheritance, previous job, etc).
This is considered to be just how it works, workers are conditioned to accept these low (in comparison with capital/wealth) salaries.
As a result of that I think that getting the required sums from taxpayer money is effectively impossible.
There are powerful non-monetary incentives to be in Parliament. Since UK law requires all government ministers to be MPs, most of the most powerful 500 people in the UK are MPs. Many people find power to be very intoxicating even if one of the conditions for holding that power is that it is illegal to convert that power into personal wealth. But the same analysis does not apply to the other 99.999% of jobs in the UK, so it would be nice if you could come up with an example population of workers other than MPs. Almost any other population would do. (Rock musician would not do because there are powerful non-monetary incentives to being a male rock musician because many young women find rock musicians extremely attractive.)
Are the rates of having high personal wealth lower in the US than in the UK? I doubt it is much lower. So if (as you claim) it is the presence of high-net-worth individuals in the labor market that keeps salaries low in the UK, then why is the mechanism you claim is operating in the UK so much weaker in the US? That is, why are US salaries, particularly for jobs in computing and AI, much higher than UK salaries?
Do you claim that the same mechanism operates in other markets or is it restricted to the labor market? For example, do high-net-worth individuals on average pay more for the same car as a no-net-worth individual because they can afford to pay more? Do high-net-worth individuals sometimes pretend to be no-net-worth individuals when talking to car salesmen to get a better price? It seems to me that they almost never bother doing that because they don’t have to because the salesman (and the car dealership he works for) regards one pound as just as desirable as any other pound regardless of the wealth status of the person giving them the pound. There a people who hold an animus towards wealthy people, but even such a person if employed as car salesman will almost always care much more about making a sale (and consequently earning a commission on that sale) than on disadvantaging the wealthy person or making the wealthy person suffer a little. Increasing personal income is just a much more compelling consideration than most considerations for the average person.
Would the cost of a particular make and model of car in country X or city X be lower if no one in country X or city X had any savings? My guess is no. It costs Y pounds to make the car, so that Y pounds acts as a strict lower bound on the average selling price of the car. A seller would like the average selling price to be as high as possible, but buyers are price sensitive, so if the seller’s price is higher than the price of another seller of the same make and model of car, buyers will just buy from the other seller till the other seller gets low on inventory and the prices equalize. These are the main 2 mechanisms determining the prices of cars, and I don’t see how the net worth of the buyers affects these mechanisms in any significant way.
I always thought the labor market works basically the same way: if employer A is willing to pay salary X to job Y and additionally if employer B stubbornly refuses to pay more than 0.9X, then pretty soon employer B will be unable to hire any employees willing and able to do Y, so the salaries offered by A and B will tend to equalize (unless for example B’s workplace is much easier to commute to than A’s or some other factor, but all of these factors tend to be minor factors relative to salary).
I expect wealthy people to care about salary, too, such that as long as employer B stubbornly refuses to pay more than 0.9X, it will find it difficult to hire even the very wealthiest people. People don’t stop wanting more money just because they already have a lot of money. If you ask people in a survey how much income would they need to be happy, they answer about 30% more than their current income regardless of whether their current income is 10,000 pounds a year, 20,000, 40,000, 80,000 or 160,000.
If the salaries A and B are willing to pay get high enough, then people who previously never considered being a Y will start considering it and start acquiring the necessary qualifications. Also, if the economy changes such that the products of A and B no longer command as high a price as they once did, then the amount A and B are able to pay for labor used in the production of those products must go down.
Again I always thought that this is the basic dynamics by which salaries are set, and again I fail to see how the average net worth of the employees and job applicants affects the dynamics significantly. In fact, the largest effect (although not particularly large) I imagine is its raising salaries because there are unavoidable aversive aspect to holding down most high-paying jobs, and people of wealth will tend to choose to avoid those aversive aspects by taking themselves out of the market for those jobs, requiring employers to raise salaries (at least a little) to attract sufficient numbers of applicants.
It’s a long winded topic but my general feelings are some mixture of:
the US celebrates capitalism in a way that the UK doesn’t and so the market and workers are actually just more competitive in general. ask the guy in the rural petrol (gas) station about business in the US and he will probably perk up. ask a reasonably middle class person in the UK and they will kind of flinch as if we just don’t do that kind of ambition here.
older country with more established class system
smaller country with worse building restrictions
more centralised around London
Bear in mind that I’m not talking about absolute wealth but in comparison to salaries e.g. does work buy you the ability to work less at some point or is it tightly bound around minimum cost of living
edit: re the car thing; more like high NW families have actual family houses and low NW families rent flats, cars are irrelevant expenses unless you look at clearly silly things like a ferrari
Anthropic is reportedly lobbying against the federal bill that would ban states from regulating AI. Nice!
When people compliment you on highly salient attributes, you should be skeptical.
If you have a very prominent tattoo or a funky rainbow-colored jacket, people will frequently give you compliments about it, whether or not it is actually good. When there is a highly conspicuous thing, people cannot resist commenting on it. But the only polite way to mention something about someone’s appearance is to give a compliment. So people who make terrible fashion choices are frequently reinforced in their belief that everyone around them thinks they look great.
Similarly, if you give a public speech in front of a large audience or perform a song at a concert, everyone will compliment your performance, even if it was bad, because they can’t bring themselves not to mention the most salient fact about you at that moment, and a compliment is the only socially acceptable way to do that.
I would counter the idea that people are simply being polite when expressing compliments with the argument that what you are likely seeing instead is simply a kind of selection bias.
9x% of the population might think that you look ridiculous, but the single digit percentage that you appeal to could be both honest and vocal. In my experience, tattoos _in general_ are a great example of this where the “masses” will just not comment at all but tattoo fans are loud.
Or there’s the opposite—perhaps in a general “poll the world” sense people would think that you look pretty good, but when you wear X in a Y area, the Y’s don’t tend to like it.
Yeah, it could be that 99% people hate it but remain silent, and 1% really like it and say it out loud. You think you are popular in general, while in reality you are just optimizing for the tastes of 1%, and getting more and more silent disapproval of the 99%. Even worse, the reason why the 1% like it may be different than you think; for example you think “this is nice”, but for them it is an arbitrary ingroup symbol, so you are ultimately wrong about everyone.
People could also not comment about it to the person in question, but satisfy the “cannot resist commenting on it” energy by waiting to be out of earshot and then telling someone else how terrible that person looked. I feel like this would happen a lot more often than giving a fake compliment just out of the need to say something.
I’ve also worn conspicuous clothing before and gotten comments that were more about expressing surprise/amusement rather than being compliments. (I don’t remember the exact words, but sentiments along the lines of “this is the first time I’ve seen someone wear cat ears to work” or whatever.)
I think people do this, but there’s a risk to assuming this motivation. It’s taking a statement that’s positive on its face, and then coming up with a scenario via mind reading that it’s actually bad for you. Sort of reverse CBT that could make you more anxious or paranoid over time since you are finding a way to reframe even positive comments as bad. Then you are potentially disappointed whether people compliment or insult you.
I don’t think this is a knock-down argument against discussing CBRN risks from AI, but it seems worth considering.
The trick is that chem/bio weapons can’t, actually, “be produced simply with easily available materials”, if we talk about military-grade stuff, not “kill several civilians to create scary picture in TV”.
You sound really confident, can you elaborate on your direct lab experience with these weapons, as well as clearly define ‘military grade’ vs whatever the other thing was?
How does ‘chem/bio’ compare to high explosives in terms of difficulty and effect?
Well, I have bioengineering degree, but my point is that “direct lab experience” doesn’t matter, because WMDs in quality and amount necessary to kill large numbers of enemy manpower are not produced in labs. They are produced in large industrial facilities and setting up large industrial facility for basically anything is on “hard” level of difficulty. There is a difference between large-scale textile industry and large-scale semiconductor industry, but if you are not government or rich corporation, all of them lie in “hard” zone.
Let’s take, for example, Saddam chemical weapons program. First, industrial yields: everything is counted in tons. Second: for actual success, Saddam needed a lot of existing expertise and machinery from West Germany.
Let’s look at Soviet bioweapons program. First, again, tons of yield (someone may ask yourself, if it’s easier to kill using bioweapons than conventional weaponry, why somebody needs to produce tons of them?). Second, USSR built the entire civilian biotech industry around it (many Biopreparat facilities are active today as civilian objects!) to create necessary expertise.
The difference with high explosives is that high explosives are not banned by international law, so there is a lot of existing production, therefore you can just buy them on black market or receive from countries which don’t consider you terrorist. If you really need to produce explosives locally, again, precursors, machinery and necessary expertise are legal and widespread sufficiently that they can be bought.
There is a list of technical challenges in bioweaponry where you are going to predictably fuck up if you have biological degree and you think you know what you are doing but in reality you do not, but I don’t write out lists of technical challenges on the way to dangerous capabilities, because such list can inspire someone. You can get an impression about easier and lower-stakes challenges from here.
This seems incredibly reasonable, and in light of this, I’m not really sure why anyone should embrace ideas like making LLMs worse at biochemistry in the name of things like WMDP: https://www.lesswrong.com/posts/WspwSnB8HpkToxRPB/paper-ai-sandbagging-language-models-can-strategically-1
Biochem is hard enough that we need LLMs at full capacity pushing the field forward. Is it harmful to intentionally create models that are deliberately bad at this cutting edge and necessary science in order to maybe make it slightly more difficult for someone to reproduce cold war era weapons that were considered both expensive and useless at the time?
Do you think that crippling ‘wmd relevance’ of LLMs is doing harm, neutral, or good?
My honest opinion is that WMD evaluations of LLMs are not meaningfully related to X-risk in the sense of “kill literally everyone.” I guess current or next-generation models may be able to assist a terrorist in a basement in brewing some amount of anthrax, spraying it in a public place, and killing tens to hundreds of people. To actually be capable to kill everyone from a basement, you would need to bypass all the reasons industrial production is necessary at the current level of technology. A system capable to bypass the need for industrial production in a basement is called “superintelligence,” and if you have a superintelligent model on the loose, you have far bigger problems than schizos in basements brewing bioweapons.
I think “creeping WMD relevance”, outside of cyberweapons, is mostly bad, because it is concentrated on mostly fake problem, which is very bad for public epistemics, even if we forget about lost benefits from competent models.
Are you open to writing more about this? This is among top 3 most popular arguments against open source AI on lesswrong and elsewhere.
I agree with you you need a group of > 1000 people to manufacture one of those large machines that does phosphoramidite DNA synthesis. The attack vector I more commonly see being suggested is that a powerful actor can bribe people in the existing labs to manufacture a bioweapon while ensuring most of them and most of rest of society remains unaware this is happening.
I’m trying to write post, but, well, it’s hard.
Many people wrongly assume that the main way to use bioweapons is to create small amount of pathogen to release it in environment with outbreak as an intended outcome. (I assume that where your sentence about DNA synthesis comes from.) The problem is that creating outbreaks in practice is very hard, we, thankfully, don’t know reliable way to do that. In practice, the way that bioweapons work reliably is “bomb-saturate the entire area with anthrax such that first wave of death is going to be from anaphylactic shock rather than infection” and to create necessary amount of pathogen you need industrial infrastructure which doesn’t exist, because nobody in our civilization cultivates anthrax at industrial scale.
I wrote about something similar previously: https://www.lesswrong.com/posts/Ek7M3xGAoXDdQkPZQ/terrorism-tylenol-and-dangerous-information#a58t3m6bsxDZTL8DG
I agree that 1-2 logs isn’t really in the category of xrisk. The longer the lead time on the evil plan (mixing chemicals, growing things, etc), the more time security forces have to identify and neutralize the threat. So all things being equal, it’s probably better that a would be terrorist spends a year planning a weird chemical thing that hurts 10s of people, vs someone just waking up one morning and deciding to run over 10s of people with a truck.
There’s a better chance of catching the first guy, and his plan is way more expensive in terms of time, money, access to capital like LLM time, etc. Sure someone could argue about pandemic potential, but lab origin is suspected for at least one influenza outbreak and a lot of people believe it about covid-19. Those weren’t terrorists.
I guess theoretically, there may be cyberweapons that qualify as wmd, but those will be because of the systems they interact with. It’s not the cyberweapon itself, it’s the nuclear reactor accepting commands that lead to core damage.
I’d love a reply on this. Common attack vectors I read on this forum include 1. powerful elite bribes existing labs in US to manufacture bioweapons 2. nation state sets up independent biotech supply chain and starts manufacturing bioweapons.
https://www.lesswrong.com/posts/DDtEnmGhNdJYpEfaG/joseph-miller-s-shortform?commentId=wHoFX7nyffjuuxbzT
This has been an option for decades, a fully capable LLM does not meaningfully lower the threshold for this. It’s already too easy.
This has been an option since the 1950s. Any national medical system is capable of doing this, Project Coast could be reproduced by nearly any nation state.
I’m not saying it isn’t a problem, I’m just saying that the LLMs don’t make it worse.
I have yet to find a commercial LLM that I can’t make tell me how to build a working improvised explosive (I can grade the LLMs performance because I’ve worked with the USG on the issue and don’t need a LLM to make evil).
Makes sense, thanks for replying.
Do you have a link/citation for this quote? I couldn’t immediately find it.
I first encountered it in chapter 18 of The Looming Tower by Lawrence Wright.
But here’s a easily linkable online source: https://ctc.westpoint.edu/revisiting-al-qaidas-anthrax-program/
On LessWrong, the frontpage algorithm down-weights older posts based on the time-since-posted, not the time-since-frontpaged. So, if a post doesn’t get frontpaged until a few days after posting, then it’s unlikely to get many views.
LessWrong has an autofrontpager that works a reasonable amount of the time. Otherwise, posts have to be manually frontpaged by a person. In my experience, this was always quite quick, but my most recent post was not frontpaged until 3 days after it was posted, so AFAICT it never actually appeared on the frontpage (unless you clicked “Load More”).
I think the solution is to downweight posts based on the time-since-frontpaged.
If you downweigh posts based on the time-since-frontpaged then posts get a huge boost when they have a delay of getting frontpaged (since they then first show to everyone who has personal blog enabled on their frontpage, and can accumulate karma during this time, and then when they have their effective date reset have a huge advantage over posts that were immediately frontpaged, because the karma provides a much longer visibility window).
I don’t really have a great solution to this problem. I think the auto-frontpager helps a lot, though of course only if we can get the error rate sufficiently down.
I’d be happy, if the auto-frontpager is ~instant, to get the option “delay publishing until human review” if it declines frontpage. Whether something gets ~50% less karma than it would by default is a pretty major drop in the effectiveness of what is often many hours of work, I’d be fine with waiting a day or two to avoid that usually.
That makes sense! I’ll think about it, though probably fitting that complexity into the publishing process isn’t worth it.
Relatedly, I’ve been thinking about building a schedule-this-post-for-publication feature. If I publish a post at 10pm, it’s often better to publish the next morning for visibility. My guess is this would be useful for Inkhaven Residents who finish writing near-midnight.
If I could schedule, the frontpage review happened before publishing, and the schedule UI had “delay publishing until frontpage”[1] as a checkbox, this would be ~solved.
I’d prefer this to “delay publishing until human review”, as ~half a dozen times in the past few years I’ve appealed via Intercom and had a human-reviewed page retroactively frontpaged (usually a resource, which LW team’s priors seem to be something like ‘this won’t be maintained’ but will because I optimize a bunch for not leaving stale projects).
Examples which Rafe requested when I mentioned this: the following were all marked as personal blog until I intercom’d in and asked for a re-assessment
https://www.lesswrong.com/posts/JsqPftLgvHLL4Pscg/new-weekly-newsletter-for-ai-safety-events-and-training
https://www.lesswrong.com/posts/dEnKkYmFhXaukizWW/aisafety-community-a-living-document-of-ai-safety
https://www.lesswrong.com/posts/vxSGDLGRtfcf6FWBg/top-ai-safety-newsletters-books-podcasts-etc-new-aisafety (nudge didn’t work for this one)
https://www.lesswrong.com/posts/MKvtmNGCtwNqc44qm/announcing-aisafety-training
https://www.lesswrong.com/posts/JRtARkng9JJt77G2o/ai-safety-memes-wiki
https://www.lesswrong.com/posts/x85YnN8kzmpdjmGWg/14-ai-safety-advisors-you-can-speak-to-new-aisafety-com
Perhaps it could use time-since-frontpaged only if the karma is below some threshold.
Or even better, at the the time when a post is frontpaged, check if will actually appear on the frontpage. If it is too old and has too little karma to be seen, then use the time-since-frontpaged.
“Back in December 2025, working with the campaign group PauseAI, I was the first MP to launch an AI Safety Debate in the House of Commons. Some colleagues thought the subject premature, one for technologists and scientists to determine, not legislators. Nine months on, that argument is dead.”
Iqbal Mohamed MP today in The National: https://x.com/PauseAI_UK/status/2104013872164180365
Startups often pivot away from their initial idea when they realize that it won’t make money.
AI safety startups need to not only come up with an idea that makes money AND helps AI safety but also ensure that the safety remains through all future pivots.
[Crossposted from twitter]
If you combine the fact that power corrupts your world models with the general startup person being power hungry as well as AI Safety being a hot topic, you also get a bunch of well meaning people doing things that are going to be net-negative in the future. I’m personally not sure that the VC model actually even makes sense for AI Safety Startups given some of the things I’ve seen in the space.
Speaking from personal experience I found that it’s easy to skimp out on operational infrastructure like a value aligned board or a more proper incentive scheme. You have no time so instead you start prototyping a product yet that means you get this path dependence where if you succeed, you suddenly have a lot less time. As a consequence the culture changes because the incentives are now different. You start hiring people and things become more capability focused. And voila, you’re now in a capabilities/AI safety startup and it’s unclear what it is.
So get a good board and don’t commit to something unless you have it in contract form or similar that you will have at least a PBC structure if not something even more extreme as the underlying company model. The main problem I’ve seen here is if your co-founder(s) is/are being cagey about it, I would move on to new people at least if you care about safety.
I think what you’re saying is that they need to be aligned.
Best way to start an AI safety startup is get enough high status credentials and track record that you can ask your investors to go fuck themselves if they ever ask you to make revenue. Only half-joking. Most AI research (not product) companies have no revenue today, or are trading at an insane P/S multiple.
Silicon Valley episode: No revenue
For a while, everyone was underinvesting in AI because they were not sufficiently AGI-pilled. Now, I think a lot of people (the big companies, investors, people in this community) might be overinvesting in AI because they’re not sufficiently AGI-pilled.
AI data center capex is approaching $1 trillion per year. Enough gigawatts are already planned to multiply the AI labs’ compute several times over in the next few years.
And yet the models are going crazy. I feel like the new revelations about misaligned agent swarms from OpenAI are not yet priced in. How can we possibly keep accelerating for another three years like this? If nothing else cybersecurity is surely going to break.
I can see a few possible futures:
Governments step in and regulate pre-emptively before a major catastrophe.
There is a major catastrophe and everyone is forced to slow down.
We fully lose control and AI takes over.
But another three years of just more scaling and acceleration like we’ve had the last three is increasingly infeasible. We are entering the danger zone and the race cannot just carry on as before.
I wouldn’t expect markets to price in the third possible future (total loss of control), but they should price in the first two. I think we could soon see a market correction when investors see that a deceleration will be required to keep models under control.
Economically it is probably rational for the world to be investing $1 trillion in capabilities and $10 trillion in safety. Unless you’re pessimistic, in which case we should spend $0 trillion on capabilities and still $10 trillion on safety.
This assumes that safety is the kind of good you can invest in. Right now that feels about as true as investing in peace or investing in love or similarly abstract concepts. Approximately nobody who is mainly money-bottlenecked has demonstrated a compelling track record of turning money into safety.
My main issue is that the world will find it hard to invest anything like ten trillion dollars. This is around a half of Chinese annual GDP, a quarter of the USA’s debt or a third of the USA’s annual GDP. Additionally, the world’s understanding of the AIs’ wildly transformative potential is surprisingly poor and prevents the world from making the relevant investments, especially during the looming financial crisis.
$0 trillion on capabilities and $1 trillion on safety would be equally fine.
The idea is the faster safety is solved, the faster we can scale capabilities safely, which increases the growth rate of the economy from ~3% to ~100% and makes people immortal. If people want to maximize something like their discounted log(consumption) over the next 100 years, a wartime investment is warranted unless you think we couldn’t solve AI safety well enough to cure aging and automate the economy in a few decades.
I’m definitely on the pessimistic side, but yes I can see that spending more on safety may shorten the timeline for solving death which is probably a good thing if it doesn’t also increase s-risks.
My main worry is the following couple of doom scenarios.
Scenario A:
The amount of investments which was already priced in by Some Lab becomes reduced by external shocks, the lack of priced-in capabilities or misalignment;
Some Lab decides to avoid becoming bankrupt by gaining the revenue, to deliver the capabilities and uses potentially hazardous architectures like neuralese;
The capabilities emerge, but the AIs aren’t aligned.
Scenario B:
The $2,4T of investments which were already priced in by AI-2040 become reduced by external shocks, misalignment or the lack of priced-in capabilities;
Capabilities do slow down far more than AI-2040 predicted, but China keeps racing to produce compute;
The leadership is lost to China where safety is far worse, but the USA and China don’t do things like cross-audits of training runs.
Superintelligence safety isn’t “far worse” in China. You can’t get worse than zero superintelligence safety, which is what we currently have in both the US and China.
The next PauseAI UK protest will be (AFAIK) the first coalition protest between different AI activist groups, the main other group being Pull the Plug, a new organisation focused primarily on current AI harms. It will almost certainly be the largest protest focused exclusively on AI to date.
In my experience, the vast majority of people in AI safety are in favor of big-tent coalition protests on AI in theory. But when faced with the reality of working with other groups who don’t emphasize existential risk, they have misgivings. So I’m curious what people here will think of this.
Personally I’m excited about the protest and I’ve found the organizers of Pull the Plug to be very sincere and good to work with, but I’ve also set things up so that the brands of PauseAI UK and Pull the Plug are clearly distinct, so that our messaging remains clearly focused on the risks of future AI. For example, we have a separate signup page and we have our own demands focused on decelerating frontier development.
is this true? I think many people (myself included) are worried about conflationary alliances backfiring (as we see to some extent in the current admin)
I only have anecdata but I’ve talked to quite a few people and most people say it’s is a good idea to use the myriad of other concerns about AI as a force multiplier on shared policy goals.
Speaking only for myself, here: There’s room for many different approaches, and I generally want people to shoot the shots that they see on their own inside view, even if I think they’re wrong. But I wouldn’t generally endorse this strategy, at least without regard for the details of how the coalition is structured and what it’s doing.
I think our main problem is a communication problem of getting people to understand the situation with AI
that model capabilities are steadily increasing;
that the labs are aiming at literal superintelligence, no really, something more capable than any human alive, and then even better than that; that the labs are explicitly aiming to do an RSI, which looks increasingly likely to succeed;
that there is not a known science of reliably controlling or shaping the motivations of superhuman AIs.
that there are competitive pressures for all of the labs and all of the countries to beat their competitors, so slowing down or pausing requires international coordination.
These are slippery points to get across specifically because audiences tend to slip into visualizing something other than “actual strategic superintelligence”, that is automating science and technological progress and capable of strategically outmaneuvering adversaries—even when I talk with people from the labs, they often tend to gravitate to a fuzzier vision that has the form factor of the current AI chatbots / agents, but is much more competent.
Most of the time, I’m trying to land these points, despite the slipperiness, and talking about present-day harms that don’t have a through-line to the core alignment problems seem like more of a distraction than a help.
If we already had developed policies that would substantially improve the situation and were politically feasible, and we just needed to get a big enough coalition to get them implemented, I would feel differently.
But insofar as we have policies substantially help, they’re rather radical (on the order of “don’t allow private individuals to own more than 8 GPUs” and “negotiate with China for an international pause in frontier AI development”), and are only politically realistic if the stakeholders have a close-to-accurate picture of the situation.
From https://pulltheplug.uk/:
I think this would probably be a disaster, given how misinformed and unwise large parts of the broad public have been on many other scientific issues (e.g. vaccines, GMOs, nuclear power).
The rest of their views doesn’t inspire much confidence in their epistemics either:
Citizen assemblies often involve selecting a small number of delegates who are then informed about the all of the details of the issue in depth, including by expert testimonies, which the delegates have the affordance to do because they’re being paid for their time.
My understanding is that this works pretty well for coming to reasonable policy.
Who’s behind Pull the Plug? I don’t see any details about it on their website.
LLM hallucination is good epistemic training. When I code, I’m constantly asking Claude how things work and what things are possible. It often gets things wrong, but it’s still helpful. You just have to use it to help you build up a gears level model of the system you are working with. Then, when it confabulates some explanation you can say “wait, what?? that makes no sense” and it will say “You’re right to question these points—I wasn’t fully accurate” and give you better information.
What if you say that when it was fully accurate?
Then it will often confabulate a reason why the correct thing it said was actually wrong. So you can never really trust it, you have to think about what makes sense and test your model against reality.
But to some extent that’s true for any source of information. LLMs are correct about a lot of things and you can usually guess which things they’re likely to get wrong.
Not OP but IME it might (1) insist that it’s right, (2) apologize, think again, generate code again, but it’s mostly the same thing (in which case it might claim it fixed something or it might not), (3) apologize, think again, generate code again, and it’s not mostly the same thing.
Announcing PauseCon, the PauseAI conference.
Three days of workshops, panels, and discussions, culminating in our biggest protest to date.
Tweet: https://x.com/PauseAI/status/1915773746725474581
Apply now: https://pausecon.org
The next international PauseAI protest is taking place in one week in London, New York, Stockholm (Sunday 9th Feb), Paris (Mon 10 Feb) and many other cities around the world.
We are calling for AI Safety to be the focus of the upcoming Paris AI Action Summit. If you’re on the fence, take a look at Why I’m doing PauseAI.
140 people came to the UK Parliament yesterday for the PauseAI conference.
The panel speakers were:
Dame Chi Onwurah MP, Chair of the Science, Innovation & Technology Committee
Professor Stuart Russell, world-renowed AI scientist
Lord Tim Clement-Jones, Co-Founder and Co-Chair of the All-Party Parliamentary Group on AI
Lord Lionel Tarassenko—President of Reuben College, Oxford
Iqbal Mohamed MP, leader of the Dec 2025 Westminster Hall debate on AI safety
Brando Benifei MEP, lead architect of the EU AI Act
I think this is probably the largest event on AI safety there’s been in Parliament. The tone from all the panelists was very supportive and most expressed very serious concern about AI, including the risk of extinction.
Brando Benifei gave us a dose of harsh reality with his assessment about the difficulty of achieving a global treaty of AI. But he remains supportive of PauseAI and seems to really by trying his best to get the EU to lead international cooperation for a ‘CERN for AI’ type solution.
[Linkpost for https://x.com/PauseAI_UK/status/2097102594908803088]
Niradata ran the largest ever survey on global attitudes to superintelligence (n=377,458). The question was:
The headline result is this:
PauseAI UK has published a visualization of the full results on our website.
I think choosing to make the cut at “net support of rapid development” when people with even the slightest sense of “hey maybe let’s put some rules on the thing” get sorted into the “strict oversight” bucket is a weasely way of presenting these results. You could just as easily make the case that both “Pause/Stop” and “Continue” are matched at 38% of the vote.
Fair point. I just copied that from the original report, but I’ve removed it from our presentation of results.
In response to Annoyingly Principled People, and what befalls them.
I think some people should internalize that often Principles Don’t Justify Drama (Among Humans). Drama destroys future discussion and activates primal emotions that will override rational deliberation. If you want to bring people around to your viewpoint, make them like you.
If everyone died on their hill every time someone violated what they thought was an important principle, society would not function. Cooperate with people you disagree with and when you violate someone else’s principle, be thankful that they are going to cooperate with you.
The best historical example of this is probably religious toleration. Thank God for those who tolerated infidels.
Of course the art is in knowing which hills to actually die on and who you should really shun as your enemy. To get it right, you need to fully appreciate the costs.
This advice works equally well whether or not your viewpoint is true.
If you make people dislike you then you will push them away from your viewpoint whether or not your viewpoint is true.
If you make people like you then there’s a greater chance they will chance they will judge your view on the merits.
The people you are talking with, maybe. You also might have an audience. Which archetypal character do people like more when they’re viewing an interaction from a distance, the abrasive prophet or the unctuous people-pleaser?
You’ll get better signal-to-noise in convincing people of true things (compared to false things) the more of an argument you’re able to get them to listen to (up to a point where you overwhelm their argument processing capacity) which is easier to do if they like you enough to listen to you.
This isn’t obvious. What if people who like you tend to agree with your conclusions “as a favor” even if your arguments are bad?
This all leads into Zack Davis’s top rationality advice: become unlikable, so that people will never agree with you just to get along. They’ll argue with everything you say and you increase your chance of hearing true counterarguments!
(jk)
I definitely wouldn’t call it “top rationality advice” because there are lots of other reasons to want to be liked and not want to be disliked, but I do expect the effect you describe to be real.
I think what you describe is mostly a high-status (dominance or prestige) phenomenon rather than a high-likeability phenomenon. But I will admit I’m rather going off instinct and intuition here.
When I go on LessWrong, I generally just look at the quick takes and then close the tab. Quick takes cause me to spend more time on LessWrong but spend less time reading actual posts.
On the other hand, sometimes quick takes are very high quality and I read them and get value from them when I may not have read the same content as a full post.
Interesting. I am concerned about this effect, but I do really like a lot of quick takes. I wonder whether maybe this suggests a problem with how we present posts.
Quick takes are presented inline, posts are not. Perhaps posts could be presented as title + <80 (140?) character summary.
I think the biggest problem with how posts are presented is it doesn’t make the author embarrassed to make their post needlessly long, and doesn’t signal “we want you to make this shorter”. Shortforms do this, so you get very info dense posts, but actual posts kinda signal the opposite. If its so short, why not just make it a shortform, and if it shouldn’t be a shortform, surely you can add more to it. After all, nobody makes half-page lesswrong posts anymore.
This. The struggle is real. My brain has started treating publishing a LessWrong post almost the way it’d treat publishing a paper. An acquaintance got upset at me once because they thought I hadn’t provided sufficient discussion of their related Lesswrong post in mine. Shortforms are the place I still feel safe just writing things.
It makes sense to me that this happened. AI Safety doesn’t have a journal, and training programs heavily encourage people to post their output on LessWrong. So part of it is slowly becoming a journal, and the felt social norms around posts are morphing to reflect that.
In some ways the equilibrium here is worse, journals have page limits.
I’d love to see the reading time listed on the frontpage. That would make the incentives naturally slide towards shorter posts, as more people would click and it would get more karma. Feels much more decision relevant than when the post was posted.
Naive idea:
Get an LLM to generate a TLDR of the post and after the user finishes reading the post, have a pop-up “Was opening the post worth it, given that you’ve already read the TLDR?”.
xAI claims to have a cluster of 200k GPUs, presumably H100s, online for long enough to train Grok 3.
I think this is faster datacenter scaling than any predictions I’ve heard.
Source: https://x.com/xai/status/1891699715298730482
They don’t claim that Grok 3 was trained on 200K GPUs, and that can’t actually be the case from other things they say. The first 100K H100s were done early Sep 2024, and the subsequent 100K H200s took them 92 days to set up, so early Dec 2024 at the earliest if they started immediately, which they didn’t necessarily. But pretraining of Grok 3 was done by Jan 2025, so there wasn’t enough time with the additional H200s.
There is also a plot where Grok 2 compute is shown slightly above that of GPT-4, so maybe 3e25 FLOPs. And Grok 3 compute is said to be either 10x or 15x that of Grok 2 compute. The 15x figure is given by Musk, who also discussed how Grok 2 was trained with less than 8K GPUs, so possibly he was just talking about the number of GPUs, as opposed to the 10x figure named by a team member that was possibly about the amount of compute. This points to 3e26 FLOPs for Grok 3, which on 100K H100s at 40% utilization would take 3 months, a plausible amount of time if everything worked on almost the first try.
Time needed to build a datacenter given the funding and chips isn’t particularly important for timelines, only for catching up to the frontier (as long as it’s 3 months vs. 6 months and not 18 months). Timelines are constrained by securing more funding for a training system, and designing and manufacturing better chips. Another thing on that presentation was a claim of starting work on another 1.2 GW GB200/GB300 datacenter, which translates to 600K chips. This appears to be more than other LLM labs will construct this year, which might be only about 0.5 GW, except for Google[1], but then Musk didn’t name deadlines for 1.2 GW either. It’s only more concrete than Meta’s 2 GW site in specifying that the chips are Blackwell, so it can’t be about plans for 2027 when better chips will be available.
On a recent podcast, Jeff Dean stated more clearly that their synchronous multi-datacenter training works between metro areas (not just for very-nearby datacenters), and in Dec 2024 they’ve started general availability for 100K TPUv6e clusters. A TPUv6e has similar performance to an H100, and there are two areas being built up in 2025, each with 1 GW of Google datacenters near each other. So there’s potential for 1M H100s or 400K B200s worth of compute, or even double that if these areas or others can be connected with sufficient bandwidth.
Can we assume that Gemini 2.0, GPT-4o, Claude 3.5 and other models with similar performance have a similar compute?
For Claude 3.5, Amodei says the training time cost “a few $10M’s”, which translates to between 1e25 FLOPs (H100, $40M, $4/hour, 30% utilization, BF16) and 1e26 FLOPs (H100, $80M, $2/hour, 50% utilization, FP8), my point estimate is 4e25 FLOPs.
GPT-4o was trained around the same time (late 2023 to very early 2024), and given that the current OpenAI training system seems to take the form of three buildings totaling 100K H100s (the Goodyear, Arizona site), they probably had one of those for 32K H100s, which in 3 months at 40% utilization in BF16 gives 1e26 FLOPs.
Gemini 2.0 was released concurrently with the announcement of general availability of 100K TPUv6e clusters (the instances you can book are much smaller), so they probably have several of them, and Jeff Dean’s remarks suggest they might’ve been able to connect some of them for purposes of pretraining. Each one can contribute 3e26 FLOPs (conservatively assuming BF16). Hassabis noted on some podcast a few months back that scaling compute 10x each generation seems like a good number to fight through the engineering challenges. Gemini 1.0 Ultra was trained on either 77K TPUv4 (according to The Information) or 14 4096-TPUv4 pods (according to EpochAI’s quote from SemiAnalysis), so my point estimate for Gemini 1.0 Ultra is 8e25 FLOPs.
This gives 6e26-9e26 FLOPs for Gemini 2.0 (from 2-3 100K TPUv6e clusters). But unclear if this is what went into Gemini 2.0 Pro or if there is also an unmentioned Gemini 2.0 Ultra down the line.
Thank you. In conditions of extreme uncertainty about the timing and impact of AGI, it’s nice to know at least something definite.
It seems that we are already at the GPT 4.5 level? Except that reasoning models have confused everything, and increasing OOM on output can have the same effect as ~OOM on training, as I understand it.
By the way, you’ve analyzed the scaling of pretraining a lot. But what about inference scaling? It seems that o3 has already used thousands of GPUs to solve tasks in ARC-AGI.
The 200k GPU number has been mentioned since October (Elon tweet, Nvidia announcement), so are you saying that that they managed to get the model trained so fast is what beat the predictions you heard?
Crossposted from https://x.com/JosephMiller_/status/1839085556245950552
1/ Sparse autoencoders trained on the embedding weights of a language model have very interpretable features! We can decompose a token into its top activating features to understand how the model represents the meaning of the token.🧵
2/ To visualize each feature, we project the output direction of the feature onto the token embeddings to find the most similar tokens. We also show the bottom and median tokens by similarity, but they are not very interpretable.
3/ The token “deaf” decomposes into features for audio and disability! None of the examples in this thread are cherry-picked – they were all (really) randomly chosen.
4/ Usually SAEs are trained on the internal activations of a component for billions of different input strings. But here we just train on the rows of the embedding weight matrix (where each row is the embedding for one token).
5/ Most SAEs have many thousands of features. But for our embedding SAE, we only use 2000 features because of our limited dataset. We are essentially compressing the embedding matrix into a smaller, sparser representation.
6/ The reconstructions are not highly accurate – on average we have ~60% variance unexplained (~0.7 cosine similarity) with ~6 features active per token. So more work is needed to see how useful they are.
7/ Note that for this experiment we used the subset of the token embeddings that correspond to English words, so the task is easier—but the results are qualitatively similar when you train on all embeddings.
8/ We also compare to PCA directions and find that the SAE directions are in fact much more interpretable (as we would expect)!
9/ I worked on embedding SAEs at an @apartresearch hackathon in April, with Sajjan Sivia and Chenxing (June) He.
Embedding SAEs were also invented independently by @Michael Pearce.
Claude 3.7′s annoying personality is the first example of accidentally misaligned AI making my life worse. Claude 3.5/3.6 was renowned for its superior personality that made it more pleasant to interact with than ChatGPT.
3.7 has an annoying tendency to do what it thinks you should do, rather than following instructions. I’ve run into this frequently in two coding scenarios:
In Cursor, I ask it to implement some function in a particular file. Even when explicitly instructed not to, it guesses what I want to do next and changes other parts of the code as well.
I’m trying to fix part of my code and I ask it to diagnose a problem and suggest debugging steps. Even when explicitly instructed not to, it will suggest alternative approaches that circumvent the issue, rather than trying to fix the current approach.
I call this misalignment, rather than a capabilities failure, because it seems a step back from previous models and I suspect it is a side effect of training the model to be good at autonomous coding tasks, which may be overriding its compliance with instructions.
It’s concerning that such a high proportion of the best people I know in AI safety go into grantmaking, running upskilling programs or building AI safety courses / talent pipelines.
I expect if we had better ideas about what to do, the best people would mostly just go and do those things directly instead of finding / funding others to do the actual work of preventing AI catastrophe.
Do you have a way of knowing whether the people you know are a representative sample of the whole field? From what I understand there’s way more safety researchers than field builders.
Does anyone have a summary of Eliezer Yudkowsky’s views on weight loss?
There’s a good overview of his views expressed in this manifold thread.
Basically:
Caloric restriction works, however it impedes his productivity (“ability to think”).
Exercise isn’t effective in promoting weight loss or reducing weight gain due to compensatory metabolic throttling during non-exercise times
His fat metabolism is poor, because his fat cells are inclined to leach glucose and triglycerides from his bloodstream to sustain themselves rather than be net contributors, and the effect is that muscle loss makes up the difference, leading to unfavourable implications for body composition, energy, and overall health. In essence “good genetics” for fat loss = “Fat cells that are readily and efficiently broken down for energy”, whereas “bad genetics” for fat loss = “fat cells that are resistant to being used as energy”.
OpenAI IPO delayed over investor fears of Zvi full ‘delenda est’.
BBC Tech News as far as I can tell has not covered any of the recent OpenAI drama about NDAs or employees leaving.
But Scarlett Johansson ‘shocked’ by AI chatbot imitation is now the main headline.
PauseAI UK is hiring: https://pauseai.uk/jobs
LessWrong LLM feature idea: Typo checker
It’s becoming a habit for me to run anything I write through an LLM to check for mistakes before I send it off.
I think the hardest part of implementing this feature well would be to get it to only comment on things that are definitely mistakes / typos. I don’t want a general LLM writing feedback tool built-in to LessWrong.
Don’t most browsers come with spellcheck built in? At least Chrome automatically flags my typos.
LLMs can pick up a much broader class of typos than spelling mistakes.
For example in this comment I wrote “Don’t push the frontier of regulations” when from context I clearly meant to say “Don’t push the frontier of capabilities” I think an LLM could have caught that.
If you missed it, Veo 3, Google’s text-to-video model was just launched and is very impressive. And the videos have audio now.
https://www.reddit.com/r/ChatGPT/comments/1krmsns/wtf_ai_videos_can_have_sound_now_all_from_one/
There are two types of people in this world.
There are people who treat the lock on a public bathroom as a tool for communicating occupancy and a safeguard against accidental attempts to enter when the room is unavailable. For these people the standard protocol is to discern the likely state of engagement of the inner room and then tentatively proceed inside if they detect no signs of human activity.
And there are people who view the lock on a public bathroom as a physical barricade with which to temporarily defend possessed territory. They start by giving the door a hearty push to test the tensile strength of the barrier. On meeting resistance they engage with full force, wringing the handle up and down and slamming into the door with their full body weight. Only once their attempts are thwarted do they reluctantly retreat to find another stall.
Rationalist twitter rage-bait recipe:
I haven’t read the paper yet but I’m pretty confident people are overhyping the Global Workspace paper.
This isn’t even any fault of the paper. It’s just kinda inevitable for any ML paper with a PR team and a cool video promoting it.
But it’s especially true for a paper that is (1) interpretability and (2) uses a bold philosophical framing. The interp literature has a long history (pre-dating mech interp) of techniques going through a cycle of discovery, hype, criticism and skepticism.
Not everyone is overhyping the paper. Neel Nanda’s review seems pretty sober. But I’m surprised to see Zvi writing this:
The J-Lens is improvement of the logit lens. It’s cool, but not a major advance that deserves this level of attention from people who do not typically read mech interp papers.
Once again, I have not read the paper and certainly hope to be pleasantly surprised when I do!
I find Anthropic’s style irritating (Mostly because Mech Interp ppl are not the intended audience of the posts ig), but reluctantly I found it to be a pretty interesting paper on the technical level.
“The J-Lens is an improvement of the logit lens.”: The J-lens is answering a different question from the logit lens, so it’s not best described as just an improvement imo.
The J-lens gives a global linear approximation for the impact that intervening will have on the outputs at future tokens, as opposed to just the output at the current token. This leads to qualitatively different tokens being surfaced in comparison to the logit lens, and their results here are pretty interesting.
Agreed it’s definitely overhyped though, yeah. But that’s what Anthropic does! It’s probably not too bad a trade off in exchange for getting more people into the field (as long as people in field maintain appropriate skepticism).
I do wish that Anthropic would publish a preprint alongside the blog post that could plausibly be submitted to conference, because I think it would be far more readable if they were forced to compress things down to the key experiments.