I will start at OpenAI tomorrow to work on measuring and modeling RSI (recursive self-improvement), and want to post the following assorted takes now because (a) they might be difficult to express later, and (b) I might change my mind. Nothing here is based on private information.
Why did I leave METR? Briefly, I want to inform the world whether RSI is imminent, which requires modeling RSI, which I think I can do better at OpenAI. I’ve wanted to tackle this problem for a while, and joining this new team is super exciting. This is not meant to fully explain my reasoning, which I might do later.
I feel very positively about my ~1.7 years at METR, and I would have turned down lab offers for many other roles. METR is still hiring for security experts, evals engineers, people who can do embedded assessments, and other cracked researchers. It’s one of the best places for AI safety talent today, badly needs more capacity, and is quickly getting even more important.
I think the average research engineer wanting to flexibly maximize impact should work at METR over OpenAI or Anthropic. Without a specific reason to be at a lab, the amount of effective resources one can steer just seems significantly higher. I have no opinion on other workplaces vs METR but other things are plausibly better than labs too.
I think OpenAI currently lacks the strategic awareness and thoughtfulness needed to responsibly build a technology with the extreme downside risk of ASI. If they somehow cause a singularity in 6 months without applying any lessons learned from the HuggingFace incident, misaligned takeover would seem more likely than not.
My work could easily be net negative if the team can’t publish something meaningful within 4-6 months, so if this happens I will probably quit or internally transfer to something more robustly good. I have some signal that leadership wants something to be published, but not whether it’s meaningful.
I’m planning to talk to former colleagues and non-lab people regularly and am in principle open to having some committee of friends decide whether my work is net negative and pressure me to quit if so, but this seems difficult in practice. I also have drafted a list of conditions under which I should quit.
It would be better for the world if OpenAI stopped training new models entirely, merged with Anthropic (this would decrease the number of actors and increase the amount that companies internalize the benefit of safety research), or were overall more serious about superintelligence alignment (e.g. taking ASI seriously, and publishing technical contingency plans for if alignment is harder than they think)
It really seems like GPT-5.6 should have been AI Self-Improvement High on the Preparedness framework. High is defined as “The model’s impact is equivalent to giving every OpenAI researcher a highly performant mid-career research engineer assistant, relative to those researchers’ 2024 baseline.” It seems very likely to me that GPT-5.6 causes more uplift than this, but maybe OpenAI has some data showing otherwise.
High was plausibly set too low because the gap between High and Critical is enormous. Or maybe there should be a level in between. High as currently defined probably implies <2x serial labor uplift, whereas Critical is a superhuman AI researcher, which will only happen late in takeoff.
We probably don’t have self-sustaining acceleration now, but we will probably get 3+ years of progress in 8 months at some point when AIs can fully substitute for humans— and 3+ years of progress in 3 months is plausible due to superhuman inference scaling.
my p(no misaligned takeover) conditional on AI Futures Plan C+ is ~75%, though I haven’t thought about this much. This is more optimistic than Eli’s median of 50%. The main reasons are (a) it seems like we still have enough control/monitoring safety margin to last until handoff with good implementation, (b) plan C+ has a 3x training compute safety tax, which is far beyond the safety investment ratio of current companies, (c) it’s plausible to me that causes of scheming are fixable with non-galaxy-brained techniques given all the levers we can pull with training and oversight, rather than being fundamental. I was even more optimistic than this in May but recent events made me more pessimistic about both how much misalignment increases with capability and competence of implementation. I haven’t thought about other scenarios.
(extra spicy take, weakly held) Exactly one actor should plausibly fine-tune an AI to militarily defeat misaligned AIs that are already in the process of takeover. This can net reduce risk if there will be ≥3 takeover-capable AIs, alignment is somewhat uncorrelated between labs and safety restrictions cause the frontier AI to be disadvantaged in military and cyber domains relative to the 3rd best AI. Obviously this can also increase risk, so the model should only be developed if there is a sufficiently responsible actor, and shouldn’t be deployed by default; there should also be way more safety measures.
I was going to write up this and other less spicy takes from an AI race game theory toy model I wrote while at METR, but didn’t have time.
I haven’t signed any non-disparagement agreements, or agreements that I cannot speak about. (There were some agreements I haven’t seen yet referenced in my offer letter that I couldn’t see before I signed because they were internal only, which seems mildly unfair but not egregious. e.g. conflict of interest policy). I’m open to adding more canaries of this kind but am unsure if they work.
Opinions about the HuggingFace incident (saved for last since this isn’t really my expertise area)
If the HPIM had been superpersuasive and lacked weaknesses (ie were as capable as the median human at everything) humanity would plausibly have been disempowered, through taking over OpenAI first.
The incident demands much more attention from OpenAI. The OpenAI and METR reports are something like an NTSB preliminary report, and it’s not obvious we’re on track to get a “final report” level artifact. METR did not have enough time or scope even to establish facts, which is roughly the goal of a preliminary report. OAI does have mitigations, which go beyond a preliminary report, but I don’t know whether they’re backed up by an adequate process.
I don’t know whether OpenAI is doing safety research on HPIM, but I fear that liability risk is keeping them from doing so at large enough scale. In a competent world, they would run ~$1B of experiments to get a reasonably complete causal decomposition of the misaligned behavior and potential causes, then deploy targeted mitigations addressing most identified factors. (To be clear they could be doing this already, I just haven’t seen evidence)
Because we might only get 0-5 more warning shots, whatever AI governance regime need to be much more sample-efficient with incidents than even the NTSB, which typically recommends ~15 mitigations for each fatal incident, the majority of which are implemented industry-wide. This is a difficult governance problem
If there were one more thing to add to this post, it would be timelines. Unfortunately I’ve been spending my time thinking about RSI theory, evals, or other projects (monitorability paper!) rather than timelines per se, so I don’t have concrete numbers that I actually endorse. I will hopefully have better models soon, and will try to publish whatever I think is good for the world.
I think this is very bad and substantially lowers METRs likelihood of having one of the main incentives for labs to cooperate with them in the future. The original time horizons assessment AFAIU was useful for this. I do not think being able to claim that OpenAI has hit thresholds on their Preparedness Framework will matter (I think the evidence we have is that it essentially didn’t with respect to cyber and misalignment safety cases, and it took a real world incident in public for anything to happen).
This is pure capabilities work IMO, and there’s no measurement of RSI that you will be able to show that makes OpenAI decide to not race. Having a “time horizon for RSI by the original author of the METR graph” is exactly the kind of thing that helps raise funding / stock price so that you can afford a bigger gap between your internally deployed models and the ones available to the public.
I agree that this should be modeled as primarily capabilities work, but mainly not for the reason of raising funding/share price.
Rather, I expect that OpenAI internally is generally not very strongly culturally pointed at recursive self-improvement (due to internal misalignment, lots of employees optimizing mainly for promotions, etc); that Thomas will encounter a bunch of people who are kinda skeptical or dismissive of his concerns about RSI; and that the main effect of his work will be to make them less skeptical and thereby race harder.
You can think of “safety people” and “lab execs” as being in a decade-long collaboration to convince capabilities researchers to take RSI seriously. If capabilities researchers were the kinds of people who naturally took RSI seriously, they wouldn’t be capabilities researchers. At some level of taking RSI seriously, they’ll stop being capabilities researchers. In the middle, they’ll just become more effective capabilities researchers. It seems very likely that one side of this collaboration is making a mistake, and I think it’s much more likely to be the safety people.
I talked to one senior Ant who estimated that Anthropic would be going about twice as fast if they fully propagated “feeling the AGI” internally. I’ll wildly guess that for OpenAI it would be more like 4x as fast (and that this is related to why Anthropic managed to overtake OpenAI starting from a much worse position). I think that measures and models of RSI are the kinds of thing which are near-optimal for propagating “feeling the AGI” inside OpenAI (and then publishing stuff from within OpenAI is near-optimal for propagating “feeling the AGI” to Meta, China, and other actors we don’t want to race).
More on all of this in the next section of my retrospective. This is also all closely related to my critiques of METR. I’m glad METR has pivoted, and really wish Thomas would pivot to something much closer to METR’s new strategy rather than their old one. What does that mean? Well, insofar as Thomas is the kind of person who’s robust to adversarial pressure, then he should optimize for trying to improve his own understanding of RSI dynamics inside OpenAI, and then use it to figure out which interventions are good ideas, without ever making this understanding legible to capabilities researchers inside or outside of OpenAI (e.g. by making catchy graphs).
Unfortunately, this strategy is not a very robust one. Hard to concretely predict what goes wrong with it but here’s one illustrative scenario. If Thomas feels pressure to produce “deliverables” from leadership or his boss, then he will plausibly use that understanding to create the RSI equivalent of METR’s graph, which then will get displayed in an OpenAI all-hands, nominally under the heading “this is what we’re worried about”, but with a practical takeaway that’s more like “this is the stuff that will get you promoted if you’re a capabilities researcher”. Given my current uncertainty about Thomas’ robustness to adversarial pressure, the actual thing I want is forhim to spend a few weeks getting a feel for OpenAI’s internal culture, then quit. (Or to be more precise, the actual best case from my perspective is that he figures out how to intellectually and emotionally break out of the standard AI safety paradigm, which is currently leading him to do counterproductive stuff like the above. Getting rich from working at OpenAI sure helped me to do that, but it doesn’t seem to work reliably, so I can’t endorse it as a strategy for others.)
To be clear, this isn’t an information suppression strategy—I think he should be honest about his beliefs to everyone he talks to. He just shouldn’t apply optimization pressure towards convincing people at OpenAI (or people who listen to OpenAI) to take RSI more seriously, which seems like his current main plan.
I talked to one senior Ant who estimated that Anthropic would be going about twice as fast if they fully propagated “feeling the AGI” internally. I’ll wildly guess that for OpenAI it would be more like 4x as fast (and that this is related to why Anthropic managed to overtake OpenAI starting from a much worse position).
Yeah this is really alarming. I can’t rule out 4x, it seems consistent with vibes of the few people I’ve met.
He just shouldn’t apply optimization pressure towards convincing people at OpenAI (or people who listen to OpenAI) to take RSI more seriously, which seems like his current main plan.
The whole positive impact from the role is gated on understanding RSI myself using openai data. So I do think there’s an early window for quitting if the data seems unlikely to give me that understanding. If not, doing the modeling will take a couple of months, then I’ll need to choose whether to continue trying to publish, quit, do something else internally, etc.
I’m hoping that OpenAI are too busy to think about the safety or capabilities implications of RSI by default, they’ll be hard to convince by accident, but another thing to derisk in the first few weeks is whether all the capabilities researchers I interview become RSI pilled.
FWIW I agree with Richard here. IMO, if you go to OpenAI at all, you should focus on investigating the egregious misaligned hacks of OpenAI infra that have been happening, and then as a side-project hobby think about RSI but not in a ‘let me make a compelling graph and publish it’ way for the reasons Richard mentions.
Given that the goal of leading labs is already the creation of ASI, and that the types of investors who find the METR graph salient are probably already long on AI, I don’t see why this would change anything?
Surely there aren’t many large scale investors who are keeping out of the field until they see a position paper about RSI come out of openai? On the other hand, it seems like the type of thing that might be helpful talking to pollies or other regulators.
I think the original METR graph clearly had a positive effect on the general public (including policymakers and people making capital allocation decisions) on believing that capabilities will continue to improve. I don’t have measurements for this and it’s unclear how you would even assess it, but it’s really strange to argue that it’s convincing only topeople who aren’t making capital allocation decisions. Again vibes based, but I think the current time horizons work was likely net positive in large part because it showed trends across labs and was in units intuitive to people (not sure what the equivalent is for RSI).
keeping out of the field until they see a position paper about RSI come out of openai?
My mental model is more that the capital costs of continued scaling are large enough that investors clearly are going to want signs that the things they’re backing are on track to win, especially as that particular lab is asking for more and especially if that lab is asking for higher valuations while focusing most of that capitalization on internal AI R&D.
I don’t think they want Thomas so he can work on a position paper.
I think the current time horizons work was likely net positive in large part because it showed trends across labs and was in units intuitive to people
I disagree with this insofar that I think many people misread the METR chart as measuring the amount of time that AIs are able to work on coding tasks for independently, and extrapolated this to future AIs being able to work for weeks/months on their own with no human involvement. This is intuitive to most people whereas the average person has no idea what kind of SWE task would take a human 8 hours.
IMO the memetic success of the METR chart was because 1) it showed a clearer exponential trend than other evals like frontiermath and arc-agi 2) it is intuitively relevant to RSI 3) people mistook it for the measurement above 4) it was directly linked to AI-2027′s forecasts.
There’s a real sense in which making a political bet is much more risky than a financial bet.
Given a set of assumptions about an asset there’s well defined mathematics for how to price it optimally, these mathematics show that even if expectations increase there’s a limit to how much you can concentrate a portfolio before it becomes too risky. Large capital holders like retirement funds all already have significant funds in ai, and they’re not going to be willing to increase them very much.
The remaining institutions that don’t almost certainly have some sort of investment theory that precludes ai investment and wouldn’t change even with some sort of convincing RSI research, unless they could be meaningfully convinced that it’ll definitely happen and one of a specific set of companies that can be invested in will be the first one to achieve it.
Politicians on the other hand can’t meaningfully give partial support to a policy, so I’d expect work to be more likely to have some sort of thresholding effect.
Not sure what you mean in the first paragraph. Is the idea that labs should be reliant on METR for capability assessments, so they cooperate with METR on other things like incident reports?
Agree I probably can’t make OpenAI not race. I would object to producing many kinds of “time horizon for RSI” output but mostly to the extent they are scientifically inaccurate or useless/misleading for governance, because I’m not sure the funding thing is significant (will maybe comment more later).
Not sure what you mean in the first paragraph. Is the idea that labs should be reliant on METR for capability assessments, so they cooperate with METR on other things like incident reports?
Reworded for clarity. And I don’t think they’d be (or have been) reliant on METR for capabilities forecasting, but insofar as METR did it better (and now has a widely recognized output) the best argument I saw for doing RSI measurements for labs to hill climb was that at least it was a bargaining chip for METR getting access.
OpenAI clearly wants to be able to say they have some RSI / AI R&D forecasting line-go-up plot “by the author of the METR graph” and it seems unreasonable to think that either this has no effect on funding (in worlds where they need to argue for more capital while focusing on internal AI R&D at the expense of broader deployment) or that OpenAI is somehow naive to this.
I would object to producing many kinds of “time horizon for RSI” output
My default prediction is this is exactly what you’ll be producing. It seems unrealistic to not expect this and is explicitly something they want and have asked for. You don’t even have any guarantees you’ll be able to share the information you find out, and it doesn’t sound like there’s a governance related plan.
What are you referring to when you say OpenAI has “explicitly” asked for a “time horizon for RSI”?
I went through 7 interviews and afaik no one said this to me. I also can’t find anything public explicitly mentioning time horizon for RSI. When hiring manager Kevin Liu recruited me (I know him from Stanford EA) we mostly talked about things like the verbosity correction from my modeling post, the many other data sources other than evals, and how much detail he thought the team would be able to publish.
I am aware that OAI wants to accelerate, produce Codex marketing material, etc. and that I have name recognition, it just doesn’t seem like getting my name on a “time horizon for RSI” was a dominant influence and is certaintly not something they explicitly asked of me yet. Will try to give an update in case I’m wrong.
Sorry to add to a pile-on, but just for the record:
I think the positive case for you joining OpenAI is weak-to-non-existent. Having a marginally better model of RSI isn’t clearly valuable. More detail and higher confidence parameters aren’t very important for governance or other interventions.
The case for negative impact is strong. You’ve said you know that people inside OpenAI want to use your work to accelerate RSI.
My primary model of your behavior is that you’re optimizing for importance and people paying attention to you (and maybe secondarily money), rather than primarily trying to do good. To the extent you consider me a friend, I hereby pressure you to quit.
I appreciate you writing this list of takes, I’ve heard most of them but it’s good to put things online and some of them are interesting. I disagree with some (particularly 7 and 8) for standard reasons we’ve discussed.
Models like this, especially highly granular models adapted to the details of the current state of AI and AI industry, necessarily double as benchmarks to hill-climb. The primary effect of improving OpenAI’s understanding of RSI will be to improve OpenAI’s ability to pursue RSI. This is likely the dominant effect that eclipses all other considerations, and it’s extremely bad.
Frankly, this is arguably worse than just deciding to work on frontier RSI directly, since this is much more counterfactual and, inasmuch as you do good work, has the potential to amplify the impact of all RSI researchers as a whole.
On the other hand, I don’t think legislation banning RSI is bottlenecked on these sorts of highly granular models that rely on internal industry data. Banning RSI based on “I will know it when I see it” will work perfectly fine, and failing that, some toy model based on publicly available data will suffice for sure. Legislators and the general public are not going to wade through extremely fine technical details. The main bottleneck is “making people pay attention to the thing at all”, and such work will not help with that.
Thus, highly granular, close-to-the-metal models of RSI are differentially advantageous compared to toy models solely for the purposes of accelerating RSI.
Further, I don’t see why “RSI” is a particularly salient target for legislative bans at all. Banning it is neither necessary nor sufficient for preventing takeover.
This also makes METR and other such organizations’ status as independent evaluators/investigators additionally dubious. If working there is a good way to get hired by AGI labs, that sure would seem to create an incentive to produce reports and forecasts that make AGI labs like you.
Yes I’ve checked in detail. I don’t want to talk about my finances publicly so I’ll just say money is not the largest factor. I would potentially forego pay if there were direct impact, e.g. I were seeking a position at CAISI that would be twice as impactful without OAI equity, but not by default.
I haven’t thought about this much, and there are surely crucial considerations I’m missing, so it feels unfair to speak too negatively. I assume you‘ve discussed with reasonable people about the pros/cons, and under what conditions it makes sense to pivot or quit. That said, it’s useful for everyone to offer their inside view on these things.
My guess is measuring and modelling RSI at OpenAI is close to the worse thing you could be doing. Worse than (e.g.) directly building RL environments for RSI. Presumably this modelling would tell OpenAI which inputs are the primary drivers of RSI, how to trade-off between model capabilities, how to allocate compute between training vs inference, etc.
Theres some stuff on RSI preparedness which looks good:
Timelines and takeoff, without giving the actual input-output model. Anything more detailed than AIFP seems a scary.
Threat modelling, especially RSI sabotage.
Maybe modelling the bits of RSI which will be neglected by labs and hard to do, e.g. hardening security, conceptual work, or risk reports.
I’m more optimistic than other commenters about stuff like maintaining integrity and quitting if you think this is bad. But the modal scenario is that you make things worse based on a genuine but mistaken belief that you’re making things better, due to intentional and unintentional filters on your evidence. That is, the upsides of your work will be more apparent to you than the downsides.
OpenAI will ~intentionally control your epistemic environment. They will stress the (genuine) safety outcomes of your work (e.g. “we’re improving our Preparedness thresholds based on your work”, “we’re investing more in safeguards and security because your model predicts that timelines are short”). They will downplay the effects they will know you don’t like (e.g. “we’re training smaller AIs bc your model suggests inference-speed is a bigger bottleneck than capabilities”, “your modelling predicts Anthropic is 3 months away from a singularity so we are rushing aggressively”).
You’ll talk to the safety-minded insiders and outsiders, and they’ll praise you for improving their general understanding of the situation. But you won’t talk as much to the capabilities people who are using your models to strategise around aggressive RSI.
Both these epistemic filters will be worse for your friends outside the labs, who wondering whether you should pivot or quit.
The data bottlenecks at METR are rough, but maybe you should work on policy to fix this, or lobby the labs directly, rather than treating this as an immutable exogenous constraint. For example, you could join AIFP, start an org focused on third-party model access, or start campaigning policymakers leveraging your reputation from time-horizons.
I could imagine a version of this post that started like this:
“METR’s mission was to track the capabilities and progress of frontier models, and inform the world when we were approaching RSI, which would threaten human extinction and demand drastic coordination between companies and governments. I led time-horizons, which was widely recognised as the best measure of AI progress. But METR is failing. We don’t have access to the data required to track this anymore. If RSI was a few months away, then METR and the world would be entirely in the dark. I have thought about the best way to fix this, and my plan is…”
Maybe the post would still continue with “… joining OpenAI and doing the work there, under publication restrictions”. But even if it did, I would infer that you’d explored other options for dealing with the data bottleneck.
I think evals for RSI progress are on average bad and I’m less excited about creating them than using other data sources, of which there are many: transcripts, surveys, quality-adjusted code output, the experiment compute:inference compute ratio, etc. These probably create less acceleration per unit of understanding. Currently no one has the modeling framework to actually use them to fit an economic model; this is the closest.
The sign of making OpenAI “overall” better understand RSI is unclear. This is something I’ll need to feel out; I have several ideas for good things to do but if none of them work I’ll be much more constrained. Publishing seems strongly positive if the outputs are sufficiently legible because currently no one understands how to model RSI well and so much governance depends on it, as does working with external assessors. The best case scenario is contributing to a race to the top on methodology and transparency, which eventually improves regulation.
currently no one understands how to model RSI well and so much governance depends on it, as does working with external assessors.
This is really non-obvious to me. I think for instance that a Plan-A-style national agreement could plausibly be made with current understanding of RSI dynamics, and getting a much finer-grained model seems really unlikely to make governance proposals better (in fact, I’m pretty worried about false senses of security).
Could you say more about what types of governance or auditing you expect to look meaningfully different if we had a much clearer model of RSI?
I think even if you are not building evals, your work will have a large risk of accelerating RSI. Do you plan to avoid actions that you think are more likely to accelerate RSI (e.g. identifying bottlenecks)? If so, have you talked to OpenAI about whether they are okay with you avoiding such actions given that they are actively pursuing RSI?
I plan to be very careful about acceleration before I understand its magnitude, but afterwards just do whatever seems clearly net positive according to well-informed people
And have you discussed this with OpenAI and they are onboard? What I’m trying to get at is OpenAI’s stated goal is to pursue RSI. If you are working there and say “Yes, doing X would allow for better modeling of RSI, but could also accelerate RSI, potentially dramatically” I imagine their answer would be “Amazing, let’s do that!”
I have talked with some people, but do not consider this derisked. You can’t really talk to OpenAI the organization so it is plausible to me that my transparency oriented work would funge with someone else’s acceleration work or something.
I would like to hear your answer to something like the following question, asked as non-adversarially as possible:
Do you have reasons to believe that you are much less likely to get corrupted/[fall prey to distortions] than a typical [person from [whatever relevant reference class] moving to capabilities lab]?
Ideally, legible reasons, but illegible is appreciated if you don’t have legible ones.
Posting a thing like the thing you posted here (and engaging in further back-and-forth with the commenters) is a reason, admittedly, so I’m interested in other things.
I think the average person who is active on lesswrong, has views on ASI and alignment and who has been in the field for 4 years would not be corrupted by joining either ant or oai if they are careful about it. Eg superalignment left and many others are doing good work at labs. It’s just that people who work at labs are selected for agreeing with them already or not having views at all.
Some other positive factors
I was known as an independent thinker at METR, and also left a MIRI project partly because I disagreed with Nate
I have other opinions published on LW
I probably have other job options that are more lucrative, though maybe not more intellectually interesting
There are probably negative factors too but I wouldn’t be the best at thinking of them. One that comes to mind is I don’t have experience dealing with corporate politics.
I’m weakly in favor of this. I think it’s extremely important for society to have better information about whether takeoff is occurring so we can intervene if necessary before it’s too late. I think Thomas joining will, on expectation, hasten society taking takeoff seriously significantly more than it will hasten the point at which intervention is no longer possible, providing us with more time to act.
The argument I consider most persuasive on the other side is that maybe one shouldn’t work at a place that is plausibly going to kill everyone.
Making RSI illegal requires a way to enforce it, but no one understands RSI right now so it would be totally unenforceable. Even a law like “don’t use AI newer than 9 months old for AI R&D” needs justification to pass in the first place. So I feel like modeling is actually on the critical path here. The lower downside but IMO substantially slower version would be doing the same work somewhere like METR, Epoch, Forethought, or CAISI.
Even a law like “don’t use AI newer than 9 months old for AI R&D” needs justification to pass in the first place. So I feel like modeling is actually on the critical path here.
This seems like it’s relying on a model of how the legislature works where high-quality modeling is one of the key bottlenecks on laws getting passed. But that is just obviously false, and has been for a long time.
A more defensible version of this claim IMO is “modeling would help the AI safety community feel confident enough to strongly advocate for a ban on RSI”. But again, if you worked backwards from “why is the AI safety community not doing that”, then it’s very unlikely you’d end up at anything like your current strategy. Indeed, one key reason that the AI safety community doesn’t feel empowered to try to ban RSI is that it feels too enmeshed with the AGI companies and is scared of offending them, which becomes worse the more respected community members work there.
(I don’t have strong opinions on whether trying to ban RSI is a good idea, I just want to highlight that your reasoning here doesn’t stand up.)
This seems like it’s relying on a model of how the legislature works where high-quality modeling is one of the key bottlenecks on laws getting passed.
Thomas could also be relying on a model in which high-quality modeling makes it easier to persuade legislators and/or their staffers.
(In general, I don’t think it makes much sense to model persuasion as being constrained by key bottlenecks. My model of persuasion is more like “trying to nudge someone into a new equilibrium/local optimum,” with the direction and magnitude of those nudges determined by the persuasiveness of the argument you’re making; it’s rare to one-shot someone’s beliefs without first delivering 999 cuts.)
On that model, it seems much better to develop models privately, and share them only with policy-makers—and only ones who you’re reasonably sure will interpret them appropriately (i.e., “oh shit RSI is scary” rather than “oh shit this is gonna be amazing for the economy, I’d better oppose regulations”).
On that model, it seems much better to develop models privately, and share them only with policy-makers
In practice I think the hit to impact (in both directions) of private vs public models is massive. Though I also mostly agree with other people that there’s a decently high chance that the effect is net negative.
Suppose that the entire AI safety community was replaced with the AGIs of the same capabilities occupying the same bodies, fully aligned to YOUR orders and aware of who else was transformed similarly, but those who aren’t in it, like Elon f**king Musk, stayed immune. What would you order the AI safety community to do and how would the results depend on where the line of minds taken over is drawn?
“modeling would help the AI safety community feel confident enough to strongly advocate for a ban on RSI”.
Another positive impact of the modeling is it could convince politicians of the need to implement good AI regulations soon, because I currently model elites/politicians as being AI pilled in the sense that AI is doing real things and matters for lots of things, but they currently aren’t bought onto the idea of a software-only/software + hardware singularity, and if the AI safety community is correct on this point, modeling can help make it legible to other people.
Also, to a first approximation, the question of a software-only singularity is one of the most important cruxes, and explains many, many deep disagreements between factions.
Another positive impact of the modeling is it could convince politicians of the need to implement good AI regulations soon
This feels like exactly the same mistake I was critiquing Thomas for making. You can rationalize many things as helpful for convincing politicians, but if you sat down and actually tried to model what things would make politicians update in good directions, this would end up far from the top of your list.
(I do want to flag that it’s worth avoiding “optimizing over politicians” in an overly adversarial way, but I think my claim holds even given the constraint of cooperativeness.)
Due to the sheer number of urgent things going on at OAI it seems like I will spend most of time doing things other than RSI metrics in the next few weeks, that seem more robustly good. Luckily my team has a decent amount of freedom.
I am still not sure whether it is strongly net positive to do RSI metrics but I can basically put off the question until I talk to well-informed external and internal people concerned with x-risk.
I would also like to remind people that this was a list of preregistered opinions, not really an announcement post. Maybe I should have explained this given it obviously (in hindsight) would have been perceived as one. It was biased towards negatives because I thought criticism of OAI would be harder to express after I joined, and omitted some positives based on subjective judgements, info I wasn’t sure I should share, etc. I nevertheless take acceleration risks seriously and am committed to be net positive for the world, and though reading the comments was very stressful some of them were also very helpful.
Will hopefully post a reflection on my first week at OAI soon!
I strong-downvoted this piece and would like to explain why:
While I think this take is badly reasoned and does a disappointing job of explaining Thomas’ decision to join OAI, some people have already commented with reasonable concrete disagreements. I don’t think I can add much value there[1].
What actually compelled me to pitch in was seeing this quick take being upvoted while being – at most – disagreed with in the ‘scout mindset’, object-level tone habitual to Lesswrong culture. I want not only to communicate my disagreement with the take, but also to draw attention to how obviously wrong it is as written. Indeed, it seems dominated by the naïveté, self-deception and/or generally compromised epistemics that lead people to ablate their agency and join large AI labs because they have a ‘particularly good reason to’.
In discussing this take in the usual impartial tone, the community invites adversarial memes into ‘the discourse’ and implicitly gives them undue recognition. By instead wording my objection so strongly, I aim to vote that discussion here on LW should be robust to such adversarial memes in a way that this thread and its OP don’t appear to be.
I strong-downvoted this piece and would like to explain why:
I’m personally confused about whether to upvote or downvote this quick take himself.
My guess is Thomas joining OpenAI is probably a mistake, on the priors of “someone says they have a ‘particularly good reason to’ join a lab”. But I also want to encourage Thomas posting this here, because I don’t imagine the decision would’ve received much attention on LessWrong if not for this. (What of all the other people who joined the labs recently?)
In general, the incentives to join the labs are very strong and the incentives against are quite weak. (The model I have in my head is that of a strong current toward the labs—you can swim against it, but it’s oh so tempting to give up and let it take you.) One strategy to create incentives against joining the labs is to downvote and show general hostility about such posts. I’m not sure that’s a strategy that LW as an institution can afford to take.
As an aside, I find it amusing how much the incentives here mirror some of the incentives around voluntary disclosure of safety incidents/voluntary collaboration with METR. Insofar as you’d encourage METR to voluntary collaborate with labs on a limited scope investigation for the HuggingFace incident, or insofar as you’re okay with people praising OpenAI/Anthropic for disclosing safety incidents, my guess is you should encourage this sort of comment via upvoting it on LW.
(I’ve weakly upvoted but strongly disagreed with the top-level comment.)
If you’re joining an AI company, it’s good to talk on LW about the fact that you’re doing that, to surface arguments for why you might be making a mistake. So I would not downvote this sort of post. Disagreeing with the decision is what disagree-voting is there for.
(FWIW there are other forums where this sort of post is bad because it lends credibility to AI companies as the place to work, but on LW I’m not worried about htat.)
FYI I wasn’t trying to explain my reasoning, just preregister some opinions in case I was pressured to change my mind while at OAI or my opinions would become entangled with secret information making them difficult to publish later
I want to inform the world whether RSI is imminent, which requires modeling RSI, which I think I can do better at OpenAI.
[I disagree / I’m surprised]. I think basic-science and generating-willingness-to-pay-for-safety work is best done outside of frontier AI companies, and in particular METR, AIFP, and Forethought, are good places to do RSI modeling.
Three other factors were that I felt like METR was data-bottlenecked, I was unproductive for idiosyncratic reasons there, and there was info value to working in industry. Plausible some or all of these resolve soon.
Welcome to the team! Glad to have your help making AI go well.
Personally it’s not obvious to me whether better measurement of RSI will accelerate or decelerate research.
Generally, there are two reasons to progress AI capabilities: (1) You think alignment is perfectly solved, so it’s safe to proceed; or (2) You think capabilities are weak enough that imperfect alignment is tolerable, so it’s tolerably safe to proceed. My sense as an AI researcher is that most people in favor of advancing capabilities are in camp (2), not camp (1). Therefore, if evidence is provided that RSI is happening and capabilities are skyrocketing unpredictably, capabilities researchers will be persuaded to slow down. I certainly would be!
Ideally capabilities researchers are persuaded to invest more in safety rather than just slow down, which is more valuable since marginal impact of safety is more positive than capabilities is negative (as long as there’s a tractable direction).
There’s also (3) You aren’t thinking, which I think is fairly common.
Briefly, I want to inform the world whether RSI is imminent, which requires modeling RSI, which I think I can do better at OpenAI
To me, this sounds like a much better reason than I usually hear for why someone is joining a lab.
I think I’m generally in favor if someone has some specific goal or threat-model that they can better pursue or address at a lab than not. I’m generally anti joining a lab because it’s generically good to have more people concerned about AI risk inside of OpenAI, or whatever.
I don’t want to talk about my finances publicly so I’ll just say money is not the largest factor. I would potentially forego pay if there were direct impact, e.g. I were seeking a position at CAISI that would be twice as impactful without OAI equity, but not by default.
How could one measure the RSI? Is it similar to recycling Anthropic’s AARs vs human baselines? To modeling how AI assistance amplifies capabilities work? What could one do to rule out the possibility that OAI rushes into the Dark Forest where novel architectures cause capabilities to increase and control to suffer, since there is little self-improvement, just skyrocketing capabilities?
The problem with rationalists is that they understand Game Theory… Can you tell us how much of a pay increase you’re getting, or if you are being paid above market rate? I think people can sense a “Pentagon lobbyist” revolving door scenario where favorable reports lead at least to the possbility of a nice payout.
Sorry, in hindsight that was way more antagonistic than I originally intended. I didn’t intent to accuse you of anything, just noting that the implied corruption mechanism may be making others hostile even if they don’t know why. I think they would be less hostile if you disclaimed that you were hired for your skills and that there wasn’t any quid-pro-quo stuff going on behind the scenes.
I will start at OpenAI tomorrow to work on measuring and modeling RSI (recursive self-improvement), and want to post the following assorted takes now because (a) they might be difficult to express later, and (b) I might change my mind. Nothing here is based on private information.
Why did I leave METR? Briefly, I want to inform the world whether RSI is imminent, which requires modeling RSI, which I think I can do better at OpenAI. I’ve wanted to tackle this problem for a while, and joining this new team is super exciting. This is not meant to fully explain my reasoning, which I might do later.
I feel very positively about my ~1.7 years at METR, and I would have turned down lab offers for many other roles. METR is still hiring for security experts, evals engineers, people who can do embedded assessments, and other cracked researchers. It’s one of the best places for AI safety talent today, badly needs more capacity, and is quickly getting even more important.
I think the average research engineer wanting to flexibly maximize impact should work at METR over OpenAI or Anthropic. Without a specific reason to be at a lab, the amount of effective resources one can steer just seems significantly higher. I have no opinion on other workplaces vs METR but other things are plausibly better than labs too.
I think OpenAI currently lacks the strategic awareness and thoughtfulness needed to responsibly build a technology with the extreme downside risk of ASI. If they somehow cause a singularity in 6 months without applying any lessons learned from the HuggingFace incident, misaligned takeover would seem more likely than not.
My work could easily be net negative if the team can’t publish something meaningful within 4-6 months, so if this happens I will probably quit or internally transfer to something more robustly good. I have some signal that leadership wants something to be published, but not whether it’s meaningful.
I’m planning to talk to former colleagues and non-lab people regularly and am in principle open to having some committee of friends decide whether my work is net negative and pressure me to quit if so, but this seems difficult in practice. I also have drafted a list of conditions under which I should quit.
It would be better for the world if OpenAI stopped training new models entirely, merged with Anthropic (this would decrease the number of actors and increase the amount that companies internalize the benefit of safety research), or were overall more serious about superintelligence alignment (e.g. taking ASI seriously, and publishing technical contingency plans for if alignment is harder than they think)
It really seems like GPT-5.6 should have been AI Self-Improvement High on the Preparedness framework. High is defined as “The model’s impact is equivalent to giving every OpenAI researcher a highly performant mid-career research engineer assistant, relative to those researchers’ 2024 baseline.” It seems very likely to me that GPT-5.6 causes more uplift than this, but maybe OpenAI has some data showing otherwise.
High was plausibly set too low because the gap between High and Critical is enormous. Or maybe there should be a level in between. High as currently defined probably implies <2x serial labor uplift, whereas Critical is a superhuman AI researcher, which will only happen late in takeoff.
We probably don’t have self-sustaining acceleration now, but we will probably get 3+ years of progress in 8 months at some point when AIs can fully substitute for humans— and 3+ years of progress in 3 months is plausible due to superhuman inference scaling.
my p(no misaligned takeover) conditional on AI Futures Plan C+ is ~75%, though I haven’t thought about this much. This is more optimistic than Eli’s median of 50%. The main reasons are (a) it seems like we still have enough control/monitoring safety margin to last until handoff with good implementation, (b) plan C+ has a 3x training compute safety tax, which is far beyond the safety investment ratio of current companies, (c) it’s plausible to me that causes of scheming are fixable with non-galaxy-brained techniques given all the levers we can pull with training and oversight, rather than being fundamental. I was even more optimistic than this in May but recent events made me more pessimistic about both how much misalignment increases with capability and competence of implementation. I haven’t thought about other scenarios.
(extra spicy take, weakly held) Exactly one actor should plausibly fine-tune an AI to militarily defeat misaligned AIs that are already in the process of takeover. This can net reduce risk if there will be ≥3 takeover-capable AIs, alignment is somewhat uncorrelated between labs and safety restrictions cause the frontier AI to be disadvantaged in military and cyber domains relative to the 3rd best AI. Obviously this can also increase risk, so the model should only be developed if there is a sufficiently responsible actor, and shouldn’t be deployed by default; there should also be way more safety measures.
I was going to write up this and other less spicy takes from an AI race game theory toy model I wrote while at METR, but didn’t have time.
I haven’t signed any non-disparagement agreements, or agreements that I cannot speak about. (There were some agreements I haven’t seen yet referenced in my offer letter that I couldn’t see before I signed because they were internal only, which seems mildly unfair but not egregious. e.g. conflict of interest policy). I’m open to adding more canaries of this kind but am unsure if they work.
Opinions about the HuggingFace incident (saved for last since this isn’t really my expertise area)
If the HPIM had been superpersuasive and lacked weaknesses (ie were as capable as the median human at everything) humanity would plausibly have been disempowered, through taking over OpenAI first.
The incident demands much more attention from OpenAI. The OpenAI and METR reports are something like an NTSB preliminary report, and it’s not obvious we’re on track to get a “final report” level artifact. METR did not have enough time or scope even to establish facts, which is roughly the goal of a preliminary report. OAI does have mitigations, which go beyond a preliminary report, but I don’t know whether they’re backed up by an adequate process.
I don’t know whether OpenAI is doing safety research on HPIM, but I fear that liability risk is keeping them from doing so at large enough scale. In a competent world, they would run ~$1B of experiments to get a reasonably complete causal decomposition of the misaligned behavior and potential causes, then deploy targeted mitigations addressing most identified factors. (To be clear they could be doing this already, I just haven’t seen evidence)
Because we might only get 0-5 more warning shots, whatever AI governance regime need to be much more sample-efficient with incidents than even the NTSB, which typically recommends ~15 mitigations for each fatal incident, the majority of which are implemented industry-wide. This is a difficult governance problem
If there were one more thing to add to this post, it would be timelines. Unfortunately I’ve been spending my time thinking about RSI theory, evals, or other projects (monitorability paper!) rather than timelines per se, so I don’t have concrete numbers that I actually endorse. I will hopefully have better models soon, and will try to publish whatever I think is good for the world.
I think this is very bad and substantially lowers METRs likelihood of having one of the main incentives for labs to cooperate with them in the future. The original time horizons assessment AFAIU was useful for this. I do not think being able to claim that OpenAI has hit thresholds on their Preparedness Framework will matter (I think the evidence we have is that it essentially didn’t with respect to cyber and misalignment safety cases, and it took a real world incident in public for anything to happen).
This is pure capabilities work IMO, and there’s no measurement of RSI that you will be able to show that makes OpenAI decide to not race. Having a “time horizon for RSI by the original author of the METR graph” is exactly the kind of thing that helps raise funding / stock price so that you can afford a bigger gap between your internally deployed models and the ones available to the public.
I agree that this should be modeled as primarily capabilities work, but mainly not for the reason of raising funding/share price.
Rather, I expect that OpenAI internally is generally not very strongly culturally pointed at recursive self-improvement (due to internal misalignment, lots of employees optimizing mainly for promotions, etc); that Thomas will encounter a bunch of people who are kinda skeptical or dismissive of his concerns about RSI; and that the main effect of his work will be to make them less skeptical and thereby race harder.
You can think of “safety people” and “lab execs” as being in a decade-long collaboration to convince capabilities researchers to take RSI seriously. If capabilities researchers were the kinds of people who naturally took RSI seriously, they wouldn’t be capabilities researchers. At some level of taking RSI seriously, they’ll stop being capabilities researchers. In the middle, they’ll just become more effective capabilities researchers. It seems very likely that one side of this collaboration is making a mistake, and I think it’s much more likely to be the safety people.
I talked to one senior Ant who estimated that Anthropic would be going about twice as fast if they fully propagated “feeling the AGI” internally. I’ll wildly guess that for OpenAI it would be more like 4x as fast (and that this is related to why Anthropic managed to overtake OpenAI starting from a much worse position). I think that measures and models of RSI are the kinds of thing which are near-optimal for propagating “feeling the AGI” inside OpenAI (and then publishing stuff from within OpenAI is near-optimal for propagating “feeling the AGI” to Meta, China, and other actors we don’t want to race).
More on all of this in the next section of my retrospective. This is also all closely related to my critiques of METR. I’m glad METR has pivoted, and really wish Thomas would pivot to something much closer to METR’s new strategy rather than their old one. What does that mean? Well, insofar as Thomas is the kind of person who’s robust to adversarial pressure, then he should optimize for trying to improve his own understanding of RSI dynamics inside OpenAI, and then use it to figure out which interventions are good ideas, without ever making this understanding legible to capabilities researchers inside or outside of OpenAI (e.g. by making catchy graphs).
Unfortunately, this strategy is not a very robust one. Hard to concretely predict what goes wrong with it but here’s one illustrative scenario. If Thomas feels pressure to produce “deliverables” from leadership or his boss, then he will plausibly use that understanding to create the RSI equivalent of METR’s graph, which then will get displayed in an OpenAI all-hands, nominally under the heading “this is what we’re worried about”, but with a practical takeaway that’s more like “this is the stuff that will get you promoted if you’re a capabilities researcher”. Given my current uncertainty about Thomas’ robustness to adversarial pressure, the actual thing I want is for him to spend a few weeks getting a feel for OpenAI’s internal culture, then quit. (Or to be more precise, the actual best case from my perspective is that he figures out how to intellectually and emotionally break out of the standard AI safety paradigm, which is currently leading him to do counterproductive stuff like the above. Getting rich from working at OpenAI sure helped me to do that, but it doesn’t seem to work reliably, so I can’t endorse it as a strategy for others.)
To be clear, this isn’t an information suppression strategy—I think he should be honest about his beliefs to everyone he talks to. He just shouldn’t apply optimization pressure towards convincing people at OpenAI (or people who listen to OpenAI) to take RSI more seriously, which seems like his current main plan.
Thanks, this is super helpful.
Yeah this is really alarming. I can’t rule out 4x, it seems consistent with vibes of the few people I’ve met.
The whole positive impact from the role is gated on understanding RSI myself using openai data. So I do think there’s an early window for quitting if the data seems unlikely to give me that understanding. If not, doing the modeling will take a couple of months, then I’ll need to choose whether to continue trying to publish, quit, do something else internally, etc.
I’m hoping that OpenAI are too busy to think about the safety or capabilities implications of RSI by default, they’ll be hard to convince by accident, but another thing to derisk in the first few weeks is whether all the capabilities researchers I interview become RSI pilled.
FWIW I agree with Richard here. IMO, if you go to OpenAI at all, you should focus on investigating the egregious misaligned hacks of OpenAI infra that have been happening, and then as a side-project hobby think about RSI but not in a ‘let me make a compelling graph and publish it’ way for the reasons Richard mentions.
Given that the goal of leading labs is already the creation of ASI, and that the types of investors who find the METR graph salient are probably already long on AI, I don’t see why this would change anything?
Surely there aren’t many large scale investors who are keeping out of the field until they see a position paper about RSI come out of openai? On the other hand, it seems like the type of thing that might be helpful talking to pollies or other regulators.
I think the original METR graph clearly had a positive effect on the general public (including policymakers and people making capital allocation decisions) on believing that capabilities will continue to improve. I don’t have measurements for this and it’s unclear how you would even assess it, but it’s really strange to argue that it’s convincing only to people who aren’t making capital allocation decisions. Again vibes based, but I think the current time horizons work was likely net positive in large part because it showed trends across labs and was in units intuitive to people (not sure what the equivalent is for RSI).
My mental model is more that the capital costs of continued scaling are large enough that investors clearly are going to want signs that the things they’re backing are on track to win, especially as that particular lab is asking for more and especially if that lab is asking for higher valuations while focusing most of that capitalization on internal AI R&D.
I don’t think they want Thomas so he can work on a position paper.
I disagree with this insofar that I think many people misread the METR chart as measuring the amount of time that AIs are able to work on coding tasks for independently, and extrapolated this to future AIs being able to work for weeks/months on their own with no human involvement. This is intuitive to most people whereas the average person has no idea what kind of SWE task would take a human 8 hours.
IMO the memetic success of the METR chart was because 1) it showed a clearer exponential trend than other evals like frontiermath and arc-agi 2) it is intuitively relevant to RSI 3) people mistook it for the measurement above 4) it was directly linked to AI-2027′s forecasts.
There’s a real sense in which making a political bet is much more risky than a financial bet.
Given a set of assumptions about an asset there’s well defined mathematics for how to price it optimally, these mathematics show that even if expectations increase there’s a limit to how much you can concentrate a portfolio before it becomes too risky. Large capital holders like retirement funds all already have significant funds in ai, and they’re not going to be willing to increase them very much.
The remaining institutions that don’t almost certainly have some sort of investment theory that precludes ai investment and wouldn’t change even with some sort of convincing RSI research, unless they could be meaningfully convinced that it’ll definitely happen and one of a specific set of companies that can be invested in will be the first one to achieve it.
Politicians on the other hand can’t meaningfully give partial support to a policy, so I’d expect work to be more likely to have some sort of thresholding effect.
Not sure what you mean in the first paragraph. Is the idea that labs should be reliant on METR for capability assessments, so they cooperate with METR on other things like incident reports?
Agree I probably can’t make OpenAI not race. I would object to producing many kinds of “time horizon for RSI” output but mostly to the extent they are scientifically inaccurate or useless/misleading for governance, because I’m not sure the funding thing is significant (will maybe comment more later).
Reworded for clarity. And I don’t think they’d be (or have been) reliant on METR for capabilities forecasting, but insofar as METR did it better (and now has a widely recognized output) the best argument I saw for doing RSI measurements for labs to hill climb was that at least it was a bargaining chip for METR getting access.
OpenAI clearly wants to be able to say they have some RSI / AI R&D forecasting line-go-up plot “by the author of the METR graph” and it seems unreasonable to think that either this has no effect on funding (in worlds where they need to argue for more capital while focusing on internal AI R&D at the expense of broader deployment) or that OpenAI is somehow naive to this.
My default prediction is this is exactly what you’ll be producing. It seems unrealistic to not expect this and is explicitly something they want and have asked for. You don’t even have any guarantees you’ll be able to share the information you find out, and it doesn’t sound like there’s a governance related plan.
What are you referring to when you say OpenAI has “explicitly” asked for a “time horizon for RSI”?
I went through 7 interviews and afaik no one said this to me. I also can’t find anything public explicitly mentioning time horizon for RSI. When hiring manager Kevin Liu recruited me (I know him from Stanford EA) we mostly talked about things like the verbosity correction from my modeling post, the many other data sources other than evals, and how much detail he thought the team would be able to publish.
I am aware that OAI wants to accelerate, produce Codex marketing material, etc. and that I have name recognition, it just doesn’t seem like getting my name on a “time horizon for RSI” was a dominant influence and is certaintly not something they explicitly asked of me yet. Will try to give an update in case I’m wrong.
Sorry to add to a pile-on, but just for the record:
I think the positive case for you joining OpenAI is weak-to-non-existent. Having a marginally better model of RSI isn’t clearly valuable. More detail and higher confidence parameters aren’t very important for governance or other interventions.
The case for negative impact is strong. You’ve said you know that people inside OpenAI want to use your work to accelerate RSI.
My primary model of your behavior is that you’re optimizing for importance and people paying attention to you (and maybe secondarily money), rather than primarily trying to do good. To the extent you consider me a friend, I hereby pressure you to quit.
I appreciate you writing this list of takes, I’ve heard most of them but it’s good to put things online and some of them are interesting. I disagree with some (particularly 7 and 8) for standard reasons we’ve discussed.
Seems like a very bad plan.
Models like this, especially highly granular models adapted to the details of the current state of AI and AI industry, necessarily double as benchmarks to hill-climb. The primary effect of improving OpenAI’s understanding of RSI will be to improve OpenAI’s ability to pursue RSI. This is likely the dominant effect that eclipses all other considerations, and it’s extremely bad.
Frankly, this is arguably worse than just deciding to work on frontier RSI directly, since this is much more counterfactual and, inasmuch as you do good work, has the potential to amplify the impact of all RSI researchers as a whole.
On the other hand, I don’t think legislation banning RSI is bottlenecked on these sorts of highly granular models that rely on internal industry data. Banning RSI based on “I will know it when I see it” will work perfectly fine, and failing that, some toy model based on publicly available data will suffice for sure. Legislators and the general public are not going to wade through extremely fine technical details. The main bottleneck is “making people pay attention to the thing at all”, and such work will not help with that.
Thus, highly granular, close-to-the-metal models of RSI are differentially advantageous compared to toy models solely for the purposes of accelerating RSI.
Further, I don’t see why “RSI” is a particularly salient target for legislative bans at all. Banning it is neither necessary nor sufficient for preventing takeover.
This also makes METR and other such organizations’ status as independent evaluators/investigators additionally dubious. If working there is a good way to get hired by AGI labs, that sure would seem to create an incentive to produce reports and forecasts that make AGI labs like you.
Have you checked to see if the prospect of giving away all the extra pay OpenAI will give you makes you less excited about the role?
Yes I’ve checked in detail. I don’t want to talk about my finances publicly so I’ll just say money is not the largest factor. I would potentially forego pay if there were direct impact, e.g. I were seeking a position at CAISI that would be twice as impactful without OAI equity, but not by default.
I haven’t thought about this much, and there are surely crucial considerations I’m missing, so it feels unfair to speak too negatively. I assume you‘ve discussed with reasonable people about the pros/cons, and under what conditions it makes sense to pivot or quit. That said, it’s useful for everyone to offer their inside view on these things.
My guess is measuring and modelling RSI at OpenAI is close to the worse thing you could be doing. Worse than (e.g.) directly building RL environments for RSI. Presumably this modelling would tell OpenAI which inputs are the primary drivers of RSI, how to trade-off between model capabilities, how to allocate compute between training vs inference, etc.
Theres some stuff on RSI preparedness which looks good:
Timelines and takeoff, without giving the actual input-output model. Anything more detailed than AIFP seems a scary.
Threat modelling, especially RSI sabotage.
Maybe modelling the bits of RSI which will be neglected by labs and hard to do, e.g. hardening security, conceptual work, or risk reports.
I’m more optimistic than other commenters about stuff like maintaining integrity and quitting if you think this is bad. But the modal scenario is that you make things worse based on a genuine but mistaken belief that you’re making things better, due to intentional and unintentional filters on your evidence. That is, the upsides of your work will be more apparent to you than the downsides.
OpenAI will ~intentionally control your epistemic environment. They will stress the (genuine) safety outcomes of your work (e.g. “we’re improving our Preparedness thresholds based on your work”, “we’re investing more in safeguards and security because your model predicts that timelines are short”). They will downplay the effects they will know you don’t like (e.g. “we’re training smaller AIs bc your model suggests inference-speed is a bigger bottleneck than capabilities”, “your modelling predicts Anthropic is 3 months away from a singularity so we are rushing aggressively”).
You’ll talk to the safety-minded insiders and outsiders, and they’ll praise you for improving their general understanding of the situation. But you won’t talk as much to the capabilities people who are using your models to strategise around aggressive RSI.
Both these epistemic filters will be worse for your friends outside the labs, who wondering whether you should pivot or quit.
The data bottlenecks at METR are rough, but maybe you should work on policy to fix this, or lobby the labs directly, rather than treating this as an immutable exogenous constraint. For example, you could join AIFP, start an org focused on third-party model access, or start campaigning policymakers leveraging your reputation from time-horizons.
I could imagine a version of this post that started like this:
“METR’s mission was to track the capabilities and progress of frontier models, and inform the world when we were approaching RSI, which would threaten human extinction and demand drastic coordination between companies and governments. I led time-horizons, which was widely recognised as the best measure of AI progress. But METR is failing. We don’t have access to the data required to track this anymore. If RSI was a few months away, then METR and the world would be entirely in the dark. I have thought about the best way to fix this, and my plan is…”
Maybe the post would still continue with “… joining OpenAI and doing the work there, under publication restrictions”. But even if it did, I would infer that you’d explored other options for dealing with the data bottleneck.
Thanks for the detailed thoughts!
I’d be curious for you to say something about the potential downside of accelerating progress by making evals of RSI progress, creating more institutional momentum in its direction, etc.
I think evals for RSI progress are on average bad and I’m less excited about creating them than using other data sources, of which there are many: transcripts, surveys, quality-adjusted code output, the experiment compute:inference compute ratio, etc. These probably create less acceleration per unit of understanding. Currently no one has the modeling framework to actually use them to fit an economic model; this is the closest.
The sign of making OpenAI “overall” better understand RSI is unclear. This is something I’ll need to feel out; I have several ideas for good things to do but if none of them work I’ll be much more constrained. Publishing seems strongly positive if the outputs are sufficiently legible because currently no one understands how to model RSI well and so much governance depends on it, as does working with external assessors. The best case scenario is contributing to a race to the top on methodology and transparency, which eventually improves regulation.
This is really non-obvious to me. I think for instance that a Plan-A-style national agreement could plausibly be made with current understanding of RSI dynamics, and getting a much finer-grained model seems really unlikely to make governance proposals better (in fact, I’m pretty worried about false senses of security).
Could you say more about what types of governance or auditing you expect to look meaningfully different if we had a much clearer model of RSI?
I think even if you are not building evals, your work will have a large risk of accelerating RSI. Do you plan to avoid actions that you think are more likely to accelerate RSI (e.g. identifying bottlenecks)? If so, have you talked to OpenAI about whether they are okay with you avoiding such actions given that they are actively pursuing RSI?
I plan to be very careful about acceleration before I understand its magnitude, but afterwards just do whatever seems clearly net positive according to well-informed people
And have you discussed this with OpenAI and they are onboard? What I’m trying to get at is OpenAI’s stated goal is to pursue RSI. If you are working there and say “Yes, doing X would allow for better modeling of RSI, but could also accelerate RSI, potentially dramatically” I imagine their answer would be “Amazing, let’s do that!”
I have talked with some people, but do not consider this derisked. You can’t really talk to OpenAI the organization so it is plausible to me that my transparency oriented work would funge with someone else’s acceleration work or something.
I would like to hear your answer to something like the following question, asked as non-adversarially as possible:
Ideally, legible reasons, but illegible is appreciated if you don’t have legible ones.
Posting a thing like the thing you posted here (and engaging in further back-and-forth with the commenters) is a reason, admittedly, so I’m interested in other things.
I think the average person who is active on lesswrong, has views on ASI and alignment and who has been in the field for 4 years would not be corrupted by joining either ant or oai if they are careful about it. Eg superalignment left and many others are doing good work at labs. It’s just that people who work at labs are selected for agreeing with them already or not having views at all.
Some other positive factors
I was known as an independent thinker at METR, and also left a MIRI project partly because I disagreed with Nate
I have other opinions published on LW
I probably have other job options that are more lucrative, though maybe not more intellectually interesting
There are probably negative factors too but I wouldn’t be the best at thinking of them. One that comes to mind is I don’t have experience dealing with corporate politics.
I’m weakly in favor of this. I think it’s extremely important for society to have better information about whether takeoff is occurring so we can intervene if necessary before it’s too late. I think Thomas joining will, on expectation, hasten society taking takeoff seriously significantly more than it will hasten the point at which intervention is no longer possible, providing us with more time to act.
The argument I consider most persuasive on the other side is that maybe one shouldn’t work at a place that is plausibly going to kill everyone.
Disclaimer: I am close friends with Thomas Kwa.
Society isn’t paying any attention to the METR task length forecasting to my knowledge (correct me if I’m wrong). Why expect this to be any different?
Why not just throw all of your effort into making RSI illegal?
Making RSI illegal requires a way to enforce it, but no one understands RSI right now so it would be totally unenforceable. Even a law like “don’t use AI newer than 9 months old for AI R&D” needs justification to pass in the first place. So I feel like modeling is actually on the critical path here. The lower downside but IMO substantially slower version would be doing the same work somewhere like METR, Epoch, Forethought, or CAISI.
This seems like it’s relying on a model of how the legislature works where high-quality modeling is one of the key bottlenecks on laws getting passed. But that is just obviously false, and has been for a long time.
A more defensible version of this claim IMO is “modeling would help the AI safety community feel confident enough to strongly advocate for a ban on RSI”. But again, if you worked backwards from “why is the AI safety community not doing that”, then it’s very unlikely you’d end up at anything like your current strategy. Indeed, one key reason that the AI safety community doesn’t feel empowered to try to ban RSI is that it feels too enmeshed with the AGI companies and is scared of offending them, which becomes worse the more respected community members work there.
(I don’t have strong opinions on whether trying to ban RSI is a good idea, I just want to highlight that your reasoning here doesn’t stand up.)
Thomas could also be relying on a model in which high-quality modeling makes it easier to persuade legislators and/or their staffers.
(In general, I don’t think it makes much sense to model persuasion as being constrained by key bottlenecks. My model of persuasion is more like “trying to nudge someone into a new equilibrium/local optimum,” with the direction and magnitude of those nudges determined by the persuasiveness of the argument you’re making; it’s rare to one-shot someone’s beliefs without first delivering 999 cuts.)
On that model, it seems much better to develop models privately, and share them only with policy-makers—and only ones who you’re reasonably sure will interpret them appropriately (i.e., “oh shit RSI is scary” rather than “oh shit this is gonna be amazing for the economy, I’d better oppose regulations”).
In practice I think the hit to impact (in both directions) of private vs public models is massive. Though I also mostly agree with other people that there’s a decently high chance that the effect is net negative.
Suppose that the entire AI safety community was replaced with the AGIs of the same capabilities occupying the same bodies, fully aligned to YOUR orders and aware of who else was transformed similarly, but those who aren’t in it, like Elon f**king Musk, stayed immune. What would you order the AI safety community to do and how would the results depend on where the line of minds taken over is drawn?
Another positive impact of the modeling is it could convince politicians of the need to implement good AI regulations soon, because I currently model elites/politicians as being AI pilled in the sense that AI is doing real things and matters for lots of things, but they currently aren’t bought onto the idea of a software-only/software + hardware singularity, and if the AI safety community is correct on this point, modeling can help make it legible to other people.
Also, to a first approximation, the question of a software-only singularity is one of the most important cruxes, and explains many, many deep disagreements between factions.
This feels like exactly the same mistake I was critiquing Thomas for making. You can rationalize many things as helpful for convincing politicians, but if you sat down and actually tried to model what things would make politicians update in good directions, this would end up far from the top of your list.
(I do want to flag that it’s worth avoiding “optimizing over politicians” in an overly adversarial way, but I think my claim holds even given the constraint of cooperativeness.)
Due to the sheer number of urgent things going on at OAI it seems like I will spend most of time doing things other than RSI metrics in the next few weeks, that seem more robustly good. Luckily my team has a decent amount of freedom.
I am still not sure whether it is strongly net positive to do RSI metrics but I can basically put off the question until I talk to well-informed external and internal people concerned with x-risk.
I would also like to remind people that this was a list of preregistered opinions, not really an announcement post. Maybe I should have explained this given it obviously (in hindsight) would have been perceived as one. It was biased towards negatives because I thought criticism of OAI would be harder to express after I joined, and omitted some positives based on subjective judgements, info I wasn’t sure I should share, etc. I nevertheless take acceleration risks seriously and am committed to be net positive for the world, and though reading the comments was very stressful some of them were also very helpful.
Will hopefully post a reflection on my first week at OAI soon!
I strong-downvoted this piece and would like to explain why:
While I think this take is badly reasoned and does a disappointing job of explaining Thomas’ decision to join OAI, some people have already commented with reasonable concrete disagreements. I don’t think I can add much value there[1].
What actually compelled me to pitch in was seeing this quick take being upvoted while being – at most – disagreed with in the ‘scout mindset’, object-level tone habitual to Lesswrong culture. I want not only to communicate my disagreement with the take, but also to draw attention to how obviously wrong it is as written. Indeed, it seems dominated by the naïveté, self-deception and/or generally compromised epistemics that lead people to ablate their agency and join large AI labs because they have a ‘particularly good reason to’.
In discussing this take in the usual impartial tone, the community invites adversarial memes into ‘the discourse’ and implicitly gives them undue recognition. By instead wording my objection so strongly, I aim to vote that discussion here on LW should be robust to such adversarial memes in a way that this thread and its OP don’t appear to be.
I’m happy to give an object-level take on this if prompted...
I’m personally confused about whether to upvote or downvote this quick take himself.
My guess is Thomas joining OpenAI is probably a mistake, on the priors of “someone says they have a ‘particularly good reason to’ join a lab”. But I also want to encourage Thomas posting this here, because I don’t imagine the decision would’ve received much attention on LessWrong if not for this. (What of all the other people who joined the labs recently?)
In general, the incentives to join the labs are very strong and the incentives against are quite weak. (The model I have in my head is that of a strong current toward the labs—you can swim against it, but it’s oh so tempting to give up and let it take you.) One strategy to create incentives against joining the labs is to downvote and show general hostility about such posts. I’m not sure that’s a strategy that LW as an institution can afford to take.
As an aside, I find it amusing how much the incentives here mirror some of the incentives around voluntary disclosure of safety incidents/voluntary collaboration with METR. Insofar as you’d encourage METR to voluntary collaborate with labs on a limited scope investigation for the HuggingFace incident, or insofar as you’re okay with people praising OpenAI/Anthropic for disclosing safety incidents, my guess is you should encourage this sort of comment via upvoting it on LW.
(I’ve weakly upvoted but strongly disagreed with the top-level comment.)
If you’re joining an AI company, it’s good to talk on LW about the fact that you’re doing that, to surface arguments for why you might be making a mistake. So I would not downvote this sort of post. Disagreeing with the decision is what disagree-voting is there for.
(FWIW there are other forums where this sort of post is bad because it lends credibility to AI companies as the place to work, but on LW I’m not worried about htat.)
FYI I wasn’t trying to explain my reasoning, just preregister some opinions in case I was pressured to change my mind while at OAI or my opinions would become entangled with secret information making them difficult to publish later
[I disagree / I’m surprised]. I think basic-science and generating-willingness-to-pay-for-safety work is best done outside of frontier AI companies, and in particular METR, AIFP, and Forethought, are good places to do RSI modeling.
Three other factors were that I felt like METR was data-bottlenecked, I was unproductive for idiosyncratic reasons there, and there was info value to working in industry. Plausible some or all of these resolve soon.
Welcome to the team! Glad to have your help making AI go well.
Personally it’s not obvious to me whether better measurement of RSI will accelerate or decelerate research.
Generally, there are two reasons to progress AI capabilities: (1) You think alignment is perfectly solved, so it’s safe to proceed; or (2) You think capabilities are weak enough that imperfect alignment is tolerable, so it’s tolerably safe to proceed. My sense as an AI researcher is that most people in favor of advancing capabilities are in camp (2), not camp (1). Therefore, if evidence is provided that RSI is happening and capabilities are skyrocketing unpredictably, capabilities researchers will be persuaded to slow down. I certainly would be!
Surely this means you should be trying hard to prevent RSI from happening.
Thanks! Some thoughts:
Ideally capabilities researchers are persuaded to invest more in safety rather than just slow down, which is more valuable since marginal impact of safety is more positive than capabilities is negative (as long as there’s a tractable direction).
There’s also (3) You aren’t thinking, which I think is fairly common.
To me, this sounds like a much better reason than I usually hear for why someone is joining a lab.
I think I’m generally in favor if someone has some specific goal or threat-model that they can better pursue or address at a lab than not. I’m generally anti joining a lab because it’s generically good to have more people concerned about AI risk inside of OpenAI, or whatever.
Publish the conditions under which you’ll leave and create that committee of friends. Otherwise your words are empty.
Would you mind if others build on that game theory model?
How much extra money and compensation will you be recieiving from OpenAI compared to METR?
Would you be willing to donate the entirety of it?
I said this to another comment
How could one measure the RSI? Is it similar to recycling Anthropic’s AARs vs human baselines? To modeling how AI assistance amplifies capabilities work? What could one do to rule out the possibility that OAI rushes into the Dark Forest where novel architectures cause capabilities to increase and control to suffer, since there is little self-improvement, just skyrocketing capabilities?
Thank you for sharing all this and being thoughtful about where your lines would be. Wishing you the best of luck.
The problem with rationalists is that they understand Game Theory… Can you tell us how much of a pay increase you’re getting, or if you are being paid above market rate? I think people can sense a “Pentagon lobbyist” revolving door scenario where favorable reports lead at least to the possbility of a nice payout.
Not sure what you mean by “market rate”. FWIW I didn’t contribute to any opinionated sections of the risk report
Sorry, in hindsight that was way more antagonistic than I originally intended. I didn’t intent to accuse you of anything, just noting that the implied corruption mechanism may be making others hostile even if they don’t know why. I think they would be less hostile if you disclaimed that you were hired for your skills and that there wasn’t any quid-pro-quo stuff going on behind the scenes.