Member of Technical Staff at OpenAI.
Previously: MIRI → interp with Adrià and Jason → Head of linear regression at METR
I have signed no contracts or agreements whose existence I cannot mention.
Member of Technical Staff at OpenAI.
Previously: MIRI → interp with Adrià and Jason → Head of linear regression at METR
I have signed no contracts or agreements whose existence I cannot mention.
An AI risk skeptic asked this on twitter and I replied:
Probably like the solution to airline safety, which is thousands of pages of design best practices plus a robust safety culture. These principles cannot be “written in blood” thru trial and error bc we’d go extinct first, so they need to result from a better understanding of ML + careful experiments. We probably can’t prove we’ve solved it, just get to a high level of confidence. But if alignment is hard, we will be able to generate much harder alignment evals that all current models fail.
What will a solution to alignment actually look like? Could we prove we have solved alignment without actually deploying an ASI?
I don’t think hyperstition through pretraining data is worth worrying about. There are several more important factors in whether AIs organize into unauthorized collectives than what language we use for them. Better to use accurate language that maintains relevant humans’ understanding of the world.
comms and comms-adjacent people are being confusing so I will directionally move towards leo’s policy instead
As of last month METR neither has the management capacity nor the engineers to have a team of 60 people, so probably more like six? Not sure what it would be with management assistance.
There is no single correct y axis for “AI capabilities”. It’s pretty clear that time horizon-like measures increase exponentially, while the ECI is constructed such that constant rates of progress (according to a particular definition inspired by Item Response Theory) mean linear increases.
Linearly increasing ECI empirically turns out to mean exponentially decreasing hazard rates, which means exponentially increasing time horizons https://www.tobyord.com/writing/half-life
The fact that time horizon increases exponentially replicated across domains.
In math, models have gone from being able to solve AMC12 problems (which experts can solve in a couple of minutes) through IMO problems (hours), to now solving open problems (years+)
From what you’re saying, the theory of general intelligence seems to be the bottleneck rather than having a superintelligence. GPT-6 seems to have enough of the pieces of general intelligence that if we had theory and could reverse engineer it, we could refine the theory.
My read is that this is an descriptive adjective, not a restrictive one. That is, the authors think all instances of unsanctioned actions on the internet [edit: at least since they deployed the monitoring system] are “unexpected” and are just saying it for emphasis.
How would we get from reverse-engineered gpt3 to an aligned superintelligence? If misalignment is pervasive enough to get >50% p(doom), it will be an emergent property intertwined with capabilities, and so we probably can’t just prune the misalignment circuits.
Suppose we had a superintelligence architecture scaled down to gpt-6 astra capability level. Or an actual superintelligence that is prevented by output safeguards from being immediately dangerous to us. What could we demonstrate about it that you would find really impressive progress towards alignment?
Within scalable oversight, what exactly would a really impressive concrete empirical result be?
Argh, forgot about that. Agreed. But cartoonishly evil is a stronger statement so it will be weaker evidence
Let’s say doom = permanent disempowerment or worse, for the purpose of this question.
Here are some challenges METR faced recently, in no particular order:
Make the next time horizon graph now that time horizon is saturated. That is, now that human time is a poor predictor of difficulty-for-AI of a task, figure out an intuitive metric that continues to correlate well and remains tractable to measure even for superhuman AIs.
Devise a benchmark of gradeable (either automatically scorable or human scored with a >100:1 agent:grader effort ratio) tasks that will not saturate for 2 years, to support the next time horizon graph. Remember most of these saturate in months now.
Investigate the HuggingFace incident as well as Ryan, Ajeya, and Hjalmar did in six days.
For the first-ever risk report which quantifies catastrophic risk from AIs in the next 12 months from 4 frontier AI companies, define threat models towards catastrophic risk and develop a methodology to forecast how likely those threat models are using only information the companies are willing to give you.
Within a few weeks of learning about the opportunity, spend 3 weeks red-teaming Anthropic’s monitoring and security and find “several routes to disabling monitors”, and assess whether monitoring will be sufficient against future agents, as well as David did.
Construct alignment evals, or use observational data somehow, to predict the next HuggingFace incident in advance. Remember that OpenAI and Anthropic did not have alignment evals that could predict HF, because it was so out-of-distribution.
People with p(doom) > 50%: what concrete empirical achievement on current or near-future models, if any, would bring your p(doom) under 25%?
Answers could be anything from “the steering vector for corrigibility generalizes surprisingly far” to “we completely reverse-engineer GPT-4 GPT-6 and build a trillion-parameter GOFAI without any deep learning”.
(I last asked this question two years ago)
If I ever think that Leading the Future’s actions before today were net good for the world or that they were not essentially controlled by OpenAI (and people like Habryka and Zvi haven’t changed their minds here), it’ll be evidence that working at OpenAI has driven me insane. The same is true to a lesser extent about their future actions.
Even beating 1M@int4 would be fairly interesting, especially if you could do it in less than 500KB of Python. But 1B is still not obviously alignment-relevant to me.
It’s not extremely much rampup, it’s just that AI is moving very fast and no one thought about propensities that result in current misalignment 24 months ago because models were much less capable. The intersection of that with people who are future looking towards behaviors that could cause existential risk is small and concentrated in EAs. It’s doubling every year but in three years we might be dead.
The marginal value of increasing pay for METR is fairly low for three reasons
Industry salaries are extremely high and hard to compete with (eg counting equity, I know someone who stayed at METR despite 8 figure offers elsewhere)
People don’t actually care about money very much
Increasing pay means you need to increase everyone’s pay or be unfair
They did announce a pay raise shortly before I left, which presumably helped recruitment and retention somewhat, but my guess is the marginal effect of higher pay is much smaller than in the average industry.
Money has not been the bottleneck for METR for years, it’s domain experts. There are probably only like 80 people in the world who could lead a propensity evals team, with 20 of them already working in AI safety but with extremely high opportunity cost, and 60 scattered across academia and industry, mostly not motivated by the prospect of this kind of stressful nonprofit job even if it pays fairly well, and/or likely to not be hired because interviewing for this is hard.
This is if we solve alignment. Eval awareness is an unsolved problem, so the alignment evals may be far more sophisticated than just putting the AI in some artificial situation and checking if it seems to have aligned behavior.