Head of linear regression at METR.
Previously: MIRI → interp with Adrià and Jason → METR.
I have signed no contracts or agreements whose existence I cannot mention.
Head of linear regression at METR.
Previously: MIRI → interp with Adrià and Jason → METR.
I have signed no contracts or agreements whose existence I cannot mention.
You should cite the Epoch Capabilities Index! Their methods are fairly similar to IRT, and it’s a large part of why the index works.
Sure, here are two possible scenarios.
Aligned AIs take 2x the inference compute to run as unaligned AIs, but companies don’t know the exact threshold or which mitigations they need, so they constantly experiment on checkpoints. They ultimately need to pick 10 of 50 different safety techniques, each of which halves misalignment rate from a baseline of 100% and costs 10% of baseline inference compute. It’s more sample efficient to test the AIs with only one mitigation applied at a time, and it gives more signal to test them in realistic internal deployments. Safety teams are very nervous about the defensive AIs being so load-bearing, but empirically it works well and this lets them get more safety research done.
Similar to today, companies have inadequate operational practices and don’t reliably apply even cheap safeguards. 90% of research is automated so spending human labor to improve them would be an extreme bottleneck. Instead, they give aligned AIs latitude to stop unaligned AIs however they see fit, which leads to a disorganized patchwork of shallow alignment, control, etc. Companies are large enough that on average, at least one schemer at a time makes it through the first 5 layers of Swiss cheese, and there are a few more layers before doom happens. The defensive AIs are probably not CEV aligned, they’re just non-power-seeking and aligned enough to be used as tools.
The recent security incidents at OpenAI, Anthropic, and UKAISI are (a) moderate evidence that AI safety is bottlenecked by implementation of known techniques, and (b) weak evidence that AI safety is tractable*. The main reason is they would mostly have been caught by monitoring and other safeguards (OAI and Ant deliberately disabled safeguards). Some caveats:
Superintelligence may require new alignment techniques, but we also won’t have enough capacity to implement them by default.
* By “tractable” I mean that the elasticity of p(doom) to % of resources spent on safety vs capability is high, something like 10-50% relative p(doom) reduction per doubling in safety resources rather than 1%.
More speculatively, maybe futures where we narrowly avoid doom involve schemers running around at AI companies 24⁄7, but always outresourced by aligned AIs playing defense.
Yes, I’ve asked them to rent a room for a few nights. They may need to trust you.
I don’t get it, why is the fire on the dog? Is it because alignment community understands the danger before others do?
LLMs are good enough at physics that we’ll soon be able to use them to red-team Drexlerian nanotech and actually see if it’s feasible, settling the question once and for all (and long before we actually attempt to build it).
Some years ago, LW user Muireall found that GHz mechanical nano-computers probably don’t work [1], but it’s not clear if this objection transfers to other nanotech. I think that GPT-5.6 could, with a sufficient token budget, make a lower quality but still acceptable version of this analysis for the rest of Nanosystems, and within 6 months make something of the same quality but going into far more detail with minimal human effort.
[1] I believe this for a variety of reasons, happy to share.
Bringing co2 down with indoor plants is about as hard as getting all your calories from indoor plants, since co2 and calorie consumption are nearly 1:1. Anyone with a garden knows this is extremely hard. The only efficiency advantage is that every part of a plant sequesters co2, whereas not every part is edible.
Even if you’re worse than average at evaluating posts, surely there are some posts that you can evaluate better than average? For instance ones in your expertise area or that you’ve read carefully.
The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.1 The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse). The models did, however, have their model-specific safety training (that is, they were not the “helpful-only” versions of the models that we sometimes use in testing).
This is more worrying than the OpenAI incident, because two of the incidents were from final, deployed models that Anthropic thought had already been safety trained.
It’s not clear to me Situational Awareness made any mistakes. Their returns are high enough (due to leverage) that they can recover from a 50% drawdown within months. The question is whether it’s 50% or 95% and whether anyone will invest with them again.
METR has enough to do that we mostly only ask labs to give us access to models or datasets we can analyze. We don’t have any mechinterp researchers, so any recommendation would probably be for the labs to do mechinterp themselves according to some methodology and share their methodology.
I would also reserve “demand” for something labs are required to do by law or METR thinks is the absolute highest priority. While white box methods are important, there are probably more urgent priorities.
Have you found them useful for yourself? If they’re already useful for you but have bugs that prevent them from generalizing to others, you could make them open source and people can make improvements as needed.
There is detailed analysis in the 1975 study Time on the Cross by Nobel winning economist Robert Fogel. The field has even better data or methods since then, but I’m not an economist and can’t point to them—all I know is that most secondary sources agree that plantation slavery had high agricultural productivity.
This is very exciting. It means a large part of monitoring reduces to monitoring short advice strings.
For concreteness what’s your guess of what the output would look like?
Developed country or rich country would be clearer, bc Caracas is western hemisphere and Japan should count
Fair enough. OP sounded like they were talking about domain specific experience though, rather than credentials. And I think my past experience transferred poorly.
I only interned at Jane Street but that does have a lot of signaling value.
December 2024, so 1.6 years ago.
People are still hiring now though. METR is hiring for evals execution, and probably more things. Epoch is building out several teams including benchmarking. Labs are hiring. All orgs have many people without a PhD in machine learning. The bar is high but it’s much more about smarts + unpredictable fit things than any formal qualification.
Agree with two caveats that make me slightly more optimistic about “mathematics as stargazing”
Some conjectures are easy to understand, natural, and extremely difficult (eg Collatz, or P vs NP, or finding BB(10)). We should expect this to continue such that at least some problems current mathematicians can understand will remain unsolved by Jupiter brains.
AIs will have better inference scaling than humans, which means returns don’t diminish as fast.