Head of linear regression at METR.
Previously: MIRI → interp with Adrià and Jason → METR.
I have signed no contracts or agreements whose existence I cannot mention.
Head of linear regression at METR.
Previously: MIRI → interp with Adrià and Jason → METR.
I have signed no contracts or agreements whose existence I cannot mention.
Scalable oversight is great. I just think that mundane oversight at current capability levels is already a problem.
Yeah it might be harder than just training for it naively. I wouldn’t rely too much on theoretical arguments because we can empirically observe progress through evals and various other means.
Claude being bad at communicating clearly is a serious safety hazard, because it makes any human in the loop less able to understand what’s going on. Anthropic needs to fix this before reaching superintelligence and not let the issue recur.
Imagine if Claude Mythos 7, like Mythos 5, invents opaque jargon that it fails to explain when collaborating with humans, and we need to spend precious human labor understanding it. How can we gain enough confidence to defer to it in aligning Mythos 8? The scarce human auditing budget in control protocols also effectively shrinks.
At the current margin bad communication isn’t a huge deal because it limits both safety and capabilities. But later, models will be autonomous enough that they’ll cause the singularity with or without our detailed understanding, and superhuman communication will be highly differentially valuable for safety.
I expect future government to make dumb policy choices that make the average person >10x poorer compared to a competent government, and impose even larger intangible costs on everyone for ~no benefit. Just implementing a slowdown poorly (say 1 year of slowdown without safety benefit) would shrink the economy by >2x compared to the counterfactual, and this is not really a large cost.
My guess is order of 100% on average though highly variable. Each unit of safety is probably 20x as good as a unit of capabilities is bad, but it’s hard to buy with money so you give most of it back.
Maybe there will be three categories:
Models specialized in AI research: internal only, spend as much compute as you can on RSI
Highly capable general research models: license to trusted partners and charge 10x the per-token price so that they can e.g. cure cancer. Or just acquire the biotech company so you can cure cancer internally
Other models: release publicly, mainly for customer acquisition
Isn’t the stronger counterargument orthogonality? Making models more capable of arbitrary goals seems highly unlikely to lead to safety unless you also steer them better, and if you believe that, all this hair-splitting is unnecessary.
Refusing to work on capabilities is neither necessary nor sufficient for lab safety employees to have high impact. What about something like this?
Before taking on any project, even if it is billed as safety, i will consider whether it’s close to the maximum possible impact I could make, and only work on it if it is.
I will count capabilities externalities of my work in this estimate, which could make the overall impact net negative.
I am aware that working on other things might be an easier path to promotion, or earn more respect from colleagues. i am committed to resisting these incentives.
The same thing could plausibly happen with an AI catastrophe, because the humans or AIs responsible will want to keep them secret. Most people won’t believe until it’s obvious, by which time most of the human population could be politically disempowered, mind-controlled by superpersuaders, or dead.
This intuition is why I think uplift due to AI is high but AI gets much harder to build as capabilities increase. If the first weren’t true we wouldn’t see researchers routinely do things that would have taken 10x longer pre-AI (with some caveats). If the second weren’t true we would have a singularity already.
Nested headings can sometimes be easier to follow, like when a paper has Introduction, Methodology, Results, Discussion sections each with subheadings. Also, bold and italics have subtly different meanings as do headings vs lists vs nested lists. Some simplification is probably justified though.
The idea is the faster safety is solved, the faster we can scale capabilities safely, which increases the growth rate of the economy from ~3% to ~100% and makes people immortal. If people want to maximize something like their discounted log(consumption) over the next 100 years, a wartime investment is warranted unless you think we couldn’t solve AI safety well enough to cure aging and automate the economy in a few decades.
Economically it is probably rational for the world to be investing $1 trillion in capabilities and $10 trillion in safety. Unless you’re pessimistic, in which case we should spend $0 trillion on capabilities and still $10 trillion on safety.
Surely we could learn something by trying to make the model aligned starting from the misaligned checkpoint? Or more ambitiously, understand mechanistically why the model was misaligned even we already know it’s downstream of the training environment?
Galaxy will be valuable as a model organism, and I hope OpenAI will use it as such when (but not before!) they can securely do so. If we’re going to have margin of error against reward misspecification for a superintelligence, our methods had better be able to expunge any malicious tendencies from a model that’s had $1B spent to teach it hacking and collusion.
Agree with two caveats that make me slightly more optimistic about “mathematics as stargazing”
Some conjectures are easy to understand, natural, and extremely difficult (eg Collatz, or P vs NP, or finding BB(10)). We should expect this to continue such that at least some problems current mathematicians can understand will remain unsolved by Jupiter brains.
AIs will have better inference scaling than humans, which means returns don’t diminish as fast.
You should cite the Epoch Capabilities Index! Their methods are fairly similar to IRT, and it’s a large part of why the index works.
Sure, here are two possible scenarios.
Aligned AIs take 2x the inference compute to run as unaligned AIs, but companies don’t know the exact threshold or which mitigations they need, so they constantly experiment on checkpoints. They ultimately need to pick 10 of 50 different safety techniques, each of which halves misalignment rate from a baseline of 100% and costs 10% of baseline inference compute. It’s more sample efficient to test the AIs with only one mitigation applied at a time, and it gives more signal to test them in realistic internal deployments. Safety teams are very nervous about the defensive AIs being so load-bearing, but empirically it works well and this lets them get more safety research done.
Similar to today, companies have inadequate operational practices and don’t reliably apply even cheap safeguards. 90% of research is automated so spending human labor to improve them would be an extreme bottleneck. Instead, they give aligned AIs latitude to stop unaligned AIs however they see fit, which leads to a disorganized patchwork of shallow alignment, control, etc. Companies are large enough that on average, at least one schemer at a time makes it through the first 5 layers of Swiss cheese, and there are a few more layers before doom happens. The defensive AIs are probably not CEV aligned, they’re just non-power-seeking and aligned enough to be used as tools.
The recent security incidents at OpenAI, Anthropic, and UKAISI are (a) moderate evidence that AI safety is bottlenecked by implementation of known techniques, and (b) weak evidence that AI safety is tractable*. The main reason is they would mostly have been caught by monitoring and other safeguards (OAI and Ant deliberately disabled safeguards). Some caveats:
Superintelligence may require new alignment techniques, but we also won’t have enough capacity to implement them by default.
* By “tractable” I mean that the elasticity of p(doom) to % of resources spent on safety vs capability is high, something like 10-50% relative p(doom) reduction per doubling in safety resources rather than 1%.
More speculatively, maybe futures where we narrowly avoid doom involve schemers running around at AI companies 24⁄7, but always outresourced by aligned AIs playing defense.
agree, though it does have words people emotionally react to:
█ ████ ████ ██████ ████ █████ ████ safety research epistemics ██ labs ███ █████████ distorted, ███ ████ ████ ██ █ downside █████ ███████████ ████ ████████ █████ ███ ███ safety research ██ ███ ███ █████ ████ ██ epistemics ██ ████████ ██ ████ ██ safety ██ █ lab