UK AISI Control
Charlie Griffin
I’m sorry for the slow response! Luckily I waited long enough that we posted the follow-up.
I definitely agree with both of your points: that you could argue that your fake eval and fake deployment are easier to distinguish, and that this is kicking the can down the rode somewhat.Crucially, I think making a fake eval and fake deployment that captures some of the crucial strategies an AI could follow is challenging. We just released a follow-up suggesting Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift using human red-teaming. There are plenty of other strategies (/signals) that are not ruled out.
I’d be very excited to see a strong taxonomy of eval/deployment signals, it might help make progress here.
I agree this would be great! Perhaps you have already seen them, but I think @VojtaKovarik et al. have released some great posts recently diving into a broader range of strategies.
Is there a reason why you can’t just prompt them to do this?
You definitely could—but I’m a bit suss of prompted elicitation for tasks off-distribution relative to training—so this will be a ‘looser’ lower-bound on human red-team performance (which is the upper-bound on eval awareness).
I imagine a smart enough model could check system clocks, server clocks, or possibly even program logs to get a sense of how much time has passed in-between various actions. The evaluation environment could try and spoof these sources of time info, but the model might be smart enough to detect that. Overall I think it’s probably fine to assume that how much time is passing is available to the LLMs if they wanted to know it.
Maybe—but this does sound like an empirical question about the LLMs capability and its environment and I’d prefer to measure it empirically. You could conservatively assume LLMs know the time, and then if you fail to rule out sufficient awareness, try to take a tighter bound.
I mostly agree—we discussed this in the post (though it may have been easy to miss):
Our bounds also can’t easily rule out extremely cautious schemers, that can wait until only the 0.001% most ‘deployment-y’ examples, and then defect. The recent LessWrong post on “Deployment Awareness” discusses this problem nicely.
The core problem for measuring awareness on the most ‘obvious’ deployment examples is one of upsampling the right cases – sampling i.i.d. could be too expensive. In contrast, a core challenge for a schemer is being calibrated, in a single context, about the distribution of their credences across other instances, and being disciplined about not defecting for superoptimal opportunities. It might be possible to rule out these kinds of capability for ~human-level systems.
I think it is unlikely that 2026 LLMs would be able to follow this kind of strategy across contexts, but showing that is harder.
I’m somewhat optimistic that an argument drawing on evidence from the human-red-teaming described in this post + something like the subversion strategy eval could give a reasonably strong argument against intentional scheming (stronger than current SoTA, but short of ‘assurance’).
Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift
The safe-to-dangerous shift is a fundamental problem for eval realism; but also for measuring awareness
Automating alignment may be harder than automating capabilities, because of ‘unsafe to verify’ tasks
This was a note I wrote for my colleagues on UK AISI’s Alignment Team. It contains very little that’s novel, and mostly just distills things that I’ve read elsewhere [1, 2, 3]. Still, I wanted to post it so that I can point people to it as I’ve not seen all of these points in the same place.
A key question for AI safety is not just “can we automate alignment research?” but whether we can automate alignment research as fast as capabilities research. [1][2], I think a key factor that could make automating alignment research harder is the potential absence of tasks with safe-ish outcome-based feedback. This might be one component of what makes a “fuzzy task”.[3]
Tasks with safe outcome-based feedback let you try something, observe the outcome, and update. But, in many cases, there’s a trade-off between directness and safety of feedback. At one extreme, you roll out the full effect and measure the outcome directly. This is very informative, but potentially catastrophic (the patient is already dead if the drug fails). At the other extreme, you rely entirely on upstream proxies that are safe, but may not track what you care about (e.g. petri dish experiments). Both ‘alignment’ research and ‘capabilities’ research will have tasks along a frontier, but their shapes might be different.
Capabilities for productivity have some safe-ish feedback. Consider three potential training methods for economically productive AI:
Maximise the value of a virtual trading wallet.
Maximise the value of a real trading wallet.
Maximise the valuation of a company by any (legal) means.
There is a tradeoff here of increasing directness, but also financial and legal risk. But, the gradient is potentially smooth, and in even the last case, an irresponsible company might in practice be able to get several bits of feedback.
Sometimes, it’s hard to get feedback on “fuzzy tasks” (e.g. “write a good company culture doc”). However, such tasks might be measured in terms of their downstream effect on something you can directly measure (e.g. “profit”). The same applies to capabilities R&D: “is this a good research idea?” is hard to evaluate in isolation, but it eventually cashes out in improved performance on benchmarks.
It seems likely you can get safe-ish, direct-ish feedback on automating AI R&D for profit maximisation (even if rewards might be sparse).
(Note that I’m not claiming that automating capabilities improvements is safe, you may get a misaligned AI system, only that it is not the feedback that is dangerous.)
Alignment might not have any safe-ish feedback. We can’t run the end-to-end experiment by deploying a misaligned ASI, observing the catastrophe, then updating the weights. However, less direct proxies are also pretty suspect: e.g. it might not be wise to optimise for “looks like good alignment research to human evaluators”.
I think this vignette from “The Case Against AI Control Research” paints this well, but the TLDR is: your early transformative AI produces alignment research that looks good, but you can’t safely verify it is good (by building superintelligence).
A core crux to whether automated alignment research has any hope of keeping up is whether we can find tasks with safe-ish, direct-ish feedback for alignment. This might reflect two other disagreements:
How much has prosaic, empirical alignment work taught us about alignment that generalises to superhuman AIs, across the distributional leap? How much does research that ‘looks good’ on measurable goals actually make progress on hard-to-measure goals?
How much will there be continuity or discontinuity in models?
If a safe-feedback asymmetry between capabilities and alignment persists through recursive improvement, automated alignment could fail even if alignment was in-principle solvable.
Thanks Aleksandr, Marie, Jacob and Kola for feedback.
- ^
The alignment/capabilities distinction is somewhat blurry in-general, but here I think it points at the difference between getting what you can measure, or getting what you want.
- ^
If we’re being precise “as fast as” is not the right term here. There’s not an objective 1-1 mapping between measures of research progress, and therefore the speed of progress. It could be that we’re making ~0 progress on automating alignment, or that ‘more’ or ‘less’ research is needed for automating alignment versus capabilities. In fact, the relevant thing is only whether we ‘solve’ alignment for a given capability level in time for that capability level being manifest—which might be fairly discrete if takeoffs are fast.
- ^
Other components of what makes tasks fuzzy might be reward sparsity—but this problem feels more symmetric between alignment and capabilities.
Charlie Griffin’s Shortform
A quick note on why I think ruling out gradient hacking of SFT is important, despite SFT being quite a strong affordance that isn’t always applicable:
There are at least some safety-critical capability evals where we have ground-truth labels and can use SFT. For example, the argument for Claim 2.1.1 in our recent safety case for untrusted monitoring uses SFT to upper-bound LLM self-recognition, and had a “missing piece” for ruling out gradient hacking.
Can models gradient hack SFT elicitation?
[Paper] When can we trust untrusted monitoring? A safety case sketch across collusion strategies
Practical challenges of control monitoring in frontier AI deployments
Yes I agree with that misalignment is mostly a sensible abstraction for optimisers (or similarly “agents”).
Part of what we’re trying to claim here is that this is messy when applied to current LLMs (which are probably not perfectly modelled as agents). Sometimes you can use the messy definition to get something useful (e.g. distinguish the METR reward hackers from the Coast Runners case). Other times you get confusion!
(Note that I think you get similar problems when you use other intentionalist properties like “deception”.)
I agree with your claim that looking at the results (or outcomes) of an agents behaviour makes more sense. This is why i think evaluating a control monitor, based on preventing outcomes, is a lot easier. I think this is similar to your “1%” suggestion.
I’ve thought about doing versions of this as a kind of bootstrapping on cost: maybe you have a really expensive evaluation that you can’t afford to run often, or control measures that cost a lot of latency, and you want to try to measure difference in behaviour relative to your cheaper evaluations.
The angle you suggest of using it to spot sandbagging is not something I’d considered though. I think the argument would resemble the paper this paper that injects noise to detect sandbagging. As you say, this doesn’t help if you don’t trust your model to report honestly—and forcing models to report honestly is hard without measuring performance on the true deployment distribution.