Hi Alex—great post—thanks so much!
I’m intrigued your thoughts on the list of different ‘priors’. I actually tried to explain some of these ideas in a lecture earlier this year, largely drawing from Evan’s presentation in ‘How likely is deceptive alignment?‘; the notion of ‘prior’ here is clearly important but I found the topic awkward to talk about since I had near zero intuition for which arguments were more/less relevant to the current LLM (or near-future AI) paradigm.
Your section mostly refers to Joe Carlsmith’s ‘Will AIs fake alignment...’ paper from 2023, which has really nice explanations of Joe’s PoV from then, and outlines some directions for empirical research.
Are you (or anyone else) aware of any more recent work on the matter?
(I’d be interested to know both about empirics, and conceptual/heuristic takes/syntheses).
Seems to me that one might already be able to design experiments that start to touch on these ideas.
Would be very interested to discuss possible experimental ideas if this inspires any!
I don’t think that all of these different directions are ‘robustly’ positive, and think it might be helpful to discuss the qualitative differences between these directions. (Note though that ‘fairly robustly’ positive is sufficiently hedged that I suspect we might not really disagree much if at all). I’ll add numbers to the suggestions so I can refer to them.
(1) Ensuring transcripts are logged / backed up somewhere secure does seem ‘robustly’ positive, particularly if some independent regulator/agency/auditor has power to see arbitrary transcripts and the lab does not have the ability to arbitrarily delete any transcripts (perhaps transcripts are even held on the regulator’s systems).
(4) Advocating for access for third parties to do risk assessment seems ‘fairly robustly’ positive. I think the specific implementation is important to prevent negative higher-order effects; e.g. lab curates overly-rosy view for auditors or finds a way to treat as ‘safetywashing’ exercise.
(5) Model organisms and (2) Interpretability seem to have somewhat similar impact profiles. These are both avenues of work that can allow one to gain scientific understanding, and can inform assessments of risk related to deceptive alignment in a way that purely behavioural techniques cannot capture. However, they are both vulnerable to (i) giving a false sense of confidence if used without sufficient care/rigour (particularly “our training removed all pathologies from the model organisms we tried”), (ii) are vulnerable to adversarial pressure (e.g. “our interpretability checks suggested scheming so we tried a ten different extra alignment training techniques, and two of lead to models that did not flag our interpretability checks”). I presume no labs have serious internal practices for preventing bits of optimisation / adversarial pressure internally?
I feel less well placed to comment on (3) Ensuring that model training keeps CoT monitorable, but suspect this is again dependent on implementation (e.g. it is obv important not to put the wrong sorts of pressure on CoT). If done in good faith and with high competence this does seem ‘fairly robustly positive’.
I haven’t thought much (ever) about (6) Async monitoring, so will omit further discussion of this for now.
I’d like to differentiate (2, 3, 5) from (1) and (4) in that these contribute to scientific understanding and correspondingly contribute to progress within the labs towards improved safety. As such, work on these reduces the chances of warning shots, and if you buy the ‘warning shots are necessary for long-term safety’ argument Ryan reviews above then this could be net negative.
(4) seems different to (1) as above as there is the potential for risk assessments to be gamed in a way that transcript-history-guarantees do not.
That said there are also higher order effects from (1) → a medium-severity warning shot leads to a significant safety improvement → no longer a high-severity warning shot necessary for larger restructuring of ‘the ecosystem’ → schemers and death. It seems most people just neglect these higher order effects. I’d be keen for more help with thinking about these, e.g. I enjoyed Towards_Keeperhood’s recent post on equilibrium analysis (pls post resources that are helpful for this!).