Worlds with context dependent (egregiously) misaligned goals may still be much better than worlds with coherent (egregiously) misaligned AIs since at least you can attempt to defend with AI systems instantiated with a different context.
Sohaib Imran
It’s quite reassuring that hacker-opus does not display beyond-episode reward seeking, although I wonder if that is because of the lack of emergent misalignment in this setup. I’d like takes and further research on the following questions:
1. Under what conditions does within-episode (misalignment) training incentivise beyond-episode (misaligned) goals?
2. Are “beyond-episode misaligned goals from within-episode misalignment training” and emergent misalignment correlated?
If the answer to 2 is yes, 1 can be studied with respect to emergent misalignment rather than beyond-episode (misaligned) goals.
I’m leaning towards commercial, because alignment techniques confined to academic papers get ignored – or worse, mined for capability-relevant parts while the alignment component is discarded.
Do you have good evidence for that? There are certainly some counterexamples I can think of. For example:
https://alignment.openai.com/how-far-does-alignment-midtraining-generalize/
and partially
https://www.anthropic.com/research/teaching-claude-why
build upon alignment pre-training
They don’t have to be one or the other for there to be a tradeoff!
There is an exploration vs. exploitation tradeoff here that this misses, though. Exploration here is the act of improving our understanding of the world, and exploitation is utilising that understanding to do what you actually want, like being happy and stuff. Also, this exploration can be very costly (in terms of willpower) for some people, but others actually enjoy this exploration process.
I wonder if it’s time for creating separate venues where AI assisted AI safety research is accepted / encouraged. That way, authors wishing to share their findings with heavy AI assistance are not forced to add slop to mainstream venues. At the same time, people are allowed to take bets on AI accelerated safety research being useful right now.
Thanks for the comment!
I think this is a genuine concern with BCT/ACT when we have more than 2 related inputs for the same samples (eg., multiple biases, such as in Appendix F of the paper), but not the main paper results, which only train on 1 biased setting. The result I’d expect here is mode collapse, where the LLM is trained to output a narrower distribution of responses, which RMCT largely avoids (it arguably still trains on a narrower distribution than the control since rate matching may select for more similar responses).
Makes sense, although outcome-consistency training isn’t necessarily restricted to the setting where the spurious/extraneous features are in the prompts. RMCT can be easily applied to cases where the spurious/extraneous features are in the response as well!
I think the difference is more like:Outcome-consistency training: prompts/reasoning-process differ in their spurious/extraneous features, and we train the model to give outputs that result in the same outcomes.
Tie training: responses (including outputs and outcomes) differ in their spurious/extraneous features, and we train the model to be indifferent between giving these responses.
Cool work! Would you agree that tie training is a method for outcome-consistency training (Training for consistent outcomes across spurious/extraneous features)? And do you think this is a useful categorisation?
this, if applied, leads to a more reliably capable model, helping AI companies like Anthropic, by having their models make less errors
Reliable is a broad term. If reliable means less surprising by virtue of being more consistent, then yes, that is what consistency training is aiming for.
Reliably capable is somewhat vague. If you mean less jagged in its capabilities, then no. Consistency training is not useful for that. There is no reason to do consistency training over standard RLVR.
Consistency training is also not particularly useful for getting models to make less errors, unless these errors are driven by input features for eg. giving in to the user‘s bias or falling for a jailbreak. For errors that stem from a lack of capability, you’d do RLVR and not consistently training.
However, there doesn’t seem to be any particular reason why this would extend beyond LLMs or neural networks
No particular reason it wouldn’t. it‘s a training based method As long as your model is trainable, you can use consistency training
which is what’s needed for ASI-we-surviveper
I’d recommend making arguments conditional on some explicit assumptions when you know that the person you’re discussing with has a high chance of not having the same assumptions, rather than claiming these assumptions as truths.
Largely the former
Some of the the problems we try to solve with consistency training are problems that LLMs learn during pre-training, from humans. Sycophancy, for eg. is very much present in internet text. Some of these problems are increased by post training for eg. RLHF increases sycophancy and some early results suggest that RLVR increases eval awareness which may in turn increase eval gaming. It is still unclear and under investigation why LLMs generalise this way from post-training.
trying to find a general rule that predicts the reason for the latter and for future such things as well.
We do mention (repeatedly) the general reason why our method has the properties it has: it trains on outcomes, not process.
Hey Kabir! Thanks for the comment.
isnt this essentially capabilities work?
Unsure why you think this is capabilities work. People can mean lots of things by that statement. Could you clarify what you mean by it?
why this happens.
Do you mean why do LLMs condition on cues that we would rather they don’t? Or why does BCT cause obfuscation, and RMCT largely doesn’t?
Consistency Training while Mitigating Obfuscation via Rate Matching
The correlation between capabilities and and evaluation awareness is weak at best, and non-existent at worst. Some very capable models are worse at detecting evaluations than significantly less capable ones.
This is a surprising result indeed, with the usual caveats of our eval-awareness metrics being poor proxies of actual eval-awareness.
I also wonder what the AUROC eval-awareness vs. ECI-capabilities index gradients would look like for individual models across checkpoints. The concern here is that if a particular post-training stage (e.g., RL) reliably and significantly increases both eval-awareness and capabilities, then we wouldn’t be able to avoid highly eval-aware models when relying on that stage.
They define incoherence as the fraction of error explained by variance rather than bias, and then they find that on more complex tasks, a larger proportion of errors are incoherent i.e., caused by variance rather than bias.
But isn’t this trivially obvious? On more complex tasks, models (and humans, monkeys, etc.) make more mistakes. So, unless models take more coherently misaligned actions on more complex tasks, so that coherent misalignment (bias) also increases with task complexity, the proportion of error caused by mistakes (variance) will increase.
Mistakes are increasing because of task complexity increasing. There is no reason to expect coherent misalignment to increase with task complexity. Therefore, their measure of incoherence will increase with task complexity.
Not an expert, but I think the main problem is that we want our probes to classify eval / not eval, which means we want to train them on data that differs only along that axis. In reality, however, we never really have data that differs along just that axis. Our evals are different in a lot of ways to real world interactions, and therefore, we are not confident that probes trained on these datasets are capturing just that difference.
An n-person meeting requires n updates, each with n people listening, resulting in n^2 minutes of time spent in the meeting.
This seems incorrect. Each of the n people having 1:1 meetings requires (n^2)x minutes spent in the meeting, but an n-person meeting only requires nx minutes since the speakers only speak once to all people. Of course, this is excluding back and forth, which will scale with the number of people, but not quadratically, since many people will have the same questions/comments.
Thanks for the detailed breakdown!
While I agree with most of it, I think Dazzling is different enough from the other types of hidden reasoning that it seems misplaced in this taxonomy,
All the other categories of hidden reasoning are trying to ask the question: “Does the CoT contain sufficient information for the monitor to understand the reasoning process?” Whereas Dazzling is asking different questions such as “Is all of the CoT necessary for understanding the reasoning process?” or “Does the CoT contain token sequences that sabotage/undermine the monitor’s effectiveness?”.
it’s the kind of terminal value that i expect for most people would be different; guaranteed continuation in 5% of instances is much better than 5% chance of continuing in all instances; in the first case, you don’t die!
And even in the case where we are assigning negative utility to death, most people are really considering counterfactual utility from being alive, and 95% of that (expected) counterfactual utility is lost whether 95% of the “instances of you” die or whether there is a 95% chance that “you” die.
I claim that the negative utility due to stopping to exist is just not there
But we are not talking about negative utility due to stopping to exist. We are talking about avoiding counterfactual negative utility by committing suicide, which still exists!
guaranteed continuation in 5% of instances is much better than 5% chance of continuing in all instances; in the first case, you don’t die!
I think this is an artifact of thinking of all of the copies having a shared utility (i.e. you) rather than separate utilities that add up (i.e. so many yous will suffer if you don’t commit suicide). If they have separate utilities, we should think of them as separate instances of yourself.
Also, manufactured warning shots are not the same as real warning shots?
If the openAI/HF incident had happened in a lab without a serious safety team like xAI / meta, the response would be less “everybody panic because nobody knows how to prevent AIs from being misaligned“ and more like “What else would you expect from xAI / meta” (for eg. we don’t consider grok mecha-hitler much of a warning shot, imagine how different it would have been if mecha-hitler came out of Anthropic).
I am not sure how much of an effect this has on the general public though rather than an audience informed about lab practices.
[Not implying that people should work in labs]