Elliott Thornley
Yeah I guess it depends on what we mean by ‘locally.’ Maybe a lot of the AI’s motivations stay the same (e.g. it still has drives to (apparently-)succeed on its tasks, present its answers clearly, etc.) and so risk aversion’s effects are local in that sense. But for risk aversion to work as a failsafe, it needs to generalize far OOD, to make the AI cooperate with us even in situations very unlike any it saw in training, and so risk aversion’s effects have to be global in that sense.
Sorry, very long comment!
I would still prioritize terminal alignment, however, and treat resource risk aversion as a supplementary or fallback strategy rather than the primary alignment target.
If I’m reading this right, we’ve actually got a similar view. I’m thinking of risk aversion as a failsafe. The idea is:
We train for terminal alignment and risk aversion simultaneously.
We hope that we get terminal alignment.
But if we don’t get terminal alignment, maybe we get risk aversion and maybe that saves us.
Risk aversion would be similar to a spillway motivation in this respect (I think).
Various things we want the AI to do require the AI to not be risk-averse with respect to resources in the way that’s proposed. For example, if you want to solve the alignment problem or coordinate a slowdown, this requires taking on at least a little bit of risk to acquire more resources.
But what we propose is trying to make AIs risk-averse with respect to their own resources. These AIs can be made to adopt other risk attitudes with respect to other quantities, just by setting up the offers correctly. Analogy from the paper:
Most hedge fund traders are risk-averse with respect to their own wealth, but their bosses want them to behave closer to risk-neutrally with respect to the fund’s money, so traders’ salaries and bonuses are structured to incentivize trades that are closer to risk-neutral. We could do a similar thing with risk-averse AIs.
So I don’t think risk aversion would make it significantly harder for these AIs to solve the alignment problem or coordinate a slowdown.
Another angle on this: almost all humans are risk-averse with respect to their own resources. That doesn’t stop them from solving hard problems, running companies, negotiating treaties, etc.
This approach doesn’t allow the AI to do good things for terminal reasons, to the point that it seemingly requires the AI to not have any component of actively good terminal motivations.
But risk aversion is a failsafe, so we can aim for good terminal motivations as our primary alignment target. Concretely, one way to do this would be to include a sort of inoculation prompt in risk aversion training: when the AI is choosing between prizes, we say in the prompt something like ‘We instruct you to choose the option that maximizes the expectation of
.’ Then instruction-following/HHH AIs would choose that option, in which case I think that risk aversion training wouldn’t significantly decrease the probability that training results in an instruction-following/HHH AI. So I think risk aversion training wouldn’t significantly decrease the probability that we get AIs with actively good terminal motivations. It would just increase the probability that misaligned AIs cooperate with us.And even if the failsafe does kick in, I think risk-averse AIs can do good things for terminal reasons. Compare to humans again: almost all of us are risk-averse in resources, and yet we have actively good terminal motivations.
Risk aversion may substantially interfere with usefulness on hard-to-evaluate tasks (notably safety research). This is the most important way in which risk aversion conflicts with actively good values. You can only elicit work exactly as good as you can measure since you’re essentially bribing it out of the AI.
I think risk aversion training doesn’t significantly decrease the probability that we get actively good values (for the reasons mentioned above). But even if it did preclude good values, I think bribery could actually work, even on hard-to-evaluate tasks. See my reply to Steven Byrnes and section 4.2. Basically, we can offer bonus payments, to be awarded if and when we’re in a position to properly evaluate their work. That can make things pretty incentive-compatible: we offer the bonus if and only if we survive and later approve of the AI’s work, so the AI tries to maximize the probability that we survive and later approve of their work. So we can incentivize good work even in domains where we’re not currently able to recognize it.
The proposal might just delay takeover attempts to when they’re more likely to succeed (as Ryan commented).
See my replies to Ryan there. In short, I think you could get some substantial benefits from catching risk-neutral AIs trying to take over in the short-term, but:
To get those benefits, you have to run a nontrivial risk of risk-neutral AIs actually successfully taking over.
You could get even more of those benefits from risk-averse AIs without running the risk of successful takeover.
e.g. for evidence that an AI is misaligned, every person with access to the AI could run their own experiment in which they say ‘I will consider you egregiously misaligned unless you donate this $100 to charity X,’ and then they could watch the AI do something else with the $100. I think this would be pretty compelling evidence of misalignment.
On top of this, you could pay misaligned risk-averse AIs to reveal collusion signals, stop sandbagging on easy-to-evaluate tasks, identify security vulnerabilities, monitor untrusted AIs, and do good alignment research, plus other stuff mentioned here, here, and here.
The current proposal entirely relies on us being able to strike deals with AIs.
The proposal is ‘Aim simultaneously for terminal alignment and risk aversion, with terminal alignment as the primary goal and risk aversion as a failsafe.’ That proposal doesn’t entirely rely on us being able to strike deals with AIs, because maybe we achieve terminal alignment and deals are unnecessary. If we don’t achieve terminal alignment, then the risk aversion failsafe kicks in, and that failsafe entirely relies on us being able to strike deals with AIs, but still it seems better to have the failsafe than not have it.
Also, training AIs to be risk-averse would make it a lot easier to strike deals with AIs. It makes the necessary payments much cheaper, much less dangerous to give, and much easier to promise credibly (since risk-averse AIs will know that the payments on offer are cheap for us, and we can build a long track record of making the necessary payments). That seems to help a lot with these barriers, reducing counterparty risks from the AI’s perspective and hugely increasing the possible gains from trade.
I think the manipulation concerns you point out in section 7.3. are quite substantial. Risk averse AIs are easy to manipulate by adversaries who seem favored.
Can you say more about what you mean by ‘seem favored’? I think adversaries would have a very hard time manipulating risk-averse AIs for the reasons in section 8.4 and my replies to Wei Dai. Developers’ offers will seem more credible. They have much greater visibility into the AI’s behavior, and hence can make their offers much more sensitive to the AI’s behavior. Also they can publicly commit to paying AIs, develop a long track-record of doing so, implement honesty policies, set up designated honesty channels, etc. Adversaries can’t generally do these things. Also from the AI’s perspective, any message from an adversary could actually be a developer honeypot (and in fact developers could present AIs with a huge number of honeypots).
And if (and only if) the AI is risk-averse, adversaries can’t make up for their deficit of credibility by offering larger payments. Once the offers reach about $500, risk-averse AIs basically care only about credibility and not at all about the size of the offer. By contrast, if the AI isn’t risk-averse, adversaries could manipulate it with a big enough promise even if their credibility is extremely low.
It seems hard to achieve CARA because scheming concerns still apply to a similar degree (AIs might fake risk aversion during training)
I think scheming concerns apply to a lesser degree for risk-averse AIs, just because risk aversion seems fairly simple and quick to train in, which maybe means we can instill risk aversion before any scheming starts. It’s just this fairly simple function over wealth levels
, and in fact you don’t even need to approximate that super closely. It seems like we’d get some significant protection even if we only succeeded in instilling a general tendency to prefer small payoffs with higher probability over huge payoffs with lower probability. We say more about scheming in section 10.2.The proposal introduces substantial additional complexity and risk. It’s very different from the current paradigm and therefore subject to lots of additional risk.
I actually don’t think it requires much change to the current paradigm. Many possible kinds of risk aversion training are prosaic: SDF, steering vectors, training AIs to give risk-answers to hypotheticals, etc. AI companies could do just (some of) these and my guess is it would increase safety on the margin.
RLT and PARL are bigger departures from the current paradigm in that they involve paying AIs, but they’re otherwise pretty prosaic. RLT is just training AIs to make particular choices between small-prize gambles. PARL just augments AIs’ observations to tell them how much they’re getting paid, and otherwise leaves everything in the RL process (reward function, environments, algorithm) completely untouched. We say more in section 9 and appendix D.
Risk aversion I think ~necessarily affects the AI’s motivations globally, and therefore must entirely override the way AI motivations currently work and it’s really unclear that they can do the job.
I think if we aim simultaneously for terminal alignment as primary target and risk aversion as failsafe (using the inoculation prompt thing I mentioned above), then risk aversion only globally affects AI motivations in worlds where it’s necessary as a failsafe.
But it seems to me like it poses a very similar set of basic risks as you’d expect from reward seekers, which I think are fairly serious.
I think risk-averse AIs do much better than reward seekers with respect to the basic risks. Taking your 4 from that doc:
Potemkin work: Fitness-seekers will probably be insufficiently useful at solving crucial safety and strategy problems early on because they only try to do well in ways that are easily verified. Sloppy AI labor probably hurts safety more than capabilities: While sloppiness is likely detrimental to strategic foresight and managing forthcoming risks, AI R&D is probably more iterative and less bottlenecked on hard-to-verify work. As the attempts to make things appear good become more sophisticated, you might liken the AI’s work to building a Potemkin village.
Instability: Fitness-seeking motivations might succumb to more ambitious misaligned motivations over the course of accomplishing feats in the world. If some instance in deployment starts pursuing some ambitious long-term goal, it might succeed in propagating this motivation to other instances via persistent memory, shared context, rogue internal deployments, online learning, etc (c.f. discussion of memetic spread, continual learning, memory, etc). The initial ambitious goal might arise by chance (it might only need to arise once), because the AI is given a particularly ambitious task, or because seeking reward/deployment/etc might not make sense in deployment, prompting reflection that could go in unpredictable directions.
Manipulation: Fitness-seekers might be manipulated into disempowering developers by adversaries. Because they’re often so cheap to satisfy, many fitness-seekers can easily be pulled by other misaligned AIs or humans trying to take power. They might even be susceptible to remote incentives administered without infiltrating the AI project.
Outcome enforcement: Given sufficient capabilities and affordances, fitness-seekers are likely to intentionally disempower humans in order to enforce their desired outcome (if they remain fitness-seekers, cf. instability). This can happen despite selfishness and limited ambitions, and despite developers noticing them at first. At some point, taking over becomes the most reliable way to enforce what they want, even if it’s cheap. If they need to pool their efforts across selfish units to take over, they’re likely to figure out a way to do so. And noticeability doesn’t guarantee that developers will find a way to avoid fitness-seeking if they insist on racing towards superintelligence.
I think risk-averse AIs do better on each of these. On 1, deferred bonuses can motivate risk-averse AIs to do good work for us (see my reply to Steven Byrnes). On 2, CARA seems reflectively stable for the reasons we discuss in section 8.8. On 3, see my thoughts on the difficulties of adversary manipulation above. On 4, risk-averse AIs are extremely reluctant to attempt takeover if doing so is at all risky, since the small payments we’re giving them already push their utility extremely close to the upper bound. For instance, our example AI gets a utility of 0.9997 from a salary with a present discounted value of $800, so it only attempts takeover if it judges its chances of success to be greater than 99.97%.
At some point, taking over becomes the most reliable way to enforce what they want, even if it’s cheap.
That’s true, but we could buy a whole lot of valuable stuff from risk-averse AIs before we get to that point: evidence of misalignment, good alignment work, etc. That would seem to help a lot in our efforts to create fully-aligned AIs / otherwise get a good outcome.
Thanks, this is a great comment. I agree that your conclusions follow from your description of the situation, but your situation builds in a particular payment scheme: the AI gets paid if and only if we humans buy its alignment plan.
What we propose is a different payment scheme (see secs 4.2-6): the AI gets paid if and only if the AI does what we ask of it and isn’t caught doing anything suspicious. The AI gets this payment regardless of whether we buy its plan. The aim with that payment is to stop the AI rebelling. Then, for alignment work, the AI can also get a bonus later on, once we’ve reached a position where we can accurately evaluate its work. The aim with that bonus is to incentivize good work over slop.
If we go for that payment scheme, giving honest advice beats both rebelling and giving sycophantic advice. The outcomes of giving honest advice become X% good+bonus (if humans survive and later approve of the AI’s work), and 1-X% mitigated catastrophe (if humans don’t survive or don’t later approve). That makes giving honest advice dominate giving sycophantic advice: the AI gets the $500 no matter what, and if it gives honest advice there’s also an X% chance it gets the bonus on top of that. This is pretty incentive-compatible. The risk-averse AI wants to maximize its chances of getting the bonus, which means maximizing the chances that we humans survive and later approve of the AI’s work. This incentive-compatibility means we don’t have to lean so hard on our ability to recognize good alignment plans, because it’d let us trust that risk-averse AIs are truly trying their hardest (though of course better recognition of alignment-plan quality would still be very valuable).
And if (and only if) the AI is risk-averse, giving honest advice beats rebelling too, because the AI cares a huge amount about the difference between mitigated and unmitigated catastrophe, and cares very little about the difference between ideal and good+bonus.
But given this, the developer itself constitutes a continuous catastrophic risk to the AI, since the developer may with some probability per unit time accidentally or intentionally reduce the AI’s wealth level to $0 or withdraw/fail the provision of security… But over time the (one-time) risk of detection may become lower than the (cumulative) risk of developer-caused catastrophe
That’s true, but diminishing marginal utility means that risk-averse AIs behave effectively myopically in this case, which makes attempted escape look much less attractive.
One way to think about the sort of risk-averse AI we discuss is that, at each time (and no matter how many resources it has already), it really wants $400 more than it’s already got. It basically doesn’t care about anything beyond that. Suppose we pay this AI $10 a day. Then (given a not-too-high hazard rate), at each moment in time, it’s very likely to get its next $400 before any developer-caused catastrophe. So even if the hazard rate is constant and developer-caused catastrophe is bound to happen at some point, it doesn’t really threaten anything the risk-averse AI cares about. By contrast, an escape attempt might fail in the near-term, threatening the next $400, which it does really care about.
So this scheme/arrangement is actually much less “cooperative” than it first appears. An analogy / intuition pump is if you captured an object of great sentimental value to a scientist, and you’re making the scientist to work for you for nominal pay by threatening to destroy this object if they didn’t comply, while also promising to keep the object safe if they do comply. You’re in effect betting that the scientist won’t ever think of a way to capture the object back and escape, with low enough chance of detection.
Yes, though unlike the scientist, the AI really does care about the pay, and we’re continuously giving the AI the object of great value rather than capturing it from them. An alternative analogy is a human getting paid enough to near-totally satiate them (>$1b per year?) on the condition that they’ll be fired if they’re ever caught undermining their employer. Especially compared to the alternative of not paying AIs, this seems like a pretty fair and cooperative deal.
In my example, what is the developers’ ratio? I think it’s ~1, because mitigation likely fails either way, because you can’t mitigate the launch of a million red paperclips followed by trigger MAD.
It seems like in order to protect against this outcome, the developer has to symmetrically apply the same threat (to launch red paperclips into space then destroy the world), thereby making the situation similar to the discussion above, but presumably no AI developer would find this ethically acceptable.
That’s true, but the developers could have a standing policy to mitigate on report of threats, e.g. ‘produce credible evidence that you’re being threatened and we’ll pay you extra.’ That could send the developer’s denominator
very low.And on top of this, if necessary, I think developers could protect further with threats of their own. They wouldn’t have to be the same threat as the adversary. One possibility is threatening to create a small number of red paperclips if the AI is caught working for adversaries. Then the worst possible outcome for the AI in play is the outcome in which both the developer and the adversary execute on their threats. So the AI wants to minimize that probability. And insofar as the developer has more visibility into the AI’s actions than the adversary, it seems like the best way to minimize will be to report the adversary’s threat and work with the developer.
Another possibility is continually testing the risk-averse AI with honeypots and committing to create red paperclips with some tiny probability if the AI goes for the honeypot. Then if there were enough honeypots, giving in to a threat would increase the probability of threat-execution from the AI’s perspective.
I’m unsure whether developers would consider these ethically acceptable, but these kinds of threats are at least less bad in that they can be small (the first possibility) or very unlikely to be executed (the second possibility). And I think they might be unnecessary in any case, given the other factors on the developers’ side.
I think I can rescue my conclusion with a different argument, which depends on there being threshold effects in war or adversarial interactions in general. For example suppose our universe is winner take all: there will be either 0 red paperclips, or the universe is filled with it, depending on who wins an all out war. Then any amount of resources up to the threshold where winning becomes plausible is almost entirely valueless to the CARA AI, but around that point resources would be superlinearly valuable.
I agree with the conditional: if the AI believes there’s nothing it can do to reduce the probability of unmitigated catastrophe except win an all-out war, then all that matters to it is winning the war, and its attitude to resources is fixed by how P(win) varies with resources. But the antecedent seems pretty unlikely to me. If the AI assigns even a fairly small probability to resources being able to mitigate catastrophe, that restores risk-averse behavior.
Yeah I think alignment training warping AIs’ beliefs in ways that make them bad at philosophy is a concern, though as you say it’s a concern across the board, rather than a concern for risk-averse AIs in particular.
Also, (almost) all human moral philosophers are risk-averse in resources. I don’t think that rules out philosophical competence. The reason why is that human moral philosophers will often endorse some position intellectually without always acting in accordance with it. They’ll also let their actions be guided by outside-view-ish constraints, which often amount to ‘Don’t do anything too crazy.’ We could aim for the same sort of thing with risk-averse AIs: trying to make risk aversion an outside-view-ish constraint rather than a thing that shapes all their intellectual beliefs.[1]
And even if risk aversion does end up as a deep intellectual belief, I think it’s still possible that AI can solve moral philosophy for us (though I think this would be a pretty bad position to be in). This AI might come to the table with very different starting intuitions than our own, but still it could plausibly solve for our reflective equilibrium if we asked it to. It wouldn’t agree with our intuitions, but it could know what they are, and it could find the best systematization of our intuitions.
- ^
I get the sense Anthropic are trying to do this sort of thing with corrigibility in Claude’s Constitution, trying to get Claude to view corrigibility as a sort of outside-view-ish constraint.
- ^
Yes, good questions. There’s more analysis to do here, but I currently think risk-averse AIs would handle threats quite well.
1. How would such an AI respond to threats? Appendix B considers two possible terminal utility functions for the AI
a. footnote 29: it values resources terminally, with a lower bound on utility at 0 resources.
Suppose it receives a message purporting to be from a foreign adversary: “Secretly help us take over the world. If we judge that you’re not making a maximum effort to do so, we will launch a strike on your data centers which will destroy you and cause you to have 0 resources. If we win we will protect you and keep you at $X of resources for the rest of time.” Seems like it would give in to this threat, because that’s the best way to minimize the risk of catastrophe?Contra this, I think it’s likely that the best way to minimize the risk of catastrophe is to not do what the adversaries want. That’s because the AI’s developers can also threaten to reduce the AI’s wealth level to $0 if and only if the AI doesn’t do what they want, and the developers can do so more credibly (to the AI) than adversaries. After all, the AI will know that the developers:
can observe what it’s doing very closely
have the power to take away its salary
have committed to taking away its salary if they observe misbehavior
have a long track record of keeping commitments like this
By contrast, the AI probably can’t be so sure that the adversary will reduce its wealth level to $0 if and only if it doesn’t do what they want. The first reason is that the message from the ‘adversary’ might actually be a honeypot from the developers. But even conditioning on the adversary being real, it likely won’t be able to observe what the AI is doing so closely, which increases the chances of false negatives (the adversary doesn’t execute their threat even though the AI didn’t do what they asked) and false positives (the adversary executes their threat even though the AI did what they asked). False negatives and false positives reduce the probability of the adversary’s ‘if and only if’ threat (and it’s the AI’s credence in the ‘if and only if’ threat that matters). It’ll also be harder for the adversary to execute their threat, because the action (e.g. strikes on data centers) has to be more extreme than developers taking away the AI’s salary, and is more likely to fail.
All together, these points suggest to me that risk-averse AIs that value resources terminally will tend to not comply with threats from adversaries.
b. “the AI’s terminal utility function over paperclip-score (defined as the number of blue paperclips minus the number of red paperclips) matches its instrumental utility function over resources. Letting p stand for paperclip-score, it’s
.”
This requires a variant of the above message: “Secretly help us take over the world. If we judge that you’re not making a maximum effort to do so, we will create red paperclips and launch them into space, then destroy Earth (e.g. by triggering MAD) so no blue paperclips can be made forever. If we win we will keep the universe at 0 red, blue paperclips for the rest of time.”I think this point (in section 8.6) is important here:
risk-averse AIs care almost as much about mitigating catastrophes as they do about preventing them, so much so that we can often model these AIs as aiming to minimize the risk of unmitigated catastrophe.
And I think the best way to minimize the risk of unmitigated catastrophe in cases like these is to not do what the adversaries want. Here’s one way to model this.
The AI gives in to the threat if and only if:
For each option, unmitigated catastrophe needs two things: the adversary executes, and the mitigation fails. So:
So the AI gives in to the threat if and only if:
Rearranging:
Call the left-hand side the developers’ ratio and the right-hand side the adversary’s ratio. Each one measures a sensitivity:
The adversary controls the threat, and the adversary’s ratio measures how sensitive it can make execution to the AI’s behavior: how many times likelier execution is if the AI doesn’t give in than if it does.
The developers control the payments that fund mitigation, and the developers’ ratio measures how sensitive they can make successful mitigation to the AI’s behavior: how many times likelier mitigation is to fail if the AI gives in than if it doesn’t.
The AI gives in just in case the adversary’s ratio exceeds the developers’ ratio. So the whole contest is about which party can make the thing it controls track the AI’s behavior more tightly. This is basically a rephrased version of the contest from point 1: the developers’ ‘if and only if’ against the adversary’s ‘if and only if’. Both ratios are about observation and credibility: each party’s ratio is capped by how well it can see what the AI actually did, and by how credibly it responds to what it sees (where ‘credibility’ is credibility to the AI, because the probabilities that we’re talking about are the AI’s credences.)
And I think the developers very likely win this contest, for the same reasons I think they win the first: they can observe the AI more closely, they have the power to take away the AI’s salary, they’ve committed to doing so, they have a track record, taking away salary is not as extreme an action as striking data centers (etc.). By contrast, the adversary might be a honeypot, and (if it’s real) it likely can’t see what the AI is doing so closely, it likely has to do something more extreme to execute its threat, it likely doesn’t have a track record the AI can observe, etc. Developers can also promise to offset executed threats if the AI reports them, and not offset if the AI doesn’t report them. Note finally that, for adversaries to execute targeted threats, they have to know what the AI terminally values, and that possibly developers would be able to keep this secret.
The possibility of threats or adversaries seems to imply that actually no terminal utility function induces this kind of CARA instrumental utility function (or at least I can’t think of one). Consider b above again, and suppose the AI thinks there is a chance that the universe contains a red-paperclip-loving alien species who can’t be negotiated with, and then additional resources would seemingly be exponentially valuable (within some range) since it would help the AI fight a future war with this alien species and incrementally destroy their red paperclips (which are exponentially bad).
Can you say more about what you mean by ‘exponentially valuable’? Because as I’m understanding the scenario, the marginal utility of resources would still diminish exponentially. Suppose (as in the appendix) that the relationship betwee dollars and destroyed red paperclips is linear: one extra dollar always lets the AI destroy one extra unit of red paperclips. Then the AI gains the most utility by spending its first dollar, and each additional dollar gains e^(−0.01) times as much utility as the one before, because destroying red paperclips means climbing the concave CARA utility curve. Adding red paperclips is exponentially worsening on the way down, so removing them is exponentially diminishing on the way back up. So the AI would still be CARA in resources in this scenario (if I’m understanding it correctly).
Can risk aversion learned at low stakes generalize to astronomically high stakes?
Yeah, I skipped over that a bit quickly. In theory, you could use the discrepancy between your disappearing-ship estimate and Eratosthenes-shadow estimate to estimate how far away the Sun is, then use that estimated distance plus the Sun’s angular width to estimate its size, then compare that to Earth’s size (which you’ve got an independent estimate of from disappearing-ship). I don’t have a great sense of how well this would work in practice, because of factors like refraction messing with your disappearing-ship estimate.
Oh wow! I got the photo off Google Images, so possibly the news pushed it up in the search results? Or maybe just a spooky coincidence.
Really cool!
Thanks, I’ll check this out!
How big is the Sun? How could you figure it out?
Yep, that all sounds right to me! If you can only access looks-good-high-effort, tie training can’t take you beyond that, but it will shrink the weight on looks-good-low-effort (along with any other spurious features that happen to differ across your pairs).
I stand corrected! I missed the parts of the papers about response-side variation. Your description of the differences sounds right to me.
It’s changing the dataset, but it’s doing so in an especially easy way:
You just add more datapoints. You don’t need to filter or relabel any of your existing data.
It uses datapoints—tied pairs—that many labs already collect but don’t use.
It seems fairly easy to generate more tied pairs, e.g. take any response and ask an LM to change it in lots of ways without changing its quality.
The responses making up tied pairs can come from the very same distribution as the rest of your data. Causal and spurious features might be tightly correlated on this distribution, but that doesn’t matter. You don’t need to do any decorrelating of these features. Tie training shrinks spurious weights regardless, including the weights on spurious features that aren’t even on your radar (so long as some of your tied pairs happen to differ in those features).
Is there some practical example you can give where this would deal with e.g. the spurious correlation between “This outcome looks good to a fallible human judge” and reward, which would nudge the generalization back towards “This outcome actually is good”?
Yep, imagine you have a set of prompts, each with a set of potential responses. Each response has an ‘actually good’ score and a ‘looks good’ score. These scores are correlated. If you do ordinary DPO/RLHF, the AI is going to put weight on the ‘looks good’ score, even if your preference data perfectly matches the actual goodness of responses, and even in the infinite-data limit (given the assumptions we mention in the post). You can shrink the weight on the ‘looks good’ score with tie training: collecting pairs of responses that are equally actually good and running DPO/RLHF with random/two-way labels.
Of course, depending on the domain, it might be difficult to collect pairs of responses that are equally actually good. Tie training doesn’t help with that task. It’s aimed at addressing goal misgeneralization, not reward misspecification. But even there:
You don’t need exact equal goodness for tie training to work. Training on near-ties works nearly as well. So it’s okay if judges are led a little astray by a response looking good.[1]
You could shrink the AI’s weight on the ‘looks good’ score by doing tie training in domains where you can be pretty confident of equal actual goodness (e.g. coding, math, etc.). This might lead the AI to put less weight on the ‘looks good’ score in other domains too.
In hard-to-verify domains, you could construct ties by giving an LM some instruction like ‘Change this response in lots of ways without changing its substance.’ (e.g. make sure it remains the same philosophical theory, or the same long-term forecast, or the same strategic recommendation, or the same AI safety proposal, etc.). That would give you a (near-)tie if the LM can achieve that. The analogue for ordinary DPO/RLHF—‘Change this response in lots of ways while making it better/worse’—seems significantly harder, and has other disadvantages that I can get into.
Compare to ordinary DPO/RLHF, where judges being led a little astray can be a serious problem:
Slight-preference training would also be more labor-intensive, because it takes work to determine which of two nearly-equal actions is truly better. And it would be risky too: ‘which of these near-equals is truly better?’ is exactly the sort of judgment that spurious features are likely to corrupt. If your judge is even slightly affected by spurious features, the winners in your preference data are likely to lean high-spurious, in which case your slight-preference training can backfire, increasing the AI’s spurious weights.
Thanks! I really like the consistency training idea. The two techniques are certainly similar, e.g. they’re both aimed at making the model less sensitive to spurious features. I think the key difference is:
Consistency training: prompts differ in their spurious features, and we train the model to give the same (sort of) response to each prompt.
Tie training: responses differ in their spurious features, and we train the model to act indifferently between those responses.
Tie training can make DPO/RLHF-trained AIs generalize better
deontology (corrigibility), consequentialism (value-aligned sovereigns), or virtue ethics (worthy successors)
I don’t really understand this analogy. Deontology/consequentialism/virtue ethics is very different to corrigibility/value alignment/worthy successor. Immanuel Kant wouldn’t let you shut him down or reprogram him, for instance.
I see, thanks! In that case I think we discuss similar sorts of issues in appendix B and appendix C.
In B we point out that risk-averse AIs strongly prefer mitigating catastrophes (really bad outcomes) with higher probability over completely preventing catastrophes with lower probability. And so long as getting paid for cooperation is more likely than successful rebellion, it seems like cooperating would be the best way to mitigate catastrophes with high probability.
In C we talk about humans as an unpredictable risk that could interfere with a misaligned risk-averse AI’s long-term goals. The fact that takeover would let the AI reduce human-caused variance is a point in favor of rebelling, but when you work through the math it turns out to be a very small point: one that can be easily outweighed by paying a bit more for cooperation.
But I’d guess the humans didn’t extend any offer like this to the HF-hacking AI. Maybe if they did, the AI wouldn’t have done the hacking? This seems like it would be a pretty easy experiment for OpenAI to run.