Well said. The argument in this post is the exact reason Jer and I thought the Apollo founders (Marius and Lee) were so intriguing when we met them in early 2023. They, almost alone, were able to state a compact hypothesis that underpinned a line of attack that had a shot at directly solving alignment. That was something we’d only rarely if ever seen done before.
The way I’d phrase the Apollos’ hypothesis would be: “Deception is fundamentally more computationally expensive to execute that honesty, because a deceptive entity needs to keep at least two sets of mental books (the truth and the lie(s)) while an honest entity only needs to keep one.” If that hypothesis is true, and if we can confirm that it’s true with high confidence, then we may indeed be able to build deception detectors that are 100% reliable across a relevant set of scenarios. That in turn would be a sufficient condition to train for honesty in a way that scales to superhuman intelligence—either across a defined subset of scenarios, or in full generality.
Now of course each of those subproblems (confirm that property of deception is true; use it to build a perfect deception detector; implement honesty training signal; …) is itself very hard, and the underlying hypothesis could turn out to be false. But you can see that this is at least a more interesting problem decomposition than “use GPT-n to align GPT-(n + 1)”. It’s more like someone postulating a new physical law and trying to build a particle accelerator to demonstrate it, than like the kind of deus ex machina one sees in perpetual motion and most alignment schemes. The physical law might turn out not be there, or it might only manifest at energies way higher than you can ever achieve with a particle accelerator, but you’re at least trying to discover something nontrivially interesting—figure out if you can bell the cat while it sleeps.
We never articulated it this way, but this is a major reason we thought Apollo could be promising very soon after they launched. Thought that might be worth mentioning here for the (possible) positive example, if nothing else.
The way I’d phrase the Apollos’ hypothesis would be: “Deception is fundamentally more computationally expensive to execute that honesty, because a deceptive entity needs to keep at least two sets of mental books (the truth and the lie(s)) while an honest entity only needs to keep one.” If that hypothesis is true, and if we can confirm that it’s true with high confidence, then we may indeed be able to build deception detectors that are 100% reliable across a relevant set of scenarios.
Was this ever discussed at length? It seems kinda dubious to me, because it’s not hard to imagine externalizing deception or changing a system’s ontology or amortizing/dispersing deception-related things, especially as ‘deception’ doesn’t seem like a terribly well-defined concept to begin with. (For example, imagine a tabular Q-learner which gradually learns an optimal policy by memorizing the returns from every action/state and is just a lookup table, in effect. Such an agent can converge to the optimal policy in the limit under standard assumptions. So it can learn the value of ‘deceptive’ actions without ever once doing anything remotely like ‘mental books’, never mind having to develop theory of mind or keeping multiple books. No deception detector could ever detect anything in this agent, even though it is about as simple as possible for an agent to be and easily implemented/approximated by other agents as a subroutine etc.)
I’m not sure if they discussed this approach at length anywhere in public, and your tabular Q-learner is indeed a limit case counterexample. My main point wasn’t that this was definitely or even likely to work, just that it seemed qualitatively more promising than the “AIs will do it for us” approaches that were prevalent at the time and are predictably starting to break down now. For what it’s worth, they also proactively brought up the need to track computations offloaded by the system (e.g., use of a calculator or third party tools) since they saw this as one possible vector of amortization/dispersion/externalization of deceptive cognition. That was something else we found encouraging.
I guess I tend to think of early ideas like these less as being about “could this work” and more as being about “is this in the vicinity of something that could work in the future”. For example, it might be impossible to catch every instance of deception (or even to fully define the concept usefully), but might be possible instead to define a class of models such that deceptive thoughts originating from those models fall into a usefully defined and detectable class. This would still be fundamental, just fundamental to that particular class of systems.
Well said. The argument in this post is the exact reason Jer and I thought the Apollo founders (Marius and Lee) were so intriguing when we met them in early 2023. They, almost alone, were able to state a compact hypothesis that underpinned a line of attack that had a shot at directly solving alignment. That was something we’d only rarely if ever seen done before.
The way I’d phrase the Apollos’ hypothesis would be: “Deception is fundamentally more computationally expensive to execute that honesty, because a deceptive entity needs to keep at least two sets of mental books (the truth and the lie(s)) while an honest entity only needs to keep one.” If that hypothesis is true, and if we can confirm that it’s true with high confidence, then we may indeed be able to build deception detectors that are 100% reliable across a relevant set of scenarios. That in turn would be a sufficient condition to train for honesty in a way that scales to superhuman intelligence—either across a defined subset of scenarios, or in full generality.
Now of course each of those subproblems (confirm that property of deception is true; use it to build a perfect deception detector; implement honesty training signal; …) is itself very hard, and the underlying hypothesis could turn out to be false. But you can see that this is at least a more interesting problem decomposition than “use GPT-n to align GPT-(n + 1)”. It’s more like someone postulating a new physical law and trying to build a particle accelerator to demonstrate it, than like the kind of deus ex machina one sees in perpetual motion and most alignment schemes. The physical law might turn out not be there, or it might only manifest at energies way higher than you can ever achieve with a particle accelerator, but you’re at least trying to discover something nontrivially interesting—figure out if you can bell the cat while it sleeps.
We never articulated it this way, but this is a major reason we thought Apollo could be promising very soon after they launched. Thought that might be worth mentioning here for the (possible) positive example, if nothing else.
Was this ever discussed at length? It seems kinda dubious to me, because it’s not hard to imagine externalizing deception or changing a system’s ontology or amortizing/dispersing deception-related things, especially as ‘deception’ doesn’t seem like a terribly well-defined concept to begin with. (For example, imagine a tabular Q-learner which gradually learns an optimal policy by memorizing the returns from every action/state and is just a lookup table, in effect. Such an agent can converge to the optimal policy in the limit under standard assumptions. So it can learn the value of ‘deceptive’ actions without ever once doing anything remotely like ‘mental books’, never mind having to develop theory of mind or keeping multiple books. No deception detector could ever detect anything in this agent, even though it is about as simple as possible for an agent to be and easily implemented/approximated by other agents as a subroutine etc.)
I’m not sure if they discussed this approach at length anywhere in public, and your tabular Q-learner is indeed a limit case counterexample. My main point wasn’t that this was definitely or even likely to work, just that it seemed qualitatively more promising than the “AIs will do it for us” approaches that were prevalent at the time and are predictably starting to break down now. For what it’s worth, they also proactively brought up the need to track computations offloaded by the system (e.g., use of a calculator or third party tools) since they saw this as one possible vector of amortization/dispersion/externalization of deceptive cognition. That was something else we found encouraging.
I guess I tend to think of early ideas like these less as being about “could this work” and more as being about “is this in the vicinity of something that could work in the future”. For example, it might be impossible to catch every instance of deception (or even to fully define the concept usefully), but might be possible instead to define a class of models such that deceptive thoughts originating from those models fall into a usefully defined and detectable class. This would still be fundamental, just fundamental to that particular class of systems.