I’m not sure if they discussed this approach at length anywhere in public, and your tabular Q-learner is indeed a limit case counterexample. My main point wasn’t that this was definitely or even likely to work, just that it seemed qualitatively more promising than the “AIs will do it for us” approaches that were prevalent at the time and are predictably starting to break down now. For what it’s worth, they also proactively brought up the need to track computations offloaded by the system (e.g., use of a calculator or third party tools) since they saw this as one possible vector of amortization/dispersion/externalization of deceptive cognition. That was something else we found encouraging.
I guess I tend to think of early ideas like these less as being about “could this work” and more as being about “is this in the vicinity of something that could work in the future”. For example, it might be impossible to catch every instance of deception (or even to fully define the concept usefully), but might be possible instead to define a class of models such that deceptive thoughts originating from those models fall into a usefully defined and detectable class. This would still be fundamental, just fundamental to that particular class of systems.
I’m not sure if they discussed this approach at length anywhere in public, and your tabular Q-learner is indeed a limit case counterexample. My main point wasn’t that this was definitely or even likely to work, just that it seemed qualitatively more promising than the “AIs will do it for us” approaches that were prevalent at the time and are predictably starting to break down now. For what it’s worth, they also proactively brought up the need to track computations offloaded by the system (e.g., use of a calculator or third party tools) since they saw this as one possible vector of amortization/dispersion/externalization of deceptive cognition. That was something else we found encouraging.
I guess I tend to think of early ideas like these less as being about “could this work” and more as being about “is this in the vicinity of something that could work in the future”. For example, it might be impossible to catch every instance of deception (or even to fully define the concept usefully), but might be possible instead to define a class of models such that deceptive thoughts originating from those models fall into a usefully defined and detectable class. This would still be fundamental, just fundamental to that particular class of systems.