I think like probably others, the degree to which the model can experiment and learn to control itself (for say the coding example) is a crux. For example let/help the model do a process like:
Generate 100 alternative instructions/self-instructions;
Run fresh copies under each;
Have the model itself classify which still exhibit the unwanted pattern;
Optimize those instructions over repeated rounds;
Perhaps separate controller and coding-agent contexts;
See whether a reliable intervention emerges.
The J-space/global workspace results also seem relevant. There apparently is some relatively low bandwidth shared internal state which can be read and causally manipulated. Maybe the problem is less “the Talker isn’t in control” and more that current models have pretty poor self-model/executive control over their own learned policies. (So less actual Talker/Doer split)
The situations
It takes a very large amount of prompting effort to achieve eventual success
No amount of effort ever succeeds
Seem substantially different. (1) is a self modelling issue, (2) is like the analogy where the diplomat just obviously can never succeed in changing the countries action, only explain it plausibly but incorrectly.
OK however that could be a point against AI takeover. Its now likely that when “Sable” comes it will be in a Lean verified sandbox for which there is no hope of hacking escape. AI Doom is very path dependent. While your scenario apparently didn’t depend on hacking skills, the general point applies.