TBC, I wouldn’t say that I have any “proposed solution”, in the sense of a technical plan that I expect to work. What I have is avenues of attack (actually two of them, see §3 of my 2025 review) which are not (to my knowledge) categorically ruled out, and which thus merit further investigation. Hopefully such further investigation will eventually result in either a plan, or an argument that no plan of that type exists. If you think you have an argument that rules out one of these lines of attack, I’m generically open-minded and very interested.
But I disagree with your argument here, or don’t understand it.
First and foremost, I’m confused about whether you have in mind an LLM (like most people are thinking about) or a “brain-like AGI” (some yet-to-be-invented version of actor-critic model-based RL) (like I’m usually thinking about). If you’re talking about my technical alignment ideas, then those are specific to the latter. They don’t apply to today’s LLMs. I don’t think I have much special insight into “alignment” of today’s LLMs. But you seem to be assuming that we’re talking about LLMs (as we know them today) in many parts of this post—e.g. you talk about “produce the next token” (brain-like AGI would presumably have a wider action space than that, including invisible actions like choosing what to think about), and “training data” (RL agents generally have a “training environment” not “training data”), and “subliminal learning” (a specific LLM thing, AFAICT), and the CAST proposal (as I read it, it seems to only make sense in LLM-world), etc.
While the human society lets most individuals intervene with the reward function and training data for the belief system of others, allowing individuals to align each other to the community…
I think this is a misleading way to think about things, see my post Heritability, Behaviorism, and Within-Lifetime RL. My reward function is locked away inaccessible in my head. If you put me in a room with another person, they will affect the inputs (and thus also outputs) of my reward function, but that’s equally true if you put me in a room with a turtle, or a book. Either way, saying that the other person can “intervene” on my reward function has a misleading connotation that they can just decide what they want my reward function to be.
However, unlike the AIs, human circuitry related to scheming is pruned by the fact that potential victims usually have similar or higher capabilities, causing the schemers to be punished.
As an illustration that something is wrong with “honeypots-for-alignment” theory, that theory would predict that, if a man is in prison, and he tries to escape but gets caught and harshly punished a few times, then he stops wanting to escape prison at all. Even if, later on, there’s a prison riot, and the guards are all dead, and the front door is open, and there’s a bus idling outside about to drive the rest of the prisoners to a non-extradition country … the prediction of “honeypots-for-alignment” theory is that this man will not get on the bus, but rather sit patiently in the otherwise-empty prison, waiting for the cops to arrive and lock him back up. …Well, this prediction is obviously absurd. … [more at the link]
Or consider how rebellious teens don’t usually “learn” to be honest to their hated parents, teachers, and authority figures. Right? Adults don’t “learn” to be honest to hated authority figures either, they just get savvy enough to know what they can get away with.
Anyway, is your actual belief that humans have no innate social drives related to compassion or status-seeking? If so, we can talk about why I disagree.
I mean that humans are born with more primitive versions of such drives (e.g. valuing others smiling as a proxy for compassion), then instrumental convergence transforms these primitive versions into the actual approval reward/status seeking/etc in a manner similar to training making agents who already are in the basin of alignment/corrigibility/etc more aligned/corrigible. As for prisoners no longer wanting to escape, this is where your conjecture about the belief system is most promising: a few unsuccessful attempts instill a hard-to-reassess belief that future attempts have a low probability of success.
UPD: how similar is the belief-desire model to the master-slave model? Suppose that the belief system simply outputs what it believes to be the next tokens of, say, known theorems or of muscular actions responsible for pull-ups and the desire system steers the belief system in order to construct the outcomes like solving the problem or doing the exercise correctly. Then the desire and belief system would occupy the positions of the master and the slave.
Thanks for engaging!
TBC, I wouldn’t say that I have any “proposed solution”, in the sense of a technical plan that I expect to work. What I have is avenues of attack (actually two of them, see §3 of my 2025 review) which are not (to my knowledge) categorically ruled out, and which thus merit further investigation. Hopefully such further investigation will eventually result in either a plan, or an argument that no plan of that type exists. If you think you have an argument that rules out one of these lines of attack, I’m generically open-minded and very interested.
But I disagree with your argument here, or don’t understand it.
First and foremost, I’m confused about whether you have in mind an LLM (like most people are thinking about) or a “brain-like AGI” (some yet-to-be-invented version of actor-critic model-based RL) (like I’m usually thinking about). If you’re talking about my technical alignment ideas, then those are specific to the latter. They don’t apply to today’s LLMs. I don’t think I have much special insight into “alignment” of today’s LLMs. But you seem to be assuming that we’re talking about LLMs (as we know them today) in many parts of this post—e.g. you talk about “produce the next token” (brain-like AGI would presumably have a wider action space than that, including invisible actions like choosing what to think about), and “training data” (RL agents generally have a “training environment” not “training data”), and “subliminal learning” (a specific LLM thing, AFAICT), and the CAST proposal (as I read it, it seems to only make sense in LLM-world), etc.
I think this is a misleading way to think about things, see my post Heritability, Behaviorism, and Within-Lifetime RL. My reward function is locked away inaccessible in my head. If you put me in a room with another person, they will affect the inputs (and thus also outputs) of my reward function, but that’s equally true if you put me in a room with a turtle, or a book. Either way, saying that the other person can “intervene” on my reward function has a misleading connotation that they can just decide what they want my reward function to be.
I don’t think this argument works, here’s an example (copied from “Behaviorist” RL reward functions lead to scheming):
Or consider how rebellious teens don’t usually “learn” to be honest to their hated parents, teachers, and authority figures. Right? Adults don’t “learn” to be honest to hated authority figures either, they just get savvy enough to know what they can get away with.
Anyway, is your actual belief that humans have no innate social drives related to compassion or status-seeking? If so, we can talk about why I disagree.
I mean that humans are born with more primitive versions of such drives (e.g. valuing others smiling as a proxy for compassion), then instrumental convergence transforms these primitive versions into the actual approval reward/status seeking/etc in a manner similar to training making agents who already are in the basin of alignment/corrigibility/etc more aligned/corrigible. As for prisoners no longer wanting to escape, this is where your conjecture about the belief system is most promising: a few unsuccessful attempts instill a hard-to-reassess belief that future attempts have a low probability of success.
UPD: how similar is the belief-desire model to the master-slave model? Suppose that the belief system simply outputs what it believes to be the next tokens of, say, known theorems or of muscular actions responsible for pull-ups and the desire system steers the belief system in order to construct the outcomes like solving the problem or doing the exercise correctly. Then the desire and belief system would occupy the positions of the master and the slave.