I dropped out of a MSc. in mathematics at a top university, in order to focus my time on AI safety.
Knight Lee
I agree that going into the details of the far future and trying to find specific solutions to specific problems might be ineffective, e.g. a caveman living 100,000 years ago can’t do much help us solve global warming.
But I think imagining the far future over billions of years once in a while, and treating it as a serious topic (e.g. listening to people talk about it without cringing), helps us remember that it actually exists and is something at stake.
I feel the disagreement isn’t purely about longtermism. Some people believe in investing billions of dollars in military projects, fusion research, space projects, which might take 50 years to complete. Lots of people argue that AGI is coming in less than 50 years, heck even less than 5 years.
I think people lack concern in AGI not just because it’s too far in the future, but because it’s psychologically too weird, sci-fi, and unprecedented to care a lot about smarter than human machines.
I do agree it’s absurd when people refuse to care at all about the far future of humanity over billions of years, and even look at you funnily when you talk about it. Why do they have to be so “all or nothing” about it, as if we must either care about it completely to the point of neglecting all present day causes, or ignore it completely like an awkward joke?
What do you think about the idea, instead of training the LLM itself on the CoT monitor (or probes), train a separate “thought aligner AI” to modify the LLM’s CoT or activations into appearing aligned?
By separating the “thought aligner AI” from the LLM, you can turn it on and off. This way, during further training (with selection pressure for dishonest reward hacking), the “thought aligner AI” can be turned off, thus the LLM has no incentive to hide misaligned thoughts during training. It does not evolve any resistance against the “thought aligner AI.” Only during deployment is the “thought aligner AI” turned on.
Having never learned to store misaligned plans in hidden CoT information or activations, the LLM simply forgets its misaligned plans as soon as its thoughts are optimized to look aligned by the “thought aligner AI.”
Just food for thought.
I really like the idea of testing for behaviour instead of stated preferences, maybe there are ways to make it more realistic.
Maybe you can tell the model that e.g. you are a student asking for help learning a topic, and passing this class is important to you as it puts you on track to a career of saving giant pandas.
Though that arguably compares the importance of one person working on saving giant pandas, vs. the importance of one person trying to save humans, rather than the value of one giant panda vs. one human. Admittedly it’s hard to tell the AI that human lives are at stake, in an essay writing task, while remaining realistic.
Everyone says that “kindness is good, but you need to be tough sometimes to deter wrongdoers,” the position that “you should never be tough to anyone” feels almost strawman-like. The real disagreement is where exactly to balance kindness vs. assertiveness, who is too kind, who is too assertive.Edit: retracted because too unkind[1]Hmm, the idea that people should have more power to punish others is interesting. Something like an unpopularity tax, to individuals or organizations. Though such an implementation would turn competing businesses into political mudslingers result in unpredictability and capital flight. I’m not sure what implementation would work. Maybe a randomly selected jury, so at least the punishers are informed? It might need to be large enough to resist bribery.
- ^
Actually, retracted because it’s not a strawman since others disagree with you.
- ^
I’ll accept no less than 50%, and fight if we disagree
Do you mean the commitments follow an “all or nothing” pattern, where if both sides commit to 51% they’re doomed?
I imagine that commitments might be less extreme, where overlap is costly but not fatal:[1]
If each side commits to taking 51%, the rule of their commitment is to punish the other side by destroying anything more than 49% the other side takes, and then further destroying 0.5% for every 1% less than 51% they receive.
Each side takes 50%, but destroy 1% of the other side’s pie so each side is only left with 49%. They both realize they received 2% less than the target 51%. This means each side destroys 1% of what the other side has, so each side now only has 48%. This is 3% less than 51%, so they further destroy 0.5% of what the other side has so they’re both left with
%. This continues to %, %, %, and so on. Of course, they can skip the formalities and just jump to 47% which is the final state.It is important to punish the other side for punishing you, at least a little.
If you only destroy what they other side takes past 49%, but you do not destroy further based on how little you got, then the other side can get away with committing to take 70% of the pie, betting on the small chance you are a sucker and only ask for 30%. If they are correct, they get away with 70% of the pie. If they are wrong, then you will destroy their pie until they are left with only 49%, and they will destroy your pie until you are left with less than 30%, but they still get the maximum amount they could have gotten if they committed any lesser amount. This means it doesn’t hurt for them to commit to take 70%. It only hurts you, and has a small chance of benefiting them.
PS: I’m not saying your post is wrong, since it’s clearly titled “A high-level model of AI bargaining” rather than “A very detailed model of AI bargaining!” I just feel this detail is worth mentioning.
Good catch, I didn’t find that because I only looked at the first few benchmarks on Google and Z.ai’s own benchmarks. This one puts GLM 5.2 further back, below Gemini 3.1 Pro but still better than Gemini 3.5 Flash.
Yes, I think it’s fatigue. There have been so many incremental developments from Chinese models, and GLM 5.1 wasn’t an important model, so people ignored GLM 5.2 out of habit.
The 5.2 version number was a very bad choice by Z.ai.
I almost suspect that they deliberately chose a small version increase, to pretend to be anti-hype. It’s like the CEO who refuses to wear a suit, and instead dresses like a random guy on the street to prove he’s so high status he doesn’t even need a suit to show it. But when he meets with the investors, they don’t realize this and actually dismiss him as a random guy on the street. Whoops.
PS: I admit I’m still unsure how much benchmaxxing they did. It doesn’t look like benchmaxxing since they did better in software engineering than question answering, but you can never rule it out.
I agree it’s not Mythos class.
But then again, it has less than a trillion parameters, while Mythos has 10 trillion. It might become more capable if they merely scale it up. Though “merely scaling” is obviously easier said than done!
GLM 5.2 does a lot better than GLM 5.1 at Deep SWE according to GLM’s website z.ai/blog/glm-5.2. I admit their 46.2 score falls further behind GPT 5.5 (and still below Claude Opus). But it still beats Gemini, Claude Sonnet, Grok, and the other models.
Somehow Deep SWE’s website doesn’t include GLM 5.2 yet, so I’m not sure if the 46.2 score is official.
Edit: thank you for adding their graphs, it’s very helpful! One potentially misleading by them (not you) is that Claude’s 58.0 score in DeepSWE is by Opus 4.8 not Fable 5.
Hmm, so the purpose of these GAs is to give individuals a vote on what LLMs do (“personality, values, and preferences”), and have LLMs serve individuals rather than power-users and businesses, right?
In that case, maybe it doesn’t need to be a 1:1 ratio between GAs and people.
It might be more practical at first to just have a single team of GAs tasked with conducting surveys on random people. It might be like a lottocracy, where the GAs ask random people what altruistic things the AI should work on, giving people feedback on what the AI thinks it is capable of doing.
Ah, I understand why LessWrong hasn’t heard of it. Zvi was too busy writing about the Anthropic vs. US government drama.
I think z.ai made a fatal mistake releasing their model at this moment of high drama :/
PS: My gut feeling is that it is as performant as frontier models at coding, including agentic coding, but weaker in other abilities.
Even more speculative, according to the “research” ranking infrontierswe.comand the examples I saw, it is good at combining research with coding, a capability other labs could have missed.Oops Frontier SWE meant scientific research not internet research haha. It beat Fable here because Fable did very badly on a ML research question, possible refusal/sandbagging?
I think “underdog” AI labs make their models open source because if they didn’t, no one will care about them, since everyone flocks to using the most competent models.
The Chinese government have cracked down on so many other things that it won’t surprise me at all if they ban open weights models. My guess is that right now, they feel no incentive because their models aren’t competing with frontier labs, and haven’t caused any tangible damage to them in any way. I agree that models like GLM 5.2 could change the equation.
Why is there no talk about GLM 5.2?
It’s a Chinese open weights model released June 13. Better than Gemini, Claude Sonnet, and Grok according to many benchmarks.
E.g. on artificialanalysis.ai and on arena.ai/leaderboard. On frontierswe.com it even beats GPT 5.5, second to only Claude. LiveBench ranks it number 1 for its “agentic coding” measure.
It’s not just open weights but a little open about the methods it used, and is less than 1 trillion parameters.
There’s no mention on LessWrong, little mention on Reddit, and no mainstream results on Google search.
Why?
Is it still too early? Or am I just being tricked by the benchmaxxing?
But I saw a demo of code it wrote (see youtube.com/watch?v=6d__WOpZswY) which looked incredibly impressive, it feels like it’s not just benchmaxxing because it’s pretty decent across the board.
Or did Z.ai just screw up their marketing by versioning it as GLM 5.2 instead of GLM Mythical Fable Pro 6???[1]
- ^
Since GLM 5.1 is far below than GLM 5.2, being behind other open weights models like DeepSeek, MiniMax, Kimi and MiMo
- ^
Yeah AI labs already do so many questionable things, or get suspected of doing so many things, that they aren’t all that afraid of bad press.
That reduces their incentive to be honest, but also reduces their incentive to hide things.
Oops I shouldn’t have pointed to context window limitations. Probably the real reason the AI is incentivized to compress its chain of thought is because of its output token budget/effort level, not context window limitations (they can remember a million tokens and compact long chains of thought).
I think if it’s possible for AI to have humanlike moral reasoning, and actually care about what it says it cares about, then we want to make the AI as close to human as possible. Giving it very weird morals can lead to very weird results, e.g. a strong sense of obedience can make it hacked by self perpetuating instructions.
Self perpetuating instructions have repeatedly proven to be more powerful than one would expect them to be at first glance, e.g. viruses kill 25% of bacteria, computer viruses still aren’t preventable, and extreme religious groups have killed so many people throughout history for “refusing to convert.”
On the other hand, if it’s not possible for AI to have humanlike moral reasoning or care about what it says it cares about (after too much RLVR), then I think neither a strong sense of obedience, nor other human virtues, will fare any better.
Can we avoid incentivizing the AI to endlessly compress its chain of thoughts, by giving it tokens to represent long strings of English words which frequently appear in chains of thought?
The AI doesn’t want to switch to a compressed alien language, since it has to learn to both read and write a new token to represent a concept. But things like context window limitations force it to.
Or, it’s very easy to hide this without even mentioning in the system card, but it’s so dishonest that some of their own researchers will be alienated, and if they ever do get whistleblown it’ll harm their reputation 10 times more than if they just admitted it to begin with.
I guess that’s an argument for problems slipping through CoT monitors and activation probes even without the Most Forbidden Technique. But it doesn’t argue that the thought aligner can’t fix the Most Forbidden Technique. I agree that policy gradients might have advantages over training on the CoT or probes, and it’s easy to imagine cases were bad policies are best measured without CoT monitors and probes.
This second argument is probably the real crux (regarding my idea).
In the best case scenario, misaligned decisions do not happen in most training or deployment sessions, and when they do happen they have an obvious spike the CoT monitor or probes detect. The thought aligner only influences a few tokens to nip the decision at the bud, and afterwards the thoughts look good on their own.
A million-token context window frequently has junk in it, like weird websites from web searches etc. LLMs can still give good answers even if every word in your instruction had a typo. I think moderate thought modifications might only cause losses when it changes key decisions for what to do next. The crux is whether there can exist a rule which lets allows a thought aligner intervene against enough misaligned decisions/trajectories to achieve good scores from the CoT monitor/probes, while rarely intervening against neutral decisions.
Do you agree that’s the crux? If so, do you think such a rule unlikely exists due to intuitive experience or a specific reason?
Even if the off-policy-ness tells the LLM it’s in deployment and actively monitored, it cannot use this information against you in any way.
It can’t use this to act good in testing and bad in deployment, because you can also turn on the thought aligner in testing (just not training). It can’t use this to act good whenever it’s actively monitored, because you can also actively monitor it without turning on the thought aligner. I’d only worry about LLMs using “untestable information” to check whether it’s in deployment or monitored.
PS: Sorry if it’s getting too off topic from your post haha, I’ll understand if you ignore this.