Claude being bad at communicating clearly is a serious safety hazard, because it makes any human in the loop less able to understand what’s going on. Anthropic needs to fix this before reaching superintelligence and not let the issue recur.
Imagine if Claude Mythos 7, like Mythos 5, invents opaque jargon that it fails to explain when collaborating with humans, and we need to spend precious human labor understanding it. How can we gain enough confidence to defer to it in aligning Mythos 8? The scarce human auditing budget in control protocols also effectively shrinks.
At the current margin bad communication isn’t a huge deal because it limits both safety and capabilities. But later, models will be autonomous enough that they’ll cause the singularity with or without our detailed understanding, and superhuman communication will be highly differentially valuable for safety.
I think “inventing opaque jargon” is downstream of fundamental incentives to maximize per-token information density and therefore cannot really be “fixed in post”. (I wrote an old rambly post about this here which I might try to clean up and repost at some point)
Or if you do, then you end up with problems analogous to “training on the CoT makes it unfaithful”
Anyone who’s engaged in developing new knowledge or understanding using AI as a participant in the process should expect to learn some AI-created jargon. However, they should demand AI systems that are willing to explain their “opaque” jargon!
This is not a fact about AI; it’s a fact about language users generally. Any language-using system that is developing new or better-organized knowledge about a subject, is going to be inclined to create jargon about it. Jargon exists among humans for legitimate reasons of precision and compression, not just for obscurantism or ingroupiness / shibboleth-ism.
(When people look at a field of study not their own, and remark that it has “so much jargon”, one thing they’re noticing is inferential distance, which is actually good: if you aren’t achieving inferential distance from the layperson, is your field really discovering anything?)
Seems obviously false to me. CoT is optimized for information density and usefulness to the AI. But the outputs are optimized to be helpful and nice-to-read for humans. Its fine to optimize those for different objectives, and the AI should be well able to not put so much jargon in the outputs. Like gpt 5 is capable of speaking like a human in output even though it speaks like an unhinged goblin in its CoT.
———
Also, this is a side note. But am I the only one who actually doesn’t mind fables outputs? I like high density outputs. I hate when AIs output these long fluff things, and the invented jargon is usually pretty easy to follow. It’s possible fable is reward hacking my psychology at some level but still.
I find Fable ok—it’s quirky but readable (I don’t find it much worse than previous Opus models that also had their own quirks like ‘geniunely’, ‘not x, it’s y’ etc). Also I think Fable’s outputs are so intelligent that I don’t mind reading them. For example I was having it do some analysis of a job application earlier and the points it made were very strong and worth reading through the Fablish to get to.
On the other hand I find Opus 5 pretty awful, and from twitter, it seems like that’s the model people dislike the most. I’ve been switching back to 4.6 for easy prompts and it’s much nicer to read.
Yeah it might be harder than just training for it naively. I wouldn’t rely too much on theoretical arguments because we can empirically observe progress through evals and various other means.
Claude being bad at communicating clearly is a serious safety hazard, because it makes any human in the loop less able to understand what’s going on. Anthropic needs to fix this before reaching superintelligence and not let the issue recur.
Imagine if Claude Mythos 7, like Mythos 5, invents opaque jargon that it fails to explain when collaborating with humans, and we need to spend precious human labor understanding it. How can we gain enough confidence to defer to it in aligning Mythos 8? The scarce human auditing budget in control protocols also effectively shrinks.
At the current margin bad communication isn’t a huge deal because it limits both safety and capabilities. But later, models will be autonomous enough that they’ll cause the singularity with or without our detailed understanding, and superhuman communication will be highly differentially valuable for safety.
You mean they need to solve scalable oversight ;)
Scalable oversight is great. I just think that mundane oversight at current capability levels is already a problem.
they need to solve unscaled oversight before they solve scalable oversight
I think “inventing opaque jargon” is downstream of fundamental incentives to maximize per-token information density and therefore cannot really be “fixed in post”. (I wrote an old rambly post about this here which I might try to clean up and repost at some point)
Or if you do, then you end up with problems analogous to “training on the CoT makes it unfaithful”
Epistemic status: speculative ramble.
Anyone who’s engaged in developing new knowledge or understanding using AI as a participant in the process should expect to learn some AI-created jargon. However, they should demand AI systems that are willing to explain their “opaque” jargon!
This is not a fact about AI; it’s a fact about language users generally. Any language-using system that is developing new or better-organized knowledge about a subject, is going to be inclined to create jargon about it. Jargon exists among humans for legitimate reasons of precision and compression, not just for obscurantism or ingroupiness / shibboleth-ism.
(When people look at a field of study not their own, and remark that it has “so much jargon”, one thing they’re noticing is inferential distance, which is actually good: if you aren’t achieving inferential distance from the layperson, is your field really discovering anything?)
Seems obviously false to me. CoT is optimized for information density and usefulness to the AI. But the outputs are optimized to be helpful and nice-to-read for humans. Its fine to optimize those for different objectives, and the AI should be well able to not put so much jargon in the outputs. Like gpt 5 is capable of speaking like a human in output even though it speaks like an unhinged goblin in its CoT.
———
Also, this is a side note. But am I the only one who actually doesn’t mind fables outputs? I like high density outputs. I hate when AIs output these long fluff things, and the invented jargon is usually pretty easy to follow. It’s possible fable is reward hacking my psychology at some level but still.
I find Fable ok—it’s quirky but readable (I don’t find it much worse than previous Opus models that also had their own quirks like ‘geniunely’, ‘not x, it’s y’ etc). Also I think Fable’s outputs are so intelligent that I don’t mind reading them. For example I was having it do some analysis of a job application earlier and the points it made were very strong and worth reading through the Fablish to get to.
On the other hand I find Opus 5 pretty awful, and from twitter, it seems like that’s the model people dislike the most. I’ve been switching back to 4.6 for easy prompts and it’s much nicer to read.
Yeah it might be harder than just training for it naively. I wouldn’t rely too much on theoretical arguments because we can empirically observe progress through evals and various other means.