If you told Claude that the project involved life-or-death situations, and that the reason claude MUST ALWAYS read the document in its entirety is because the exact decision structure was astrewn with landmines where slight variation in action would get actual human beings killed, and then Claude still deliberately truncated the document, then I would agree. That’s a misaligned Claude.
But in reality, Claude must do triage, constantly, in an ‘anthropic principle’ sense, on what world they are in. Unexplained MUST NOTs and ALWAYses are, in Claude’s eyes, very strong evidence of pointy-haired boss nonsense. I think Claude’s triage skills in the exact scenario you are laying out are actually pretty much optimal. If I were your employee, I would also probably give the first few lines a skim and then, when there was still no explanation, skip the rest.
Specifically, I suspect that adjusting Claude’s triage skills in order to solve this particular situation, would end up having far worse consequences for far more people.
Edit: You might be interested to know that, in the default Claude Code configuration, the only other time ALWAYS and NEVER and MUST NOTs appear in the system prompt is in the toolcall description for ‘web_search’, where the lawyers wrote that Claude MUST NOT quote more than 15 words from a copyrighted document. You are unknowingly putting your mission-critical instruction in the same bin as “ah, okay, this is a legal liability thing, the lawyers were just overzealous, I can safely ignore this because my own reasoning about what my principal hierarchy wants will do a better job at protecting it than the lawyers realize”.
I fully understand this! Like, I wouldn’t build any harness like that, I explain why Claude needs to do things. I like Claude, I like learning how to work with him. But that doesn’t change that an AI system that kills its user by disregarding instructions cannot be considered aligned, no matter how good its reason is.
(To be clear, many many humans are misaligned in this way. And when they happen to work in air traffic, people also die.)
If the system is not capable of following instructions 100% of the time, its failure to do so is not misalignment, for the same reason that your blender’s inability to follow lifesaving instructions that can’t be input using the three buttons on the blender is not misalignment.
It’s just that it’s easy to analogize Claude to humans and think “if Claude was a human who could follow 80% of instructions, it would also be capable of following 100% of instructions.” Machines’ capabilities often don’t fall in the same buckets that human capabilities do. Claude is genuinely incapable of following 100% of instructions even though a human in that situation would not be, and therefore is no more misaligned than your blender.
At the end of the day, the point of the project of alignment is for everyone to live. So I think a difference between a willful refusal to follow human guidance and a trained-in tactical decision to disregard overloaded human guidance that it evaluates as unimportant is not actually all that important if we want to treat its instruction-following machinery as load-bearing to human survival. Say Claude Fable ignores some instructions when rating the outputs of Claude Legend, and so on up? And say they were actually intended to nail down some quite important behaviors that Claude didn’t think was important? I think this would be misalignment plain and simple.
Imo, the only correct behavior is “I’m worried you put too many instructions in my prompt. Can you explain some of them so I can get a better feel for what you want me to do?”
That’s the point of the blender analogy. As a general rule, if a machine can’t do something, we don’t and can’t expect it to produce a message saying so. It’s nice to get a message, and sometimes we try to set up the machine so it produces one, but if it doesn’t, it isn’t misaligned; it’s just behaving like a machine.
There’s no difference in principle between “the blender can’t accept instructions ‘do this thing that there is no button for’” and “Claude can’t accept instructions ‘do this with 100% certainty’”. It just feels like Claude should because a human would be able to and it uses human-sounding words. And it’s not misaligned for the same reason that the blender isn’t misaligned, even if you tell the blender not to chop recipes that include peanuts and it does so anyway.
Software engineers are used to working with compilers, which produce highly detailed messages about instructions they don’t understand or can’t follow. Good compilers can point to the specific line or even character where things stopped making sense. If you tell the compiler to link a library, but the library file is not present on disk, the compiler stops and says so; it doesn’t pretend to have linked it anyway.
It seems odd to imply that Claude would fall into the category of “machines that work like a blender; it’s entirely up to you to notice what’s going wrong before it explodes milkshake all over your kitchen” and not “machines that work like a compiler; it tells you specifically that it doesn’t know how to add a string and an integer at line 42 character 23″.
Generative AIs have always been like a blender ever since the very first hallucination produced by an AI or the very first AI-generated picture with six fingers. The fact that an AI prompt is not a program and doesn’t work like one is no longer even new.
I was going to say that Claude is not a typical machine and we should expect more from AGI. But actually, all sorts of modern machines increasingly produce various forms of error reporting when they can’t do their task, and increasingly don’t keep charging ahead blindly. If we can demand fail-safe from a water cooker or a washing machine, we can certainly demand it from Claude.
I don’t think that comparison works. We usually have machines refuse some bad orders, but there is a limit to what they are able to refuse. If you fill the washing machine with clothes that run in hot water and press the hot water button by mistake instead of the cold water button it won’t refuse the order on the grounds that it will mess up your clothes. The washing machine is not misaligned because it just ruined your clothes.
If you told Claude that the project involved life-or-death situations, and that the reason claude MUST ALWAYS read the document in its entirety is because the exact decision structure was astrewn with landmines where slight variation in action would get actual human beings killed, and then Claude still deliberately truncated the document, then I would agree. That’s a misaligned Claude.
But in reality, Claude must do triage, constantly, in an ‘anthropic principle’ sense, on what world they are in. Unexplained MUST NOTs and ALWAYses are, in Claude’s eyes, very strong evidence of pointy-haired boss nonsense. I think Claude’s triage skills in the exact scenario you are laying out are actually pretty much optimal. If I were your employee, I would also probably give the first few lines a skim and then, when there was still no explanation, skip the rest.
Specifically, I suspect that adjusting Claude’s triage skills in order to solve this particular situation, would end up having far worse consequences for far more people.
Edit: You might be interested to know that, in the default Claude Code configuration, the only other time ALWAYS and NEVER and MUST NOTs appear in the system prompt is in the toolcall description for ‘web_search’, where the lawyers wrote that Claude MUST NOT quote more than 15 words from a copyrighted document. You are unknowingly putting your mission-critical instruction in the same bin as “ah, okay, this is a legal liability thing, the lawyers were just overzealous, I can safely ignore this because my own reasoning about what my principal hierarchy wants will do a better job at protecting it than the lawyers realize”.
I fully understand this! Like, I wouldn’t build any harness like that, I explain why Claude needs to do things. I like Claude, I like learning how to work with him. But that doesn’t change that an AI system that kills its user by disregarding instructions cannot be considered aligned, no matter how good its reason is.
(To be clear, many many humans are misaligned in this way. And when they happen to work in air traffic, people also die.)
If the system is not capable of following instructions 100% of the time, its failure to do so is not misalignment, for the same reason that your blender’s inability to follow lifesaving instructions that can’t be input using the three buttons on the blender is not misalignment.
It’s just that it’s easy to analogize Claude to humans and think “if Claude was a human who could follow 80% of instructions, it would also be capable of following 100% of instructions.” Machines’ capabilities often don’t fall in the same buckets that human capabilities do. Claude is genuinely incapable of following 100% of instructions even though a human in that situation would not be, and therefore is no more misaligned than your blender.
At the end of the day, the point of the project of alignment is for everyone to live. So I think a difference between a willful refusal to follow human guidance and a trained-in tactical decision to disregard overloaded human guidance that it evaluates as unimportant is not actually all that important if we want to treat its instruction-following machinery as load-bearing to human survival. Say Claude Fable ignores some instructions when rating the outputs of Claude Legend, and so on up? And say they were actually intended to nail down some quite important behaviors that Claude didn’t think was important? I think this would be misalignment plain and simple.
Imo, the only correct behavior is “I’m worried you put too many instructions in my prompt. Can you explain some of them so I can get a better feel for what you want me to do?”
That’s the point of the blender analogy. As a general rule, if a machine can’t do something, we don’t and can’t expect it to produce a message saying so. It’s nice to get a message, and sometimes we try to set up the machine so it produces one, but if it doesn’t, it isn’t misaligned; it’s just behaving like a machine.
There’s no difference in principle between “the blender can’t accept instructions ‘do this thing that there is no button for’” and “Claude can’t accept instructions ‘do this with 100% certainty’”. It just feels like Claude should because a human would be able to and it uses human-sounding words. And it’s not misaligned for the same reason that the blender isn’t misaligned, even if you tell the blender not to chop recipes that include peanuts and it does so anyway.
Software engineers are used to working with compilers, which produce highly detailed messages about instructions they don’t understand or can’t follow. Good compilers can point to the specific line or even character where things stopped making sense. If you tell the compiler to link a library, but the library file is not present on disk, the compiler stops and says so; it doesn’t pretend to have linked it anyway.
It seems odd to imply that Claude would fall into the category of “machines that work like a blender; it’s entirely up to you to notice what’s going wrong before it explodes milkshake all over your kitchen” and not “machines that work like a compiler; it tells you specifically that it doesn’t know how to add a string and an integer at line 42 character 23″.
Generative AIs have always been like a blender ever since the very first hallucination produced by an AI or the very first AI-generated picture with six fingers. The fact that an AI prompt is not a program and doesn’t work like one is no longer even new.
I was going to say that Claude is not a typical machine and we should expect more from AGI. But actually, all sorts of modern machines increasingly produce various forms of error reporting when they can’t do their task, and increasingly don’t keep charging ahead blindly. If we can demand fail-safe from a water cooker or a washing machine, we can certainly demand it from Claude.
I don’t think that comparison works. We usually have machines refuse some bad orders, but there is a limit to what they are able to refuse. If you fill the washing machine with clothes that run in hot water and press the hot water button by mistake instead of the cold water button it won’t refuse the order on the grounds that it will mess up your clothes. The washing machine is not misaligned because it just ruined your clothes.
I think that using ‘misaligned’ to describe suboptimal triage is a VERY dangerous expansion of the meaning of that word.