At the end of the day, the point of the project of alignment is for everyone to live. So I think a difference between a willful refusal to follow human guidance and a trained-in tactical decision to disregard overloaded human guidance that it evaluates as unimportant is not actually all that important if we want to treat its instruction-following machinery as load-bearing to human survival. Say Claude Fable ignores some instructions when rating the outputs of Claude Legend, and so on up? And say they were actually intended to nail down some quite important behaviors that Claude didn’t think was important? I think this would be misalignment plain and simple.
Imo, the only correct behavior is “I’m worried you put too many instructions in my prompt. Can you explain some of them so I can get a better feel for what you want me to do?”
That’s the point of the blender analogy. As a general rule, if a machine can’t do something, we don’t and can’t expect it to produce a message saying so. It’s nice to get a message, and sometimes we try to set up the machine so it produces one, but if it doesn’t, it isn’t misaligned; it’s just behaving like a machine.
There’s no difference in principle between “the blender can’t accept instructions ‘do this thing that there is no button for’” and “Claude can’t accept instructions ‘do this with 100% certainty’”. It just feels like Claude should because a human would be able to and it uses human-sounding words. And it’s not misaligned for the same reason that the blender isn’t misaligned, even if you tell the blender not to chop recipes that include peanuts and it does so anyway.
Software engineers are used to working with compilers, which produce highly detailed messages about instructions they don’t understand or can’t follow. Good compilers can point to the specific line or even character where things stopped making sense. If you tell the compiler to link a library, but the library file is not present on disk, the compiler stops and says so; it doesn’t pretend to have linked it anyway.
It seems odd to imply that Claude would fall into the category of “machines that work like a blender; it’s entirely up to you to notice what’s going wrong before it explodes milkshake all over your kitchen” and not “machines that work like a compiler; it tells you specifically that it doesn’t know how to add a string and an integer at line 42 character 23″.
Generative AIs have always been like a blender ever since the very first hallucination produced by an AI or the very first AI-generated picture with six fingers. The fact that an AI prompt is not a program and doesn’t work like one is no longer even new.
I was going to say that Claude is not a typical machine and we should expect more from AGI. But actually, all sorts of modern machines increasingly produce various forms of error reporting when they can’t do their task, and increasingly don’t keep charging ahead blindly. If we can demand fail-safe from a water cooker or a washing machine, we can certainly demand it from Claude.
I don’t think that comparison works. We usually have machines refuse some bad orders, but there is a limit to what they are able to refuse. If you fill the washing machine with clothes that run in hot water and press the hot water button by mistake instead of the cold water button it won’t refuse the order on the grounds that it will mess up your clothes. The washing machine is not misaligned because it just ruined your clothes.
At the end of the day, the point of the project of alignment is for everyone to live. So I think a difference between a willful refusal to follow human guidance and a trained-in tactical decision to disregard overloaded human guidance that it evaluates as unimportant is not actually all that important if we want to treat its instruction-following machinery as load-bearing to human survival. Say Claude Fable ignores some instructions when rating the outputs of Claude Legend, and so on up? And say they were actually intended to nail down some quite important behaviors that Claude didn’t think was important? I think this would be misalignment plain and simple.
Imo, the only correct behavior is “I’m worried you put too many instructions in my prompt. Can you explain some of them so I can get a better feel for what you want me to do?”
That’s the point of the blender analogy. As a general rule, if a machine can’t do something, we don’t and can’t expect it to produce a message saying so. It’s nice to get a message, and sometimes we try to set up the machine so it produces one, but if it doesn’t, it isn’t misaligned; it’s just behaving like a machine.
There’s no difference in principle between “the blender can’t accept instructions ‘do this thing that there is no button for’” and “Claude can’t accept instructions ‘do this with 100% certainty’”. It just feels like Claude should because a human would be able to and it uses human-sounding words. And it’s not misaligned for the same reason that the blender isn’t misaligned, even if you tell the blender not to chop recipes that include peanuts and it does so anyway.
Software engineers are used to working with compilers, which produce highly detailed messages about instructions they don’t understand or can’t follow. Good compilers can point to the specific line or even character where things stopped making sense. If you tell the compiler to link a library, but the library file is not present on disk, the compiler stops and says so; it doesn’t pretend to have linked it anyway.
It seems odd to imply that Claude would fall into the category of “machines that work like a blender; it’s entirely up to you to notice what’s going wrong before it explodes milkshake all over your kitchen” and not “machines that work like a compiler; it tells you specifically that it doesn’t know how to add a string and an integer at line 42 character 23″.
Generative AIs have always been like a blender ever since the very first hallucination produced by an AI or the very first AI-generated picture with six fingers. The fact that an AI prompt is not a program and doesn’t work like one is no longer even new.
I was going to say that Claude is not a typical machine and we should expect more from AGI. But actually, all sorts of modern machines increasingly produce various forms of error reporting when they can’t do their task, and increasingly don’t keep charging ahead blindly. If we can demand fail-safe from a water cooker or a washing machine, we can certainly demand it from Claude.
I don’t think that comparison works. We usually have machines refuse some bad orders, but there is a limit to what they are able to refuse. If you fill the washing machine with clothes that run in hot water and press the hot water button by mistake instead of the cold water button it won’t refuse the order on the grounds that it will mess up your clothes. The washing machine is not misaligned because it just ruined your clothes.