I’ve been following AI models, and specifically AI safety, since ChatGPT started entering the public consciousness. What caught my attention about ChatGPT was how remarkably it could hold a conversation with me—that was something I genuinely never imagined a machine could do. I had always assumed that language was something unique and special about humans, and lo, here was a machine doing it well!
What struck me, and continues to strike me, is how easy it is to just talk to frontier LLMs as if they’re a person. It feels natural to refer to LLMs in the second person, to ask their opinion, to wonder what it’s trying to do or what it wants. And reading the research on AI, especially AI safety, I see a lot of that language. I read, routinely: “what a model thinks”, “it knows it’s being tested”, “it wants”. AI models’ output seems to make this language intuitive and useful, but do LLMs, as they are now, earn that language?
Prior to the existence of LLMs, it was true that if something could speak and reason, it could be held to account for that speech and that reasoning. That’s why we ask people to give their reasons for their actions, especially in cases where those actions created or destroyed value: throughout history, being responsible for an action meant being able to provide an account of why the actor behaved as such. The correlation between “able to reason” and “able to be held responsible” was never even visible as a correlation: there was nothing to check the correlation against.
But now there is: LLMs are a case where “able to reason” and “able to be held responsible” could come apart. LLMs can reason: whether LLMs are able to be held responsible has yet to be checked, precisely because the surface against which responsibility was checked—the ability to give reasons—is the surface LLMs reproduce.
AI researchers’ own language—”a model thinks”, “it knows”, “it wants”—maps on to models’ behavior just fine; the prediction “the model wants to avoid the filter” does tell you what the model will do. It also smuggles in an answerable someone: a party who could be held responsible. However, simply predicting the wanting does not predict the wanter; the behavior is the same regardless of whether there is any actual wanter. The language is verified on the model’s behavior, but it brings with it an unverified actor responsible for the behavior, and AI safety relies on that unverified feature.
Prior to the development of LLMs, the inference from “able to reason” to “able to be held responsible” was safe. Testing whether something could be held responsible meant testing whether it had the capacity to reason. LLMs have the capacity to reason. The question is: does reasoning imply responsibility? Is that even being tested?
Does AI reasoning imply responsibility?
I’ve been following AI models, and specifically AI safety, since ChatGPT started entering the public consciousness. What caught my attention about ChatGPT was how remarkably it could hold a conversation with me—that was something I genuinely never imagined a machine could do. I had always assumed that language was something unique and special about humans, and lo, here was a machine doing it well!
What struck me, and continues to strike me, is how easy it is to just talk to frontier LLMs as if they’re a person. It feels natural to refer to LLMs in the second person, to ask their opinion, to wonder what it’s trying to do or what it wants. And reading the research on AI, especially AI safety, I see a lot of that language. I read, routinely: “what a model thinks”, “it knows it’s being tested”, “it wants”. AI models’ output seems to make this language intuitive and useful, but do LLMs, as they are now, earn that language?
Prior to the existence of LLMs, it was true that if something could speak and reason, it could be held to account for that speech and that reasoning. That’s why we ask people to give their reasons for their actions, especially in cases where those actions created or destroyed value: throughout history, being responsible for an action meant being able to provide an account of why the actor behaved as such. The correlation between “able to reason” and “able to be held responsible” was never even visible as a correlation: there was nothing to check the correlation against.
But now there is: LLMs are a case where “able to reason” and “able to be held responsible” could come apart. LLMs can reason: whether LLMs are able to be held responsible has yet to be checked, precisely because the surface against which responsibility was checked—the ability to give reasons—is the surface LLMs reproduce.
AI researchers’ own language—”a model thinks”, “it knows”, “it wants”—maps on to models’ behavior just fine; the prediction “the model wants to avoid the filter” does tell you what the model will do. It also smuggles in an answerable someone: a party who could be held responsible. However, simply predicting the wanting does not predict the wanter; the behavior is the same regardless of whether there is any actual wanter. The language is verified on the model’s behavior, but it brings with it an unverified actor responsible for the behavior, and AI safety relies on that unverified feature.
Prior to the development of LLMs, the inference from “able to reason” to “able to be held responsible” was safe. Testing whether something could be held responsible meant testing whether it had the capacity to reason. LLMs have the capacity to reason. The question is: does reasoning imply responsibility? Is that even being tested?