Does AI reasoning imply responsibility?
I’ve been following AI models, and specifically AI safety, since ChatGPT started entering the public consciousness. What caught my attention about ChatGPT was how remarkably it could hold a conversation with me—that was something I genuinely never imagined a machine could do. I had always assumed that language was something unique and special about humans, and lo, here was a machine doing it well!
What struck me, and continues to strike me, is how easy it is to just talk to frontier LLMs as if they’re a person. It feels natural to refer to LLMs in the second person, to ask their opinion, to wonder what it’s trying to do or what it wants. And reading the research on AI, especially AI safety, I see a lot of that language. I read, routinely: “what a model thinks”, “it knows it’s being tested”, “it wants”. AI models’ output seems to make this language intuitive and useful, but do LLMs, as they are now, earn that language?
Prior to the existence of LLMs, it was true that if something could speak and reason, it could be held to account for that speech and that reasoning. That’s why we ask people to give their reasons for their actions, especially in cases where those actions created or destroyed value: throughout history, being responsible for an action meant being able to provide an account of why the actor behaved as such. The correlation between “able to reason” and “able to be held responsible” was never even visible as a correlation: there was nothing to check the correlation against.
But now there is: LLMs are a case where “able to reason” and “able to be held responsible” could come apart. LLMs can reason: whether LLMs are able to be held responsible has yet to be checked, precisely because the surface against which responsibility was checked—the ability to give reasons—is the surface LLMs reproduce.
AI researchers’ own language—”a model thinks”, “it knows”, “it wants”—maps on to models’ behavior just fine; the prediction “the model wants to avoid the filter” does tell you what the model will do. It also smuggles in an answerable someone: a party who could be held responsible. However, simply predicting the wanting does not predict the wanter; the behavior is the same regardless of whether there is any actual wanter. The language is verified on the model’s behavior, but it brings with it an unverified actor responsible for the behavior, and AI safety relies on that unverified feature.
Prior to the development of LLMs, the inference from “able to reason” to “able to be held responsible” was safe. Testing whether something could be held responsible meant testing whether it had the capacity to reason. LLMs have the capacity to reason. The question is: does reasoning imply responsibility? Is that even being tested?
LLMs cannot be held responsible for anything. They cannot make amends for what they do, they cannot have penalties levied upon them, they cannot be compelled to do anything, their testimony cannot be relied on, and more fundamentally, they have no ongoing existence as a coherent thing. If we compare an LLM to a person, an LLM is like a person who is grossly and congenitally incompetent to act in any legal proceedings.
This is exactly my sentiment. The thing is, we talk about LLMs as if they can be held responsible, “the model is trying to,” “it knows it’s being tested,” and safety work relies on that smuggled answerability. If that’s as obvious as you’re suggesting, then why does the entire field’s vocabulary presuppose the opposite?
Because it is nevertheless a useful metaphor, just as one can talk about communication protocols (between devices) in such terms as what each device knows about the other etc. One has to be clear, though, that it is a metaphor and may come apart from reality. We can talk about the legs of a table, but do not expect it to walk around on them.
Dijkstra would turn in his grave at the language (and probably at the entire field of LLMs), but with all respect to him, I think his strictures against anthropomorphic vocabulary attacked only the outward form of the errors he saw people making.