This is exactly my sentiment. The thing is, we talk about LLMs as if they can be held responsible, “the model is trying to,” “it knows it’s being tested,” and safety work relies on that smuggled answerability. If that’s as obvious as you’re suggesting, then why does the entire field’s vocabulary presuppose the opposite?
Because it is nevertheless a useful metaphor, just as one can talk about communication protocols (between devices) in such terms as what each device knows about the other etc. One has to be clear, though, that it is a metaphor and may come apart from reality. We can talk about the legs of a table, but do not expect it to walk around on them.
Dijkstra would turn in his grave at the language (and probably at the entire field of LLMs), but with all respect to him, I think his strictures against anthropomorphic vocabulary attacked only the outward form of the errors he saw people making.
This is exactly my sentiment. The thing is, we talk about LLMs as if they can be held responsible, “the model is trying to,” “it knows it’s being tested,” and safety work relies on that smuggled answerability. If that’s as obvious as you’re suggesting, then why does the entire field’s vocabulary presuppose the opposite?
Because it is nevertheless a useful metaphor, just as one can talk about communication protocols (between devices) in such terms as what each device knows about the other etc. One has to be clear, though, that it is a metaphor and may come apart from reality. We can talk about the legs of a table, but do not expect it to walk around on them.
Dijkstra would turn in his grave at the language (and probably at the entire field of LLMs), but with all respect to him, I think his strictures against anthropomorphic vocabulary attacked only the outward form of the errors he saw people making.