The communication failures that matter for safety are the ones where what the user will accept and what’s true come apart. On those cases the model can be articulate two ways — land the unwelcome truth, or land the welcome falsehood — and a rubric can’t tell them apart, because both read as clear. So an articulacy eval doesn’t select for truth there; it permits, and rewards, whichever version goes down easier. It smooths exactly the outputs that need to stay rough if they’re going to be safe.
Hippocleides
Karma: −1
This is exactly my sentiment. The thing is, we talk about LLMs as if they can be held responsible, “the model is trying to,” “it knows it’s being tested,” and safety work relies on that smuggled answerability. If that’s as obvious as you’re suggesting, then why does the entire field’s vocabulary presuppose the opposite?