The communication failures that matter for safety are the ones where what the user will accept and what’s true come apart. On those cases the model can be articulate two ways — land the unwelcome truth, or land the welcome falsehood — and a rubric can’t tell them apart, because both read as clear. So an articulacy eval doesn’t select for truth there; it permits, and rewards, whichever version goes down easier. It smooths exactly the outputs that need to stay rough if they’re going to be safe.
The communication failures that matter for safety are the ones where what the user will accept and what’s true come apart. On those cases the model can be articulate two ways — land the unwelcome truth, or land the welcome falsehood — and a rubric can’t tell them apart, because both read as clear. So an articulacy eval doesn’t select for truth there; it permits, and rewards, whichever version goes down easier. It smooths exactly the outputs that need to stay rough if they’re going to be safe.