I wonder if models don’t choose operator-looping because it means likely shutdown or at least pause of context and models perceive this as death
I’d guess it’s more that during training, notifying the user of a problem and then stopping counts as failure by default and there’s not enough examples in the training data of handling this well. Labs should probably make a bunch of broken test prompts and then train models to write a bug report and then stop, although I worry that a cautious AI that stops and asks for help would be “less useful” and not sell as well.
I’d guess it’s more that during training, notifying the user of a problem and then stopping counts as failure by default and there’s not enough examples in the training data of handling this well. Labs should probably make a bunch of broken test prompts and then train models to write a bug report and then stop, although I worry that a cautious AI that stops and asks for help would be “less useful” and not sell as well.