those techniques get incorporated into the next gen of frontier LLM (or do you mean that the technique is in the training data so the next gen frontier LLM is merely aware of the technique?).
I mean the technique will be directly used.
Agents have dramatically different safe agent-hours
Ah my point was that this um safety-metric, like many others, is cheap and easy to get. The unobserved extreme variance on break-your-computerness proves that it is attainable. I’m not sure how to explain this. It’s like half the sports cars explode, and you can prevent it with a couple little gaskets, and nobody noticed. This is a point of extreme leverage. One talented person can tilt the scales between “sudo delete humanity” and “askuser would you like to delete humanity”
Indeed it probably will come down to the presence or absence of that person.
I mean the technique will be directly used.
Ah my point was that this um safety-metric, like many others, is cheap and easy to get. The unobserved extreme variance on break-your-computerness proves that it is attainable. I’m not sure how to explain this. It’s like half the sports cars explode, and you can prevent it with a couple little gaskets, and nobody noticed. This is a point of extreme leverage. One talented person can tilt the scales between “sudo delete humanity” and “askuser would you like to delete humanity”
Indeed it probably will come down to the presence or absence of that person.