I think your wording of “top n% of tasks” muddies your argument somewhat. Top n% of what? Risk? A very large portion of AI safety, at least at inference time, already consists of running classifiers against every prompt and hidden layers for possible dangerous usage. Top n% of capability (ie model class)? Nobody’s running safety checks against local 4b models as they simply don’t have the “smarts” to pull off anything actually effective. Haiku’s already not getting nearly as many controls or attention as Mythos does. I’m not sure what your point is here.
Top n% of tasks means things like “It’s really important to be robust when dealing with your production database, and even more if it’s a production database representing US nuclear weapons programs.”
You’re right that there are classifiers for obviously risky things, but these mainly refuse access to directly hazardous queries (“How can I create a bomb”).
I think your wording of “top n% of tasks” muddies your argument somewhat. Top n% of what? Risk? A very large portion of AI safety, at least at inference time, already consists of running classifiers against every prompt and hidden layers for possible dangerous usage. Top n% of capability (ie model class)? Nobody’s running safety checks against local 4b models as they simply don’t have the “smarts” to pull off anything actually effective. Haiku’s already not getting nearly as many controls or attention as Mythos does. I’m not sure what your point is here.
Top n% of tasks means things like “It’s really important to be robust when dealing with your production database, and even more if it’s a production database representing US nuclear weapons programs.”
You’re right that there are classifiers for obviously risky things, but these mainly refuse access to directly hazardous queries (“How can I create a bomb”).