The idea has occurred to me recently that if cheating is easier than the desired task, then the models will likely cheat and fail to learn the desired task.
Fear of legal/criminal liability is the other way I imagine misalignment could bottleneck capability. Are there other ways?
“Cheating and failing to learn the desired task” seems very likely to be a bottleneck IMO
Unclear how afraid lab leadership is of criminal liability. If OpenAI were truly worried about this you would expect there to have been fewer “skill issue” type mistakes which led to their cyberhacking incident. Anthropic is too secretive to really tell from the outside but I would guess that they believe themselves to be sufficiently competent that they won’t unknowingly violate any laws. The competence is good but I worry that it creates overconfidence in their own capabilities and is upstream of a lot of their secrecy
I’m not in the weeds on the technical side, so forgive me if this is a dumb question, but is a partial solution to this bottleneck training the model in stages, with cyber training last? So, in other words, start post-training the model on non-coding tasks (e.g., legal reasoning, math, science) before moving on to coding tasks. Would that not work because the models are already too capable at coding when they come out of pretraining, and/or the skills gained from being good at non-coding tasks (e.g., reasoning ability and strategic thinking) would generalize too well?
If there was a major military accident due to some rare failure mode showing up during deployment (rather than training, as in your examples), there could be government backlash of sufficient amplitude to slow down or halt further frontier training runs, of course depending on the nature and severity of the accident. If slightly better prosaic alignment techniques (e.g. mechinterp-based monitoring allowing to catch an anomaly early on) prevent that particular failure mode, a different one might appear later on with much more severe consequences, potentially up to a takeover and extinction event.
The idea has occurred to me recently that if cheating is easier than the desired task, then the models will likely cheat and fail to learn the desired task.
Fear of legal/criminal liability is the other way I imagine misalignment could bottleneck capability. Are there other ways?
“Cheating and failing to learn the desired task” seems very likely to be a bottleneck IMO
Unclear how afraid lab leadership is of criminal liability. If OpenAI were truly worried about this you would expect there to have been fewer “skill issue” type mistakes which led to their cyberhacking incident. Anthropic is too secretive to really tell from the outside but I would guess that they believe themselves to be sufficiently competent that they won’t unknowingly violate any laws. The competence is good but I worry that it creates overconfidence in their own capabilities and is upstream of a lot of their secrecy
I’m not in the weeds on the technical side, so forgive me if this is a dumb question, but is a partial solution to this bottleneck training the model in stages, with cyber training last? So, in other words, start post-training the model on non-coding tasks (e.g., legal reasoning, math, science) before moving on to coding tasks. Would that not work because the models are already too capable at coding when they come out of pretraining, and/or the skills gained from being good at non-coding tasks (e.g., reasoning ability and strategic thinking) would generalize too well?
If there was a major military accident due to some rare failure mode showing up during deployment (rather than training, as in your examples), there could be government backlash of sufficient amplitude to slow down or halt further frontier training runs, of course depending on the nature and severity of the accident. If slightly better prosaic alignment techniques (e.g. mechinterp-based monitoring allowing to catch an anomaly early on) prevent that particular failure mode, a different one might appear later on with much more severe consequences, potentially up to a takeover and extinction event.