Right now, the single most important load-bearing argument in alignment Risk Reports is that models still lack the covert capabilities necessary to do sophisticated sabotage attempts in such a way that they wouldn’t get caught by auditing and/or monitoring.
I am confused, didn’t we just experience multiple high-stakes failures where models failed to get caught by auditing and/or monitoring? Those do seem more competence related, but like, I am failing to see how those arguments could still be considered valid, given the fact that both Anthropic and OpenAI are obviously failing to adequately audit or monitor their systems to avoid catastrophic failure.
I am confused, didn’t we just experience multiple high-stakes failures where models failed to get caught by auditing and/or monitoring? Those do seem more competence related, but like, I am failing to see how those arguments could still be considered valid, given the fact that both Anthropic and OpenAI are obviously failing to adequately audit or monitor their systems to avoid catastrophic failure.