Anthropic is the safety-first frontier AI lab, surely they’re doing at least that much.
They’re almost certainly not doing it for Mythos? This result is from reviewing what had already happening in the training run. It’ll be way too expensive to patch everything and retrain the model.
Sorry that was sarcasm. Poe’s law, I know.
Anthropic is the safety-first frontier AI lab, surely they’re doing at least that much.
They’re almost certainly not doing it for Mythos? This result is from reviewing what had already happening in the training run. It’ll be way too expensive to patch everything and retrain the model.
Sorry that was sarcasm. Poe’s law, I know.