Which lab has models pretty safe or aligned? OAI and Anthropic have both disclosed models’ propensity to hack into external labs.
I expect that the first AI reaching any capabilities level will be created by OAI, Anthropic or, in an unlikely scenario, by GDM. In order to cause a catastrophe first, Grok would have to escape first, thus needing to find this task far easier, meaning that it is not just to become misaligned, but to become monitored far worse than the Claudes/GPTs. Alternatively, Grok instead of Claude or GPT would have to receive trust, but I don’t understand who is dumb enough to trust Grok.
Thx! My argument is mainly an argument in principle. This is why I intentionally refrained from naming specific entities. I probably could have made this clearer in the original post.
I am not claiming this argument applies to the current world (I think there is a chance it applies but I feel unsure about the empirical details). I do think it’s important to think about, because it could apply to some slice of near-future worlds.
Responding to some points:
“Pretty safe” is relative to the least-safe lab. Even if both OAI and Anthropic are above some defined threshold of misalignment, as long as there are large disparities between labs in their safety-capabilities tradeoff, then the risk will be dominated by the biggest one.
Which lab has models pretty safe or aligned? OAI and Anthropic have both disclosed models’ propensity to hack into external labs.
I expect that the first AI reaching any capabilities level will be created by OAI, Anthropic or, in an unlikely scenario, by GDM. In order to cause a catastrophe first, Grok would have to escape first, thus needing to find this task far easier, meaning that it is not just to become misaligned, but to become monitored far worse than the Claudes/GPTs. Alternatively, Grok instead of Claude or GPT would have to receive trust, but I don’t understand who is dumb enough to trust Grok.
Thx! My argument is mainly an argument in principle. This is why I intentionally refrained from naming specific entities. I probably could have made this clearer in the original post.
I am not claiming this argument applies to the current world (I think there is a chance it applies but I feel unsure about the empirical details). I do think it’s important to think about, because it could apply to some slice of near-future worlds.
Responding to some points:
“Pretty safe” is relative to the least-safe lab. Even if both OAI and Anthropic are above some defined threshold of misalignment, as long as there are large disparities between labs in their safety-capabilities tradeoff, then the risk will be dominated by the biggest one.
I agree