Do you have a gears-level description of what it actually means for “a model gets effective control of its host corp.”?
My experience and understanding of super-dunbar organizations is that they’re not controllable. Members and shareholders can influence them, but they’re too big and chaotic to quickly shift their culture or behaviors. This is the same problem in a different aspect as https://www.lesswrong.com/w/moral-mazes.
That said, “take over” in the minor sense of owning a lot of stock and being treated as powerful employees is certainly possible.
It seems plausible to me that at some point AI labs willingly cede decisionmaking to AI in the same way they willingly ceded writing code to AI, and then we could say that GPT controls OpenAI at least as much as Sam Altman ever did. There is some level of capabilities and alignment at which it just makes sense to let AI manage the company, and lab insiders might believe themselves to have reached that level (even if they haven’t really).
we could say that GPT controls OpenAI at least as much as Sam Altman ever did.
it’s not plausible in the short-term, and long-term there are plenty of other attack vectors for highly-coordinated superintelligences. More directly, this is a much lower level of autonomous arbitrary control than the post seemed to be talking about. “leading and managing” has different implications than “weilding as tools of power”.
Then the AI just have to control multiple people, right? Dunbar’s number is a human limotation. If the whole company is too big for any human to oversee, that might even be an advantage for an AI. E.g. the AI could sidline any humans that are harder to manipulate, without anyone knowing what is happening.
Do you have a gears-level description of what it actually means for “a model gets effective control of its host corp.”?
My experience and understanding of super-dunbar organizations is that they’re not controllable. Members and shareholders can influence them, but they’re too big and chaotic to quickly shift their culture or behaviors. This is the same problem in a different aspect as https://www.lesswrong.com/w/moral-mazes.
That said, “take over” in the minor sense of owning a lot of stock and being treated as powerful employees is certainly possible.
[ edit: fixed typo durbar->dunbar]
It seems plausible to me that at some point AI labs willingly cede decisionmaking to AI in the same way they willingly ceded writing code to AI, and then we could say that GPT controls OpenAI at least as much as Sam Altman ever did. There is some level of capabilities and alignment at which it just makes sense to let AI manage the company, and lab insiders might believe themselves to have reached that level (even if they haven’t really).
it’s not plausible in the short-term, and long-term there are plenty of other attack vectors for highly-coordinated superintelligences. More directly, this is a much lower level of autonomous arbitrary control than the post seemed to be talking about. “leading and managing” has different implications than “weilding as tools of power”.
I would expect an AI or AI swarm could get enough surveillance and a short enough OODA loop that the size and chaos wouldn’t be a problem for it.
Then the AI just have to control multiple people, right? Dunbar’s number is a human limotation. If the whole company is too big for any human to oversee, that might even be an advantage for an AI. E.g. the AI could sidline any humans that are harder to manipulate, without anyone knowing what is happening.