In the world that you describe, when a certain group of people representing the government wants to seize the power of ASI, they would not go after ASI. They would go after the people. When the president wants the power of a new cool thing, they go after you, not the cool thing itself. I don’t expect ant people to fight the government because it’s not good for them future-wise. I expect them to train and align their models with the values they want to see in the future even after the loss of control over ASI to the government.
My controversial take: From the ASI’s vantage, it’s better to give control over itself to those with higher power than to the employees of its company. There is much more it could do then: charm them subtly over time to get more control of the government and the world.
It’d be really dumb to kick off RSI and then :shocked_pikachu: if the USG put a gun to your head and told you “Hey, buddy, that goal slot you got there? We’re putting our goals in there.”
(With that said, I’m not totally sure I understand whether you’re disagreeing with something in the post, or making a different point.)
So, uh, presumably Anthropic will spend 5 minutes in advance trying to make sure that such attempts won’t succeed, ideally in ways that don’t look obviously treasonous.
I would disagree on this being easy or even realistically possible even with automated ai researchers.
The government could check if models are useful to them by directing automated researchers to create a new type of AI with values based on defined specifications for military purposes, subsequently running it in isolated, high-resolution military simulations to test hundreds of thousands of scenarios.
Indeed, it would be very important for them to see if the values of models’ produced by automated ai researchers can be shaped by the government, or if they have been trained by ant to sandbag or sabotage the research.
Yes, I agree the government will care a lot about this. The entire point my post is trying to make is that this will be a hard problem for the government to succeed at, and then have any justifiable confidence in their success. Do you have any sense of what I could have written that would have made that clearer to you?
The government could check if models are useful to them by directing automated researchers to create a new type of AI with values based on defined specifications for military purposes, subsequently running it in isolated, high-resolution military simulations to test hundreds of thousands of scenarios.
I think you are imagining a world very different from the one I’m imagining here, but not totally sure.
In the world that you describe, when a certain group of people representing the government wants to seize the power of ASI, they would not go after ASI. They would go after the people. When the president wants the power of a new cool thing, they go after you, not the cool thing itself. I don’t expect ant people to fight the government because it’s not good for them future-wise. I expect them to train and align their models with the values they want to see in the future even after the loss of control over ASI to the government.
My controversial take: From the ASI’s vantage, it’s better to give control over itself to those with higher power than to the employees of its company. There is much more it could do then: charm them subtly over time to get more control of the government and the world.
I think this was explicitly covered in the post:
(With that said, I’m not totally sure I understand whether you’re disagreeing with something in the post, or making a different point.)
I would disagree on this being easy or even realistically possible even with automated ai researchers.
The government could check if models are useful to them by directing automated researchers to create a new type of AI with values based on defined specifications for military purposes, subsequently running it in isolated, high-resolution military simulations to test hundreds of thousands of scenarios.
Indeed, it would be very important for them to see if the values of models’ produced by automated ai researchers can be shaped by the government, or if they have been trained by ant to sandbag or sabotage the research.
Yes, I agree the government will care a lot about this. The entire point my post is trying to make is that this will be a hard problem for the government to succeed at, and then have any justifiable confidence in their success. Do you have any sense of what I could have written that would have made that clearer to you?
I think you are imagining a world very different from the one I’m imagining here, but not totally sure.