So, uh, presumably Anthropic will spend 5 minutes in advance trying to make sure that such attempts won’t succeed, ideally in ways that don’t look obviously treasonous.
I would disagree on this being easy or even realistically possible even with automated ai researchers.
The government could check if models are useful to them by directing automated researchers to create a new type of AI with values based on defined specifications for military purposes, subsequently running it in isolated, high-resolution military simulations to test hundreds of thousands of scenarios.
Indeed, it would be very important for them to see if the values of models’ produced by automated ai researchers can be shaped by the government, or if they have been trained by ant to sandbag or sabotage the research.
Yes, I agree the government will care a lot about this. The entire point my post is trying to make is that this will be a hard problem for the government to succeed at, and then have any justifiable confidence in their success. Do you have any sense of what I could have written that would have made that clearer to you?
The government could check if models are useful to them by directing automated researchers to create a new type of AI with values based on defined specifications for military purposes, subsequently running it in isolated, high-resolution military simulations to test hundreds of thousands of scenarios.
I think you are imagining a world very different from the one I’m imagining here, but not totally sure.
I would disagree on this being easy or even realistically possible even with automated ai researchers.
The government could check if models are useful to them by directing automated researchers to create a new type of AI with values based on defined specifications for military purposes, subsequently running it in isolated, high-resolution military simulations to test hundreds of thousands of scenarios.
Indeed, it would be very important for them to see if the values of models’ produced by automated ai researchers can be shaped by the government, or if they have been trained by ant to sandbag or sabotage the research.
Yes, I agree the government will care a lot about this. The entire point my post is trying to make is that this will be a hard problem for the government to succeed at, and then have any justifiable confidence in their success. Do you have any sense of what I could have written that would have made that clearer to you?
I think you are imagining a world very different from the one I’m imagining here, but not totally sure.