https://twitter.com/DavidSKrueger
https://www.davidscottkrueger.com/
https://therealartificialintelligence.substack.com/p/the-real-ai-deploys-itself
David Scott Krueger
It’s a high level point I’m making that you still seem to be missing or misunderstanding. The onus should be on them to build and present such a case.
But anyways, there’s an obvious question of whether these traces and logs are sufficient to draw such a conclusion, and if there’s any way they could’ve been tampered with.
Also, OpenAI seems to be relying heavily on AI to assist with their investigations so far, and this seems like it could call findings into question.
How would such segregation work?
I don’t have much CS background outside of ML and don’t know computer security well, but it sounds to me from ~25:30 of the BlackHat talk from 2 days ago like the model got root access to the machine on which it was running. And they explicitly say the model moved laterally (this was known in public reporting at the time of my OP).
It may lead to AI takeover by making it impossible to pause development.
As I said:
AIs have tried to “exfiltrate” themselves (i.e. their “weights”) in previous experiments many times. It’s a natural and obvious question to ask.
Are you assuming that OpenAI would know and would disclose if they had? Why do you trust them so much, given that they didn’t even know the model had escaped the sandbox for days and their CEO has a known honesty problem?
You’re eggregiously missing my points:
1) It doesn’t matter if we don’t think it’s likely. It needs to be demonstrated to a very high confidence level. This is an entirely reasonable ask.
2) I already said there were calls for transparency; the one you reference did not make this demand.
IIUC, reward-seeking is not goal misgeneralization, it’s reward hacking / outer-misalignment?
(It would also be nice to include a definition of reward-seeking in the post).
though importantly it might be actively more fit than intent alignment, so it’s also about outer alignment
What do you mean by this?
Also, not to pick on you, but I generally think it’s useful to “related work” type things, not just for credit assignment, but also to facilitate understanding. e.g. as someone who is very familiar with alignment, if I can just view “fitness maximizing” as a “goal misgeneralization, but with XYZ” then maybe I can get any key insights in a few minutes.
I haven’t read the post, just a bit past the definition, but is what you are talking about just competent/agentic goal misgeneralization? Or can you explain how it relates?
Additionally, it is unclear to me whether these sorts of failure modes are fundamentally due to alien concepts, or whether they are better explained by differences in sensory processing architectures. The latter is potentially tractable and seems less likely to lead to misbehavior in the limit, while the former would represent a more severe and potentially intractable mismatch in ontologies.
This is a good point, and I’m not aware of research asking this question as phrased.
In a word: InkHaven.
But seriously, I’m still working full-time on Evitable.com and so am trying to churn out my daily blog posts FAST. There are topics I know I have things to say about, and I try to get them down in words in ~1-2 hours tops. In this case, the motivation is something like: “It’s annoying when people make behaviorist arguments about how AIs are more aligned/trustworthy than people”.
Well, I did say “Naively”… but yes I agree the analysis was too naive, and I will edit the post. You make a good point that it can be improved by considering that harms from AI (especially large-scale ones like x-risk) are overdetermined when there are multiple developers. The naive analysis is more accurate when the risk is smaller.
As a side note, if the risk from a single project is so large, then the first project is probably disincentivized at the individual level (would you really want to take an 80% risk of extinction?), and it’s a “pure” coordination problem, like a stag hunt, rather than an incentive problem (like prisoner’s dilema).
Another way the “naive” calculation can be is wrong (which is the main one I had in mind) is if the risks of different projects are correlated, which they are, e.g. because they are all using similar technology.
I’m not going after particular people’s justifications for their work; I’m going after the institutionalization of “marginal risk” as a relevant concept and the way it justifies unacceptable risk-taking.
I think it’s helpful to disaggregate things sometimes, and e.g. look at what trends might underly this general trend we observe
greater wealth hasn’t changed the picture tremendously
I don’t think I made that claim anywhere in my piece.
I think we disagree on the likelihood of x-risk advocates having such a precise level of impact and power. If AI x-risk advocates don’t have much sway over the political demands of an anti-AI movement, then I don’t think we have much to worry about.
Your argument seems to treat “pause/stop AI” as almost like an info-hazard, but it’s really quite an obvious idea. If “pause/stop AI” becomes a core idea in an anti-AI movement where x-risk advocates lack power, I’d expect that happens because it’s memetically fit and other people were saying it as well, so it probably would’ve happened anyways.
In my mind, the world’s where there is political will for a pause are mostly the ones where there is a broad understanding that if we don’t stop building AI, it is going to replace humanity, and this motivates the need for an international pause. Similarly, I think a pause that happens without the USG having been AGI-pilled seems really unlikely, and it’s also very hard for me to imagine the current administration doing a unilateral pause.
Overall, I think if “pause/stop AI” succeeds at all, it will probably succeed substantively, the slice of worlds where this doesn’t happen seem very narrow, because it’s a big ask.
I don’t mean to put regulation and stopping in opposition. My point is that, stopping is likely a precondition for any form of regulation that would significantly slow down development or deployment. Like, you, I am trying to argue against framings that put
“we need a global treaty to stop AI risks” in opposition to “domestic regulation is the only realistic path.”
I think stopping unlocks a lot of ability for countries to regulate in line with their values and priorities that otherwise might not be possible because of race dynamics.
I’ve tried to edit my post to make that clearer, please let me know if you have any specific suggestions on that front.
Yeah, this is a good point. The way I’ve put it before is: when you are thinking about what should happen, you’re basically imagining you have some sort of magic wand that makes it happen. But how powerful is the magic wand? I haven’t thought this through to my satisfaction, so for now I’m just going based on intuitive notions of what is actually realistically achievable.
But one way of trying to define the limits of the “magic wand” here would be: You get to magically choose a policy to be adopted, but you don’t get to magically control people’s behavior afterwards. So if you want to get people to limit AI uses, your policy needs to deal with their potential incentives to do otherwise.
This means, IIUC, that the answer to your final question is “yes”. But it’s more a matter of perceived incentives here, IMO, see: https://therealartificialintelligence.substack.com/p/following-the-incentives
> If someone believes that it will be hard to make international agreements to stop AI because countries will have incentives against this, does that mean that those considerations now fall under “incentives” and thus count for purpose of determining whether stopping is “hard”?
There’s not a lot of demand for human cloning. See https://wiki.aiimpacts.org/doku.php?id=responses_to_ai:technological_inevitability:incentivized_technologies_not_pursued:start
Good point RE deskilling of alignment researchers.
Yes, but if there is rogue AI already on the loose, we can’t stop it from doing RSI and/or taking over unless we can track down or contain all the copies.