https://twitter.com/DavidSKrueger
https://www.davidscottkrueger.com/
https://therealartificialintelligence.substack.com/p/the-real-ai-deploys-itself
David Scott Krueger
Warning Shots: A Theory
It’s a high level point I’m making that you still seem to be missing or misunderstanding. The onus should be on them to build and present such a case.
But anyways, there’s an obvious question of whether these traces and logs are sufficient to draw such a conclusion, and if there’s any way they could’ve been tampered with.
Also, OpenAI seems to be relying heavily on AI to assist with their investigations so far, and this seems like it could call findings into question.
How would such segregation work?
I don’t have much CS background outside of ML and don’t know computer security well, but it sounds to me from ~25:30 of the BlackHat talk from 2 days ago like the model got root access to the machine on which it was running. And they explicitly say the model moved laterally (this was known in public reporting at the time of my OP).
It may lead to AI takeover by making it impossible to pause development.
As I said:
AIs have tried to “exfiltrate” themselves (i.e. their “weights”) in previous experiments many times. It’s a natural and obvious question to ask.
Are you assuming that OpenAI would know and would disclose if they had? Why do you trust them so much, given that they didn’t even know the model had escaped the sandbox for days and their CEO has a known honesty problem?
You’re eggregiously missing my points:
1) It doesn’t matter if we don’t think it’s likely. It needs to be demonstrated to a very high confidence level. This is an entirely reasonable ask.
2) I already said there were calls for transparency; the one you reference did not make this demand.
…but have the weights left the server?
IIUC, reward-seeking is not goal misgeneralization, it’s reward hacking / outer-misalignment?
(It would also be nice to include a definition of reward-seeking in the post).
though importantly it might be actively more fit than intent alignment, so it’s also about outer alignment
What do you mean by this?
Also, not to pick on you, but I generally think it’s useful to “related work” type things, not just for credit assignment, but also to facilitate understanding. e.g. as someone who is very familiar with alignment, if I can just view “fitness maximizing” as a “goal misgeneralization, but with XYZ” then maybe I can get any key insights in a few minutes.
I haven’t read the post, just a bit past the definition, but is what you are talking about just competent/agentic goal misgeneralization? Or can you explain how it relates?
Additionally, it is unclear to me whether these sorts of failure modes are fundamentally due to alien concepts, or whether they are better explained by differences in sensory processing architectures. The latter is potentially tractable and seems less likely to lead to misbehavior in the limit, while the former would represent a more severe and potentially intractable mismatch in ontologies.
This is a good point, and I’m not aware of research asking this question as phrased.
Reflections on InkHaven
On today’s panel with Bernie Sanders
The AI x-risk lawsuit waiting to happen
On the political feasibility of stopping AI
AI might surprise itself by going rogue
Diary of a “Doomer”: 12+ years arguing about AI risk (part 3: the LLM era)
In a word: InkHaven.
But seriously, I’m still working full-time on Evitable.com and so am trying to churn out my daily blog posts FAST. There are topics I know I have things to say about, and I try to get them down in words in ~1-2 hours tops. In this case, the motivation is something like: “It’s annoying when people make behaviorist arguments about how AIs are more aligned/trustworthy than people”.
Yes, but if there is rogue AI already on the loose, we can’t stop it from doing RSI and/or taking over unless we can track down or contain all the copies.