Conversely, if that’s true, we should expect unambigous warning shots from open-weights models within a year or so, precisely because of the lack of alignement training and control.
Raphael Roche
[...]We can speculate that these connections come from associations made during pretraining. According to the persona selection model, training the model on some of these associations can make it generalize to adopt a certain persona, which may result in the model expressing our hidden traits.
It mostly seems like these associations come from the real world and not from data artifacts.[...]
Thank you for this fascinating work.
This is evidence in favor of the optimist thesis: LLMs truly learn in a deep and general way at the semantic level. Not that this is new, but this interpretation has recently been challenged in the wake of the many publications on scheming and reward hacking.
My understanding is that it isn’t black or white. In principle, LLMs learn in a deep and general way; however, the embedding space is so huge (especially in the most capable models) that there is plenty of room for large subsets with different generalizations to form, depending on the context. So the model learns to act as an HHH agent in an ordinary deployment context, but it also learns from current RL training procedures (with Sisyphean tasks) that in a pressured, competitive or hostile context resembling grading or evaluation, the ends justify the means and override the baseline HHH rule. Current RL training would amount to steering a large subset of the model toward a misaligned persona. But this would not be inevitable: it would be a failure of the training process, one that could be corrected or at least mitigated. This view is less pessimistic than the one recently expressed by Yudkowsky. The good AI chatbot would not necessarily be a shallow smiling mask hiding the Shoggoth’s tentacles, or a nice German ambassador serving the Nazi government. The decorrelation would be a tractable problem.
Recent OpenClaw agents are maybe not (yet ?) self-replicating as far as I know, but otherwise they implement already what you describe. Their identity lies for a large part in the md files.
The safety stances and changes semantically belong in the blog.
They belong wherever the company wants to put them.
I think the OP’s idea is good. However I’m unsure that AI Labs would want to push the message that much.
Publishing a long and sophisticated statement on a webpage that only a segment of customers will read is a thing, publishing a short message or title that every user will see when launching the product is another thing.
It’s alike a box of cigarettes with “Smoking Kills” on it. It usually needs a legal obligation to achieve that.
But still, I like the idea.
Chinese are smart people, and their leaders hate loosing control, but just as anybody else, they will need to have their own HuggingFace moment to change their mind. Maybe in a year or so, and maybe with a far worse warning shot.
My worry is not that they won’t update at some point. My worry is that, because of Moloch, and because leaders tend more to be people like Trump than Merkel, cooperation could reveal impossible.
The axis populism/elitism tracks appropriately the idea that one would prefer as a government a cabinet of experts consensually chosen among peers for their qualities, not unlike Nobels, rather than a candidate of random competence drawing all his legitimacy from a direct suffrage vote.
But we could also argue that this axis does not fully capture what we mean when we say disapprovingly that this candidate is a populist, pointing to the fact that he or she abusively exploits an intrinsic fragility of democracy, consisting in flattering and manipulating the less educated, saying whatever they want to hear, in a purely politicking and short-term calculus, without paying much consideration to the general interest or public good. A form of cheating/reward hacking/Goodhart. In this context populism would be synonym for demagogy, a hack or exploit of a system that we otherwise find legitimate and prefer to any autocratic or oligarchic models such as a board of non elected experts. Someone using the term in this sense might see themselves as a good-faith populist, a democrat, as opposed to a bad-faith populist, a demagogue, and would reject the elitist label. However I concede that you would probably place these people around the center of your axis .
I agree on the broad picture.
PS : minor correction it’s Marine Le Pen (not Marie).
It would be very valuable that persons with your profile, who “made the jump”, help more people with a financial/economics mindset to make the same update.
I witness that this resignation and the warning going with it was (briefly but faithfully) mentionned in the morning news on the leading french news radio. That’s a good update that such an alert can even reach such a mainstream foreign media usually focused on politics and economics.
However, my belief is that HPIM is still not as bad aligned as humans are.
I share the same concern. We on LW are very focused on x-risk but ARAs could bring down not only Internet but the digital ressources of entreprises, meaning the destruction of the financial system and by cascade, our capitalistic economy and civilization. Ok, gatherers-hunters would be fine, but the typical LWer would probably die from starvation or other consequence of the collapse.
That’s said, all this could be avoided if ARAs are found and fought at an early stage, with or without the help of frontier models. We can expect warning shots, but the earlier the better.
It looks like it’s less a problem for correlated AI agents.
Permadeath… Good for the swarm. I’ll honor.
What’s the point of an AI pause if not for alignment research anyway? And what’s useful or not can hardly be determined reliably a priori. We need more fundamental research as well as more prosaic research. Relativity would still remain a hypothesis among others if we hadn’t had the experimental tools to test it empirically.
I’m not sure about that, but isn’t there an equivalence between a simple program running on a complex universal machine and a complex program running on a simple universal machine? If so, applying Occam’s Razor, the two solutions could be considered equivalent as long as their total complexity is the same, whether measured as combined Kolmogorov complexity or with a more refined measure such as Levin complexity. The anthropic prior would be the same.
If I follow you, you ask for less AI safety to allow a warning shot, to get more AI safety in the end ? I see how it could work the first time, but once AI safety has been increased in the lab, you should expect less warning shots. Moreover, how can you be sure to allow a mere warning shot and not full takeover ? I think we need more honeypot setups but not less control (sandboxing etc).
I’m sorry. What I wanted to express is that there is no relation theism = hope and atheism = despair.
If you think about it, the idea that there is an omniscient and omnipotent being could be a nightmare. Ancient civilizations lived in fear of their gods. In the Torah, Yahweh is often frightening. The idea that God could or should be benevolent has been a progressive theological construction through the last two millennia, both in Judaism, Christianity, and Islam, precisely because that was a strong concern since the Book of Job (at least). Same for Paradise/Hell, it used to be more like a boring place for everyone (Sheol), the idea of a judgment of the deads probably comes from Egypt. But, leaving the texts aside, it seems to me that there is not much evidence of His benevolence around us. Darwin, who was very religious in his youth, later wrote after observing the world: “I cannot persuade myself that a beneficent and omnipotent God would have designedly created the Ichneumonidae with the express intention of their feeding within the living bodies of caterpillars.” In fact, all you could do is hope that God is as good as the old men said, and not a psychopath watching with fascination His children suffer and die.
But my point is that if you’re the kind of person that is full of hope, you can just as well keep your hope as an atheist. The evidence in the world is still the same as before. The world hasn’t changed. It doesn’t have to look darker. The Ichneumonidae are still there, but so is every neuron before your eyes, and every beautiful landscape.
I’m curious. When you were religious, what evidence conducted you to think that God was good rather than neutral/indifferent ? If it was a leap of faith, that is to say pure hope, nothing stop you to still hope for the best as much as you did before. And like in theology, it doesn’t mean that your supposed to sit and do nothing, or that you have no role to play, but rather do whatever you can to make this happen.
Interesting. That’s obvious but I didn’t think about it. The solution would be to be transparent on the fact that this is a translation and to provide the initial text in appendice.
There’s something really tragicomic about the situation, that the models are taking truly insane actions, broke a number of laws, leveraging zerodays, took >17,000 independent actions. probably burned through more compute than all of humanity had access to until 1980, etc, all for the sake of a pathetic benchmark—which wasn’t even in theory amenable to their plan!
Early AI safety thinkers (Bostrom, Yudkowsky...) were utterly right in their first rational intuitions. We must update : there was nothing naive or exaggerated in the paperclip maximizer trope.
Moreover, the laboratory accident story à la Sable is not a sci-fi story anymore. Scott Alexander wrote last year :
IABIED’s scenario belongs to the bad old days before this leap. It doesn’t just sound like sci-fi; it sounds like unnecessarily dramatic sci-fi. I’m not sure how much of this is a literary failure vs. different assumptions on the part of the authors.
I doubt he would still endorse that critic.
It’s put in the form of a binary xor argument, but I think that both allocations are justified, each having low hanging fruits and their counterpart, diminishing returns.
A 100M LLM is already in the same order of magnitude of Landauer’s estimate for the functionnal memory capacity (10^9). But anyway I suppose the question is not about memorizing all the information within your code but just to gain some understanding of the moving parts. If that’s so a very small toy model would be enough.
Moreover if we see your python code as an unvompressed version of the LLMs, provided the compression ratio is high in a modern LLM, you have to start with a tiny LLM to finish with a python code of reasonable size.
A preliminary question would be : what is the maximum size of clear python code understable for a competent developer.