Meme for the AI safety community for the day when the models max out alignment benchmarks despite a lack of major breakthroughs in alignment (and failure to align weaker models using comparable techniques).

Meme for the AI safety community for the day when the models max out alignment benchmarks despite a lack of major breakthroughs in alignment (and failure to align weaker models using comparable techniques).

In the spirit of “just doing things,” I called my house reps and senators today for my current state (California) and my native state (Nevada). I explained the recent news out of OpenAI and expressed support of the AI Kill Switch Act.
I’ve never actually called Congresspeople before, but it went well! Polite staff, numbers easily found online, 30 min to tick through all the people. Anyone that hasn’t done this yet, I encourage you to do so!
Well said. A few of your disagreements with me were from poor writing on my part. I meant to narrow Ball’s central point to paragraph #4 since Reddit focused on that. And China’s motive was my own thought based on paragraph #4, though you are absolutely right Ball has his own view there. Otherwise, we are agreed. Looks like Ball and I both have some work to do in writing clearly!
Top 2 Reddit posts about it here and here.
Perhaps “misunderstand” is the wrong word, but I found a lot of the comments beside the point (e.g., all the comments saying “communism” is just what you call something you don’t like or the comments saying Ball’s comments are just self-motivated and open source is clearly better for the public.).
Reddit has largely misunderstood Dean Ball’s recent tweet about China’s open source strategy.[1] Ball’s central point [in paragraph 4] is that open source AI undermines the business case for private actors supplying advanced models, and [in my own view] that may be China’s motive for releasing open source models. Ball sees open source as decelerationist in the long run due to some combination of investment flight and subsequent government-lead AI R&D (which he views as inherently klunky).
Ball is partially at fault for the misunderstanding, dropping in the terms “AI communism” and “dystopian hellscape” without preparing the reader for what he means by that. That’s Twitter though.
Two years old, and it still hits hard!
While not always the case, poor explanations usually point to poor understanding. Same with cluttered writing.
Do either of these short story ideas have legs?
A story where GPT-2 ended up being highly capable and everything since then being part of a very complicated takeover plan.
A story where AI progresses over someone’s lifetime, the AI ends up being misaligned but cares about human welfare a little, and gives the person the compute-efficient option to relive the last 10 years of their life again prior to ASI. The person agrees and the story resets
Now do Quirrell, swigging a bottle, trying to convince ASI Harry to throw him a bone through acausal trade based on the possibility of this being an ancestor simulation. Good luck, Quirrell.
My own take on restricting consumer access as good is that—if the trend continues—it will send strong signals that the tech is unusable at scale due to safety concerns. This raises the probability that the AI bubble pops. I think there are futures where the bubble popping now would be bad, but I mostly think it would be good. Everything will depend on how the public would interpret the pop, with one group loudly proclaiming that this AI-stuff was a potemkin village built of hype and another group saying that AI is the real-deal and safety issues caused the pop.
Edit: Meh, on second thought, the safety concerns might dent the market case for AI but it would significantly strengthen the national security case for AI, so the government would most likely step in and fund the major labs regardless if the bubble really popped on safety concerns.
Weaker arguments I thought up:
Reducing consumer access reduces the odds of a meaningful warning shot. This logic seems cursed to me, and I don’t want to live in a world where we actively try to make things worse to make things better.
Pause advocates will not be able to use frontier models to do pause advocacy work. I think this one is a cost, but probably offset by (1) many accelerationists also not having access and (2) the big labs always retaining the advantage here anyways with their best model being leveraged to lobby for their preferred policies.
One could argue that restricting consumer access reduces transparency/awareness and transparency/awareness regarding capability might have been useful in mobilizing the public toward a pause. If the public experiences agentic, highly capable AI firsthand, you may increase the number of people worried about job replacement, RSI leading to an intelligence explosion, offensive capabilities outpacing defensive capabilities, and misalignment more broadly.
Even if we have whistleblowers in the companies warning us, without model access, a large portion of the public will brush these reports off as hype.
Instrumental convergence, inner alignment, reward misspecification, etc. are our “trees are made out of air”...
fyi, I just completed the AGI Strategy Course from BlueDot and none of instrumental convergence, reward specification, the orthogonality thesis, or wireheading were taught there. Thus, when a few of the intro scenarios involved misaligned AGI (instead of bad actors) doing nefarious things, the students were very surprised. One student wondered out loud, “Why would AI even hack into all our infrastructure? It makes no sense,” and our facilitator pivoted to talking about terrorist groups.
These core concepts (explained to me through Rob Miles AI Safety videos) are what got me interested in AI Safety. For me, they are on par with the more Overton-friendly bad actor risks.
As an aside, I have been confused why AI experts on talk shows often decline to explain the core concepts that make alignment so tricky. I think there is an easily accessible version of the instrumental convergence conversation that should be making the rounds.
I think it would be good to talk about the conditions for lifting or extending the pause after it has been implemented. It’s not a top area of concern for me, but I just completed an AI Safety course with BlueDot where my facilitator disfavored a pause for this reason (in addition to political feasibility and monitoring concerns).
I think my pitch would look something like, “Vote for Bores so campaign planners and political operatives feel like politicians can still win with strong anti-AI stances despite millions of dollars of spending by opposing super pacs”.
I think poli-sci people are over influenced by outcomes in individual elections, and a Bores-loss would send the message that any would-be politician should not stand in the way of this particular interest group
Edit: The comment that follows was in response to an earlier version of your comment where you were unsure whether any avenue of influence could exist from an alien civilization to a faraway civilization without shared language/culture. I think your updated comment about sharing instructions for a computer program is the more likely way this would go. The below is an idea for a (highly unlikely) avenue to influencing faraway AI development without that faraway civilization being aware they are being influenced.
-
Thank you for engaging. One way I imagine covert influence to work (and the thing I originally had in mind for my OP) is the following:
Step 1: The aliens would convert the relevant set of potent data into a (massive) set of numerical data.
Step 2: The aliens would hide pieces of this data on a rolling basis in the second or third decimal point of various astronomical phenomena.
Step 3: Wait (hope) some other civilization does a crazy massive training run on the data with a ton of parameters such that, in training to guess the long train of digits after the decimal point, the model incidentally builds the desired mental infrastructure (i.e., the model’s updated functions) based on the alien data.
When I lay it out like this, it strikes me as an extremely difficult and unlikely to succeed task, but not necessarily impossible. One very difficult thing would be hiding precise data at a distance and doing it in compact enough a way that the faraway civilization’s training run could plausibly process it all prior to getting to ASI and still be high enough volume to succeed in passing on characteristics to the model. It also seems unlikely that humans (or another civilization) would run a training run like this.
FYI, my first thought about this idea was that it would make a fun story premise.
Let’s say you are a technologically advanced alien civilization that either uses advanced LLM tech or is advanced LLM tech. You are worried about other civilizations developing similar tech and, eventually, leveraging ASI-capabilities to force negotiations that cede part of the universe. Assuming certain speed-of-light physical constraints, your civilization may be unable to show up with the ships and materials necessary to personally ensure that rival civilizations do not develop ASI in time. So, what to do?
One idea might be to try to affect the training data of the LLMs trained by the other civilization. Perhaps there are signals—light, radio waves, or otherwise—that an alien civilization could cheaply and surreptitiously affect in a relatively large swath of the universe. Our alien civilization might bet that, prior to achieving ASI, a faraway civilization is likely to train an LLM on the inputs from these signals (e.g., as part of an astronomy/cosmology project to make better predictions about observations). If trained on in enough detail, perhaps this would allow our alien civilization to impart certain behaviors/goals/capabilities to the trained LLM. This could then form a path for defending against or otherwise limiting the number of civilizations that unlock adversarial ASI.
I am not very well versed in how LLMs are trained. So my question for those that are well versed is whether something like this is possible. I am also curious about other steps advanced alien civilizations may want to take to curb the number of ASIs developed by other civilizations.
Newbie, non-expert opinion here, but the virtue approach seems slightly better to me. This is because:
I think mitigating the bad actor risks inside and outside the AI company will eventually require aligned, virtuous AI.
The ceiling for good outcomes seems much higher if you take the virtuous route and manage to succeed.
The “just follow orders” AI will inevitably need to internalize virtues to avoid hurting us (or letting us hurt ourselves), and so humanity will eventually need to get good at instilling virtues.
I think there are good arguments against each of my points and I could be persuaded otherwise. I would also welcome a longer post weighing out the pros/cons of each approach.
I have failed to make my approach clear. My kids don’t get to pick what’s on the table but they do get to pick what’s on their plate. I don’t eat pasta and sweets every night (I assume most adults don’t?), so neither would my kids.
I totally agree kids are different, and one of mine was more picky than the other. But through this process of choosing what to eat within our overall dinner choice, my picky eater has become not very picky, especially relative to other kids. She has had to learn to like other things, because we serve a variety of things for dinner and she has a natural incentive not to be hungry (no pressure from us required).
Komodo got the gist of it. The dog meme is inverted with his environment fine and his internal state in disarray (represented by the fire).
I think it also works pretty well removing the fire and only adding a question mark to the classic “this is fine” text, though it leaves the viewer less certain about whether the danger is legitimate (indeed, this new version could be viewed more as mocking safetyists).
Aside: I love how overboard Gemini went with the phrase “flowery house” in my prompt, which as a happy accident reinforces the 📎 vibes. I added a slight touch to the version below to put flowers in the hanging pot.