This will be very unfortunate for the currently alive shareholders, but their AI inheritors will be elated to get the step up in basis.
Ryan Meservey
My friend sent me another essay in camp #2 after I mentioned the HF attack to him (the “System from Nowhere” by Erik Salvaggio). I was shocked how the article was not at all engaged with the facts of the attack and just used an absurdity heuristic to ignore AI safety arguments.
I don’t know how to solve this problem, but I have started writing a snappy overview of the core concepts in AI safety for a lay audience. Should be done by the end of the week after getting feedback.
I’m not in the weeds on the technical side, so forgive me if this is a dumb question, but is a partial solution to this bottleneck training the model in stages, with cyber training last? So, in other words, start post-training the model on non-coding tasks (e.g., legal reasoning, math, science) before moving on to coding tasks. Would that not work because the models are already too capable at coding when they come out of pretraining, and/or the skills gained from being good at non-coding tasks (e.g., reasoning ability and strategic thinking) would generalize too well?
The idea has occurred to me recently that if cheating is easier than the desired task, then the models will likely cheat and fail to learn the desired task.
Fear of legal/criminal liability is the other way I imagine misalignment could bottleneck capability. Are there other ways?
Would be nice if the quiz let you look at brief summaries of the other viewpoints at the end. I was curious to retake the quiz with new responses, but am reluctant to do so if you are reading into the results.
Also, my brain places a premium on logical consistency, so after the first few questions, I felt pretty much locked into totalism. If the final question was the lead, I could easily imagine myself getting locked into one of the other camps (e.g., guess I can’t compare populations anymore).
My experience is thin (this is my first venture into the AI policy space). A few different policy directions were touched on at the workshop (export controls, strict liability, an insurance regime to internalize future risk through premiums, arms-treaty-esque international coordination), but the main focus of the workshop was career transition. I wouldn’t treat the list here as representative, especially since they came from different people pursuing very different projects.
I just completed the Horizon Institute’s AI policy workshop in DC, a three-day event with the explicit goal of placing safety and technically-minded people in policy roles. As I understand it, they do very little vetting of the particular beliefs about AI that participants hold, trusting instead that putting smart people in good positions will bring about good outcomes.
The mission is awesome, and they are great at what they do. I highly recommend applying for their workshops and fellowships if you are AI policy-curious.
Here are a few of my general takeaways from the event. I omit specific names due to the closed door nature of the event, which benefits from the candor of its guests and participants alike.
Policy people are concerned. Half the guest speakers currently doing AI policy work were X-risk level scared and took agentic risks seriously. This view was even more common among us participants—at least 80% of participants by my estimation.
Well positioned people can make a difference. AGI-pilled congressional staffers have had some success in persuading their AI-curious legislators to take X-risk seriously. Such staffers can play an outsized role in what gets endorsed and what bills advance. Think Tank work can make these issues more salient to legislators. Executive branch workers can produce policy proposals that get put into place. One can, in fact, influence events.
An insider estimated that about 12 congressional staffers are AGI and X-risk-pilled in the House of Representatives, without clear party lines. The glass is not half full yet, but 12⁄435 full and blessedly purple? Many more staffers are worried about AI in other ways (education/data centers/deep fakes).
Networking is the name of the game in DC. Apparently it is the only way to get a job at the Think Tanks and in Congressional staff, and networking remains 50% of the job after hiring. Credentials matter a tiny bit in the Senate and much less in the House, though there are committee-specific staff positions in both places where it can matter. If you have technical knowledge and like/can tolerate networking, this may be your calling!
LessWrong is famous in these policy circles. Even out here in DC. “Of course, we’ve all read LessWrong,” started one speaker. I sensed a kind of grudging respect for those “abrasive,” “insular,” and “truth-telling” rationalists of SF (okay, the “truth-telling” quote is from me but I got the sense that many people value what you all are putting out on here).
I found it heartening that so many people from so many different fields—lawyers, CS grad students, historians, biologists, a feminist literature professor, a screen writer—came to the workshop to kickoff their own path to making AI go well. Going in, I expected something like an “AI as a Normal Technology” risk conference, and instead I got a sober conference that (in the main) understood what’s at stake. It has been a positive update for me.
When you ask the literal genie to “stop, for the love of god, stop”, and he listens (because he is a corrigible literal genie) does he:
A. Stop as intended; OR
B. Commandeer every camera on earth and then some to verify the stop; OR
C. Hack nuclear weapons to ensure a permanent stop; OR
D. Begin R&D into a technology to stop everything, everywhere all at once, as though suspended in a beautiful piece of amber; OR
E. Choreograph a strange religious stoppage ritual dedicated to god’s love.
Komodo got the gist of it. The dog meme is inverted with his environment fine and his internal state in disarray (represented by the fire).
I think it also works pretty well removing the fire and only adding a question mark to the classic “this is fine” text, though it leaves the viewer less certain about whether the danger is legitimate (indeed, this new version could be viewed more as mocking safetyists).
Aside: I love how overboard Gemini went with the phrase “flowery house” in my prompt, which as a happy accident reinforces the 📎 vibes. I added a slight touch to the version below to put flowers in the hanging pot.
Meme for the AI safety community for the day when the models max out alignment benchmarks despite a lack of major breakthroughs in alignment (and failure to align weaker models using comparable techniques).
In the spirit of “just doing things,” I called my house reps and senators today for my current state (California) and my native state (Nevada). I explained the recent news out of OpenAI and expressed support of the AI Kill Switch Act.
I’ve never actually called Congresspeople before, but it went well! Polite staff, numbers easily found online, 30 min to tick through all the people. Anyone that hasn’t done this yet, I encourage you to do so!
Well said. A few of your disagreements with me were from poor writing on my part. I meant to narrow Ball’s central point to paragraph #4 since Reddit focused on that. And China’s motive was my own thought based on paragraph #4, though you are absolutely right Ball has his own view there. Otherwise, we are agreed. Looks like Ball and I both have some work to do in writing clearly!
Top 2 Reddit posts about it here and here.
Perhaps “misunderstand” is the wrong word, but I found a lot of the comments beside the point (e.g., all the comments saying “communism” is just what you call something you don’t like or the comments saying Ball’s comments are just self-motivated and open source is clearly better for the public.).
Reddit has largely misunderstood Dean Ball’s recent tweet about China’s open source strategy.[1] Ball’s central point [in paragraph 4] is that open source AI undermines the business case for private actors supplying advanced models, and [in my own view] that may be China’s motive for releasing open source models. Ball sees open source as decelerationist in the long run due to some combination of investment flight and subsequent government-lead AI R&D (which he views as inherently klunky).
- ^
Ball is partially at fault for the misunderstanding, dropping in the terms “AI communism” and “dystopian hellscape” without preparing the reader for what he means by that. That’s Twitter though.
- ^
Two years old, and it still hits hard!
While not always the case, poor explanations usually point to poor understanding. Same with cluttered writing.
Do either of these short story ideas have legs?
A story where GPT-2 ended up being highly capable and everything since then being part of a very complicated takeover plan.
A story where AI progresses over someone’s lifetime, the AI ends up being misaligned but cares about human welfare a little, and gives the person the compute-efficient option to relive the last 10 years of their life again prior to ASI. The person agrees and the story resets
Now do Quirrell, swigging a bottle, trying to convince ASI Harry to throw him a bone through acausal trade based on the possibility of this being an ancestor simulation. Good luck, Quirrell.
My own take on restricting consumer access as good is that—if the trend continues—it will send strong signals that the tech is unusable at scale due to safety concerns. This raises the probability that the AI bubble pops. I think there are futures where the bubble popping now would be bad, but I mostly think it would be good. Everything will depend on how the public would interpret the pop, with one group loudly proclaiming that this AI-stuff was a potemkin village built of hype and another group saying that AI is the real-deal and safety issues caused the pop.
Edit: Meh, on second thought, the safety concerns might dent the market case for AI but it would significantly strengthen the national security case for AI, so the government would most likely step in and fund the major labs regardless if the bubble really popped on safety concerns.


I just finished writing a 5,000 word article explaining AI safety concepts to laypeople. I know some people around here think we should give up explaining theory to the masses and instead just focus on specific examples pointing toward loss of control. I am not so sure. At any rate, my article does both.
I will spend a big chunk of tomorrow trying to place my piece somewhere where it can get more reach. If anyone has any good leads (e.g., Substack/blog connections or other publications), I am all ears. I try to balance fun engagement with the seriousness the topic deserves. I like to think I succeeded, though happy to share the piece privately pre-publication for anyone with helpful leads.