Co-lead PauseAI Germany, Member Torchbearer community, Master Computer Science, PHD candidate bioinformatics, transitioning to AI safety Governance&Stratety,
Benjamin Schmidt
Hi Charbel,
Sorry, looking back now at my comment it’s pretty weird and too strongly worded. Especially asking you to change it right now and checking back. Thank you, for your very diplomatic answer!
Maybe to rephrase it how I would write it now: I agree that it would be good to have the Q&A back on the website. I think the “invisible cost” of not having “Risk of loss of control” present is often underestimated. I am not saying that as useful assistant one can’t help with the topic of the day without mentioning x-risk, but that the cost of systematically keeping it out of sight is often underestimated. And I have seen people who have been in the community for years, feel awkward talking directly about the topic.
Best regards,
Benjamin
Lots of helpful points.
Thank you for also pointing at CESIA and the website. I checked the website and have some feedback.
PLEASE CHANGE IT! I strongly suggest adding x-risk or loss of control back to the website.
Right now if a staffer or policymaker goes to the website they don’t find x-risk and loss of control. I think policymakers we talk to visiting the website might be actively bad. And I would suggest not just adding loss of control but making it clear that you are talking about extinction risk.
Perspective of staffer: “If X-risk was real, then CESIA would have obviously mentioned it on their website. Everything else would be insane. I was worried there for a bit, but I guess these guys at PAI and CAI are crazy.”
You could even link to more extensive things add a whole tab mentioning gradual disempowerment, loss of control, x-risk, concentration of power, geo-political risks … .
Minimum E.g. on the frontpage, the bottom and in the about section there is this paragraph:Our mission is to prevent and mitigate major AI risks by informing policymakers, raising public awareness, and conducting policy research.
changing it to something like: “Our mission is to prevent and mitigate major AI risks including risk of loss of control and risk of human extinction by informing policymakers, raising public awareness, and conducting policy research.”
Also looking at https://cesia.org/en/publications/les-contributions-du-cesia-au-dialogue-mondial-des-nations-unies-sur-la-gouvernance-de-lia/:
> Our consultation response advocates for clear, internationally enforced AI red lines to prevent unacceptable risks and misuses, including those related to autonomous weapons of mass destruction, biological risks, offensive cyber capabilities, harmful manipulation, and child safety.
Thank you for submitting something and again mentioning the deficit in your post and again mention x-risk, loss of control, gradual disempowerment.
Even Anthropic on their website mentions loss of control in their policy priorities:
> But transparency alone is not sufficient to safeguard against the most serious risks posed by powerful AI, including the ability to help create biological weapons or carry out cyber operations and the loss of control of AI systems. Our Advanced AI Framework lays out what we think governments should do about these risks in the near term.
I put an appointment in my calendar to check the website again next week.
---
PS: An exercise I suggest to anyone who feels in any way weird talking about x-risk: Talk to ten random people and say: “Sorry, do you have a moment to chat about something? Did you know AI might kill everyone in a few years?” (not because this is an effective way to engage but to get it out of your system). Repeat until the part of saying AI might kill everyone doesn’t feel weird anymore.
Good question to ask, thank you.
Using the AI futures model what kind of shift in actual time is this when it happens at what time. It might be included in the graphs already but I didn’t get it :).
Great post (Raymond linked it to me). Gave me some things to think about.
> By aiming for the moonshot, you miss the cheap asks that could actually pass. I don’t know, I feel like there are cheap asks (say, for example, transparency, or mandatory incident reporting) that are sometimes neglected by outsiders shooting for the moonshot Moratorium—and I feel that this is not strategic, because the cheap ask has a serious shot of being implemented, and this could be a meaningful improvement over literally nothing.Agreed. If you have concrete asks which would be relevant for EU countries like Germany it would be awesome to haven them written down in a structured way to provide to policymakers (ofc. things depend on the country).
PS: I wish we had a CeSIA in Germany.
I will answer now, Matilda/Maxime probably went to bed already.
“The right terminal goal is a binding international pause treaty with verification/enforcement” --> I agree functionally, ideally a pause for at least a couple of years, or at minimum the capacity to pause. Though I think there are much more palatable ways to present this to policymakers than “pause.
Could you expand upon this? I am partially unsure what you mean. Also interested in which ways you would present it to policymakers? (Red lines?)
I think there are also good reasons to pause earlier because later on a Pause might be more difficult/impossible to implement + gradual disempowerment + …
> Policymakers need strong enough incentives to act
I think this depends. 1)The current administration in the US is quite weird. 2) I think for anyone outside of the admin to ask for a ban of Fable specifically (not that we would have wanted to) we would have needed strong incentives for policymakers to act. The more ASI/AGI pilled policymakers are the less incentives they need and maybe things align well in other ways. Also Section II of the post.
A side thought: WuWei don’t try to fight, try to work with the system and create a system which already does what you want. Europe is also waking up to the security and sovereignty implications of strong AI systems.
Also, the better the incentives (for a good version of a pause) the better the version of a pause we will be able to implement.
> “The movement can grow fast without breaking”—I don’t knowDo you have specific scenarios in mind? The federation model is designed to help stop it but ofc. nothing is foolproof. We try to be positive EV compared to the alternative like you mention next.
More money would help as well!
> it seems to me that whoever is in contact with the White House in the community has much more leverage
If PauseAI grew in a timeline with high AI salience this might shift, though the current White House is weird.
> not that important (people in administrations)
Mostly agree, adding some thoughts. I think a big enough public movement will also influence these people but it’s very different from policymakers. This is where we need good think tanks/policy people. Because the situation emerged we also did a lot of talking to policymakers in Germany and even the bureaucracy once.
Another way to influence the bureaucracy is from the outside when policy interacts with bureaucracy e.g. right now the new German AI security institute is being created and can be shaped to be more useful for a Pause than e.g. the UK one.
Thank you. I think this could have been more clear in the post and also in my mind even if something directionally similar existed.
> e.g. internally routing certain things up the hierarchy.
This is part of the federation model (III.2) though I it’s not explicit for everything relevant here yet.
As an a practical example of how things work: In Germany right now the person for policymakers to contact is very clearly our main lobbyist (other co-lead) who does similar work as CAI building relationships with policymakers and then CAI whom we work together with closely in Germany. Just two days ago we had an event at the Bundestag we and CAI attended.
For high level enough press requests/policy questions to a local chapter they would contact me, I would contact global.
We haven’t yet figured out how to find a good equilibrium with humans in control given humans not doing the work (https://gradual-disempowerment.ai/). If humans were ever in control https://www.lesswrong.com/posts/kbezWvZsMos6TSyfj/the-eldritch-in-the-21st-century .
The biggest problem with gradual disempowerment is that we want it.
Ignoring all of that, I would try to avoid as many sign of immoral mazeness as possible!
A few reasons to stop right now:
Each new model shortens the path to ASI, makes a pause more restrictive and difficult.
There might be a huge gradual disempowerment overhang already which will continue to grow during the next months/years. Adoption, societal norm changing, other technology like robots, drones, self-driving cars and smart glasses and … and unknown unknowns are all lacking behind and thus GD is lacking behind. If we continue we might already end up in a bad GD future.
There is little overhang but the concentration of power and gradual disempowerment might make a pause more difficult
Don’t trust companies, countries to actually pause in the future.
I am curious what you think in the Gradual Disempowerment (GD) direction
Edit: Kaj’s answer wasn’t there yet probably some stuff is double.
Prayer seems like a mix of meditation and some other practices to me. Probably works for the same reason as those practices. As for how they work? No idea. People are trying to figure it out but it will take a while.
I think it’s more interesting to first figure out whether they do something at all. For that I would focus on getting the most experienced long term practitioner and check the phenomenology they report. These are often far from Placebo effects.
Ignore all their meta-physical and most of the how it works claims, just the phenomenology of them and if they are treating/working with someone their patients. At the same time record a bunch of data on their bodies etc. You can e.g. look into the studies on cessations of consciousness which have come out over the last 3 years (https://www.biorxiv.org/content/10.64898/2026.02.10.705005v1)
There is probably a lot of people out there praying who get some small benefits and then there are some few people who can just think of god and enter a state with similar phenomenology to MDMA. I would start with those.
What is your p-doom developing ASI/AGI with anything like the current methods in the current environment? My assumption is that you just think the technical problem is getting solved.
I think you mentioned most of the things I would have mentioned from my impression of Goenka from the outside: Especially not reasonable according to rationalist/post-rationalist standards, caught up in doctrine. Also, kind of narrow meditation experience/expertise because they are stuck in one tradition with doctrine.
I think the combination of some Vispassana and no good maps/letting people hardcore meditate for 10 days without telling them about most possible negative effects is pretty irresponsible. The Dark Night of the Soul after A&P seems to be a thing at least for some percentage of people, psychosis, trauma reactions …
I am curious if they mentioned A&P and Vispassana Nanas did they mention the Nanas of suffering (5-10)?
For a good teacher I would e.g. recommend Roger Thisdell, you can check him out through a weekly session from his patreon.
If you haven’t read it I can recommend meditationbook.page and or MCTB.
Awesome idea. I think this is one of the ways AI can! improve our society.
I recently tried creating something like this for Astralcodex comments.
There were quite a few commenters who wrote 90% incorrect things which were annoying to check. For these this feels very helpful. For more nuanced disagreements maybe something like providing context or nothing! would be better.
I was struggling with finding the right prompt and how much compute to put into checking things.The entire post is investigated in a single agentic call. Rather than extracting claims first and investigating them individually, we send the full post text and ask the model to identify claims, investigate them, and return structured output mapping each verdict to a specific text span. This gives the model full context (a claim’s meaning often depends on surrounding paragraphs) and reduces round-trips.
Right now this approach will give AI Slop/bad results both on the false/true positive side. If I wanted an AI to do this I would at least ask it to create a list of statements, use a sub-agent for each statement to check all of them in the context of the whole. Have another subagent double check for each comment. Possibly with different prompts and a bias towards false negatives. I know this is more expensive but the way it is right now it’s essentially bad to look at the output. (If you chose to answer to one part of this comment, this seems like the most important).
Which prompt to use (Different prompts will provide very different results which also hints at underlying issues which still exist)?
Try to get a high percentage of true corrections, false positives here can be bad.
To get a good fact check right now you need a lot of compute. Maybe people could chose to run different levels? Maybe people could send money back to the person who originally checked it? There could even be the possibility of making profit.
It doesn’t work well for cutting edge ideas especially since AIs are still often overconfident.
Which writing style? I think everyone reading the same default AI writing style is at least bad, possibly quite bad, so switching writing style between comments might be important.
Bad 2nd order effects of everyone converging to the LLM epistemology or something else.
For someone like me who uses a subagent with a specific prompt to check every comment it might be good while maybe the default user will just get whatever XAI or Meta want them to read? I am unsure whether for society this will go into a good direction/how to get it there. Similar as twitter can be great for some people but is terrible or most of society.
PauseAI, ControlAI, Torchbearers are somewhat in contact/cooperating. I don’t know the details for Global and less for US.
PauseAI is definitely trying to and talking to politicians. For the German chapter we have at least 3 conversations with German MPs this month and one of our ex-coleads works for ControlAI in Germany. At Pausecon in Europe PauseAI also just did a briefing for several EU MPs in the European parliament and gave a draft for a resolution. The hosting EMP started the meeting off with a quote from Terminator. It was not a joke.The system goes online August 4th, 1997. Human decisions are removed from strategic defense. Skynet begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, August 29th. In a panic, they try to pull the plug.
Please contact at least your local policymaker and get a meeting with them (your chances there are way higher). Torchbearers also just started a training program for this: https://www.diptraining.org/ or you could join PauseAI.
Offer to anyone reading this: If you give me your available times, name etc. I will literally organize the meeting for you. You just have to go there.
Most people/policymakers e.g. have no idea that AI developers don’t really understand their systems. They also have no idea about what scientists or the AI company CEOs think or say or what the current models can do.
The average model of AI might be something like this: AIs are hallucinating a lot, are programmed like a normal program which people perfectly understand, can be controlled, didn’t get much better the last few years.
You could join Torchbearers (https://www.torchbearer.community/) or PauseAI to inform policymakers about what’s actually going on.
Because of that, people who prevent bad outcomes often get treated as though they’ve done nothing, or even as though they were dramatic for worrying. Which is a pretty fucked up reward structure when you think about it.
This is more towards the personal for someone who does the work vs how society should act towards them:
The master does nothing yet leaves nothing undone.” (Tao Te Ching)
One interpretation: No one knows you saved the world or everyone thinks they did it themselves.
You have a right to perform your prescribed duties, but you are not entitled to the fruits of your actions. Never consider yourself to be the cause of the results of your activities, nor be attached to inaction. (Bhagavad Gita: Chapter 2, Verse 47)
When people trying to save the world are working for the recognition/results they quickly start goodhearting.
PS: Love the post title.
Awesome work. For anyone reading this: Please try to talk to your local policymakers. As a constituent it is much easier to get a meeting.
Torchbearers (which is a community adjacent to ControlAI) and/or PauseAI are happy to give you advice or coaching!https://www.diptraining.org/
https://pauseai.info/join
I have goals that benefit from having hundreds of millions to billions of dollars. So do other people. Money is for steering the world. I can use money to hire other people and get them to do things I want.
How do you stand towards pluralism or democracy? There is some tension there with people having billions of dollars of steering influence but of course there is with taking peoples money away as well. Money could also be used to steer towards a more pluralistic society etc. …
Anything to read which approximately describes you view there?
Well, they’ll probably still exist.
It seems more likely to me that the Malawi people and everyone else will be killed at some point.
Currently the system still has to consider popular opinion to some degree. Killing all the Malavi people would not be efficient right now. When that incentive disappears enough, I would expect everyone else to get eliminated (All of this assumes incentive aligned AI which I wouldn’t expect). This goes in the direction the author was mentioning that the reason for moral progress is more about which societal structures are efficient not actual “moral” progress”.
If competition between humans persists I would expect the other last humans to disappear as well having to transfer all their power to survive the competition.
All of these seem relevant but as a starting point I would put something like most humans having a lot of technical debt (https://sashachapin.substack.com/p/review-meditation-from-cold-start), not feeling safe and okay in there here and now. Maybe not being enlightened. And then we could get to work on the difficult problems you bring up. Of course, there is the trap here of going off the rails using spirituality or me already being off the rails.