I am a member of the technical staff at OpenAI, working in the alignment team, as well as a Catalyst professor of Computer Science at Harvard. See https://windowsontheory.org/ for my blog, and https://x.com/boazbaraktcs for my twitter profile.
Boaz Barak
Wrote a blog post about Math after AI https://windowsontheory.org/2026/08/24/math-after-ai/
That is an interesting perspective. So, if we use the language of our Model Spec, the “system” level (which can be overridden by a system message) correspond for corrigibility, but the “root” level that cannot be overridden does not. I tend to think of “value alignment” as being more about giving general values than instructions, rather than distinguishing between overridable vs. non overridable instructions.
(Apologies if points I am making here are already repeated in the other comments below—didn’t read them all.)
I disagree with the position that refusals/guardrails/jailbreaks are not “intent alignment”. I view the instruction hierarchy (e.g. system > developer > user) and following policies such as the Model Spec as intent alignment. “Intent” is not just that of the end user but of all the entities authorized to give the model instructions, and negotiating conflicts properly between them is part of following the intent. Indeed, I think being able to follow instructions or policies is key to corrigibility.
That said, I am not an “intent alignment absolutist” and I believe value alignment is important as well. This is because of generalization: the instructions and policies we give models could never cover all cases, and we want models to have good values in interpreting policies as well as dealing with novel unforeseen situations.
It’s a choice what level of independence or autonomy to give to agents. I think we can make the choice of a human centered world, where the economy remains about what humans value. And as more of their material needs are satisfied, what humans value might move more towards relations with other humans.
My point is that there is single objective function that corresponds to “making the economy better”- humans get to choose their own objectives. Yes, if our goal was to advance on the Kardashev scale as fast as possible, and maximize energy production, then it might make sense to just hand over control to the AIs. But we don’t have to make it our goal.
I think it’s better, though agree I’d rather not have it. To Claude’s credit, last time I tried, it was very happy to give me instructions how to switch from claude code to codex (and similarly ChatGPT was happy to instruct me how to switch in the other direction). I hope we maintain this!
Feel free to stop paying attention to me now! I am not a super forecaster nor pretending to be one!
I agree that leaders can do unpopular things in a democracy, but I think it is much harder in a democracy to pursue wildly unpopular things if people feel it in their daily lives, and it becomes the top issue for them as voters. So leaders can get away with some unpopular steps in foreign policy, until voters feel it in their pockets. As a case in point, the one aspect that governments do seem sensitive to is gas prices and stock market.
All Watched Over
As METR notes, the fact that such incidents are detected is a positive sign about our safety practices. Generally I am not arguing that misalignment never exists. But we can still rely on monitors to detect it precisely because we don’t have widespread collusion. For example, the whole reason chain of thought monitorability is useful for safety is because we can trust monitors to read the chain of thought and report what they see in it.
The reason I care about my notion of adversarial is that it is the notion that is relevant for using models to monitor other models.
If I, as a professor, had two twin brothers in my course, it would be highly imprudent to ask one of them to proctor the exam of the other. Yet we regularly have one copy of a model monitor another copy.
The reason it works, and the two models do not collude with one another is not because they are perfect and don’t make any mistake, and not because they don’t have never had a misaligned thought.
The reason is works is because current models basically do the task that was given to them. If one copy of the model was given the prompt “Do X” and another copy of the model was given the prompt “Report to me how well the other model did X” then they would each do this.
Yes, they can make mistakes, but as I write, the lack of collusion means that it is possible to amplify. For example, if I want to be really sure the model didn’t do anything bad I can run multiple copies of the monitor and can also vary some of their prompts.
All of this would not make sense if models had their own agenda independent of the prompt. If we were in that world, I could have 100 monitors watching an agent, and all 101 of them would collude together to decide what to report.It also means that by different prompting and context, I am not simulating a single “super employee” but a collection of many employees that can also check on one another.
This could change in two ways:
1. We could start to observe models behave in ways that can only be explained by models having their own persistent long term goals that are the same goals whether they are being used as an actor, monitor, or anything else.
2. We could radically change how we deploy models by deploying them as a single full context entity that everyone talks to, rather than a large collection of agents with restricted contexts.
I think it is important to do evaluations as well as monitoring in both training and deployment to watch out for 1. And I think it will be a bad idea to deploy models in the mode of 2.
It’s 2030 and we fucked up. How did it happen?
Homework zero in my AI safety course is now online
Like Ryan, its not a crux for me—I don’t want AIs to be dictators, benevolent or not, and I also don’t want to replace democracy by an unelected council of people, no matter how wise or good they are.
But FWIW I asked GPT 5.5 the same question and this is the answer:
I would not pick one person. I’d pick a council, require disagreement, make decisions legible, and build in democratic/legal constraints. But forced to name 10 living people I’d actually defer to, my list would be:
My underlying values would be: reduce suffering, preserve human freedom and dignity, protect liberal-democratic institutions, care about the worst-off, take catastrophic risk seriously, respect truth-seeking, and distrust concentrated unchecked power.
Amartya Sen — because his work joins welfare, freedom, democracy, famine prevention, and social choice. He has the right combination of moral seriousness and anti-utopian pluralism.
Martha Nussbaum — because the capabilities approach is one of the best frameworks for asking “what should power actually protect and enable?” Her work spans law, ethics, human rights, political philosophy, and animal ethics.
Danielle Allen — because she thinks deeply about democracy as a working institution, not just a slogan, and combines political philosophy, public ethics, civic education, and tech ethics.
Audrey Tang — because if immense power included technological power, I’d want someone whose instinct is to make systems more participatory, transparent, and trust-building rather than more centralized. Tang’s work in Taiwan has focused on digital democracy, open governance, and civic tech.
Ngozi Okonjo-Iweala — because she combines development, finance, global governance, vaccine/public-health experience, and practical negotiation across countries. She has led the WTO since 2021 and previously chaired Gavi.
Atul Gawande — because he is unusually good at turning humane values into operational systems: surgery, public health, checklists, end-of-life care, and global health implementation. He served as USAID’s Assistant Administrator for Global Health from 2022 to 2025.
Daron Acemoglu — because he is obsessed, in the right way, with institutions, power concentration, democracy, technology, and inequality. For immense power, I’d want someone constantly asking how it deforms incentives and institutions.
Mary Robinson — because she brings human rights, climate justice, statesmanship, and a long record of public responsibility. She was President of Ireland and later UN High Commissioner for Human Rights.
Angela Merkel — not because I agree with all her choices, but because she represents caution, scientific temperament, institutional patience, and crisis management under democratic constraints. She served as German chancellor from 2005 to 2021.
Holden Karnofsky — because for AI-like immense power specifically, I’d want someone who has spent years thinking concretely about catastrophic risk, philanthropy, cause prioritization, and governance of advanced AI. I would not make him sole decider, but I would want him in the room.
The people I’d trust most are not the most charismatic or the most ideologically pure. They are people who seem likely to say: “This power should be constrained, distributed, audited, and used first for those who are most vulnerable.”
I am not “outwitting the FBI” is a good operationalization, but I agree that whether through misalignment or misuse, none of the existing models (eg GPT 5.5, Opus 4.8, Myhtos, Gemini etc) can cause human extinction
So the Op Ed is only relevant in the 3% of the probability space where people need advice for adapting economically for AI
Glad to have inspired you. Looking at your profile, your P(doom) is 92%. You can think of this op Ed as focusing on the probability space that you assign 8% to
AI is a Meteor. Don’t Be a Dinosaur.
There is also another benefit of working in a lab that is related to the difference between “off policy” and “on policy” reinforcement learning. Even if you had passive access to all the internal information in a lab, you do not gain from that as much as you do by being able to run your own experiments and learn from them. (Or make your own queries to people that have run experiments and learn from them.)
once you write down a table with the first row and column, the result will be batshit insane no matter what numbers you put in.
I am obviously biased as a member of OpenAI’s alignment team, but the following is a position I’ve held for a very long time: you should strongly suspect any argument of the form “you should make the situation worse in the present so it can be better in the future”. People that try to do this usually succeed in the former but not in the latter.
If you have the skills to be a safety researcher in frontier lab then there are many great options of employment for you. If are not feeling that your work matters and you’re making a positive impact, that’s a good reason to look for other options. But if you feel that you are making a positive impact in terms of making AI models safer today, but somehow it would is better if they were less safe, then that’s highly unlikely to work out. It is also extremely risky, since indeed we cannot control how harmful the next “warning shot” would be.
I realize it does not look like this from the outside, but at OpenAI it does not feel at all that safety or alignment people are disempowered. My sense is that the “warning shots” have been heard loud and clear by people across all teams in the company. But I will not try to defend this claim or argue about it here. I do not expect people from outside to take our word for it, but hope and expect that we will show improvement with deeds rather than words.
I am also not claiming that OpenAI or other frontier labs are the best place to have positive impact as a safety researcher. There is no generic answer here that applies to each person. I would just urge people to choose a position where they are making measurable positive impact in the near term (e.g. O(1)years) horizon, since the future beyond that is extremely uncertain.