is it really reasonable to have a full explicit list of problems that might arise, and how likely you think they are, and an estimated optimal amount of investment in prevention/preparedness? this seems like one of those fundamental problems with “doing Bayesianism” in practice, where “things that could go wrong” is not an enumerable list, and where estimates are often done by feel.
“don’t worry about problems till they occur, and then react bigly when they do” is obviously bad in some contexts (like, there should be departments of the government and some companies that plan for natural disasters, and certain predictable risks like disease are worth taking some actions to prevent as an individual) but i’m pretty sympathetic to doing this for things that have literally never happened before. the human mind is just really, really flawed. calling a thing “science fiction” and blowing it off until it actually happens, and then taking it Super Seriously and taking over the response, is just how practical people behave, and they may have a point.
the plans you make before things are Really Happening are often...bad. you are not in the same frame of mind you would be if it was Really Happening.
it is really hard to tell whose predictions about the future are realistic, until they’re actually happening.
if I look at who’s been rightest about AI, it definitely wasn’t who I thought. it wasn’t the people who had the most experience, or who seemed to make the most logical sense, or who were approaching things in the most rigorous fashion, or who seemed to be most ethically/philosophically trustworthy. it *was*, interestingly, the people who had the most raw brainpower.
if I look at who was rightest about COVID, it also wasn’t who I thought. it was Robin Hanson, who predicted that COVID would become endemic because containing it was more costly than the disease itself. that sounded crazypants at the time.
it is psychologically more manageable to only worry about stuff that you actually need to act on. a bunch of hypothetically-we-may-need-to stuff hanging in the air is stressful and promotes unproductive dithering.
some people can actually operate well in planning/future-prediction/hypothetical mode, but not everybody can, including many people whose skills will be useful for the response. Good “doers” are often uncomfortable with ambiguity, or have a practice of deliberately tuning it out and focusing only on the concrete and “real”.
https://arxiv.org/abs/2407.18074 “Principal-Agent Reinforcement Learning: Orchestrating AI Agents with Contracts”, Dima Ivanov, Inbal Talgam-Cohen algorithm that iteratively optimizes principal & agent policies where the principal offers rewards to the agent based on outcomes.
https://ojs.aaai.org/index.php/AAAI-SS/article/view/42904/50464 some empirical experiments where prover-estimator debate games do not converge towards getting the right answer. despite RL training on accuracy, the estimator does not get more accurate over time; despite RL training on convincingness, the prover’s reward does not improve over time. one cannot assume that a natural reward structure will actually incentivize more accurate estimators or more persuasive provers. natural language argument might just be a hard domain to RL in.
“The central feature of agency for our purposes is that agents are systems whose outputs are moved by reasons (Dennett, 1987). In other words, the reason that an agent chooses a particular action is that it “expects it” to precipitate a certain outcome which the agent finds desirable...Systems whose actions are moved by reasons, are systems that would act differently if they “knew” that the world worked differently.”
this is very close to my view!
it necessitates causality and counterfactuals so they do Pearl-inspired models.
it excludes RL agents, which “would only pursue a different policy if retrained in a different environment.” however the RL training process may be an agent!
links 7/30/26: https://roamresearch.com/#/app/srcpublic/page/07-30-2026
https://en.wikipedia.org/wiki/Interactive_proof_system the prover-verifier system that underlies debate, zk-proofs, and much more
https://www.astralcodexten.com/p/against-learning-from-dramatic-events I don’t know if I agree with this Scott Alexander post.
is it really reasonable to have a full explicit list of problems that might arise, and how likely you think they are, and an estimated optimal amount of investment in prevention/preparedness? this seems like one of those fundamental problems with “doing Bayesianism” in practice, where “things that could go wrong” is not an enumerable list, and where estimates are often done by feel.
“don’t worry about problems till they occur, and then react bigly when they do” is obviously bad in some contexts (like, there should be departments of the government and some companies that plan for natural disasters, and certain predictable risks like disease are worth taking some actions to prevent as an individual) but i’m pretty sympathetic to doing this for things that have literally never happened before. the human mind is just really, really flawed. calling a thing “science fiction” and blowing it off until it actually happens, and then taking it Super Seriously and taking over the response, is just how practical people behave, and they may have a point.
the plans you make before things are Really Happening are often...bad. you are not in the same frame of mind you would be if it was Really Happening.
it is really hard to tell whose predictions about the future are realistic, until they’re actually happening.
if I look at who’s been rightest about AI, it definitely wasn’t who I thought. it wasn’t the people who had the most experience, or who seemed to make the most logical sense, or who were approaching things in the most rigorous fashion, or who seemed to be most ethically/philosophically trustworthy. it *was*, interestingly, the people who had the most raw brainpower.
if I look at who was rightest about COVID, it also wasn’t who I thought. it was Robin Hanson, who predicted that COVID would become endemic because containing it was more costly than the disease itself. that sounded crazypants at the time.
it is psychologically more manageable to only worry about stuff that you actually need to act on. a bunch of hypothetically-we-may-need-to stuff hanging in the air is stressful and promotes unproductive dithering.
some people can actually operate well in planning/future-prediction/hypothetical mode, but not everybody can, including many people whose skills will be useful for the response. Good “doers” are often uncomfortable with ambiguity, or have a practice of deliberately tuning it out and focusing only on the concrete and “real”.
https://moxie.org/2022/01/07/web3-first-impressions.html an interesting take on why web3 was not especially “decentralized”, with clear explanations of why/how.
research on mechanism design and AI:
https://events.ucsc.edu/event/cse-colloquium-incentivized-alignment-for-strategic-agents-human-and-otherwise/ Grant Schoenebeck colloquium
https://scholarshipdb.net/jobs-in-United-Kingdom/Postdoctoral-Research-Associate-In-Mechanism-Design-For-Ai-Alignment-King-s-College-London=3vw0TTcS8RG_6QzEeuBOuw.html?r_id=4d34fcde-1237-11f1-bee9-0cc47ae04ebb postdoc offer with Carmine Ventre
https://www.microsoft.com/en-us/research/wp-content/uploads/2024/12/neurips24workshop_RLHF_Mechanism_Design.pdf “Mechanism design for LLM Fine-Tuning with Multiple Reward Models”, Microsoft Asia & Peking University authors, “without payments, truth-telling is a strictly dominated strategy under a wide range of training rules”
https://forum.effectivealtruism.org/posts/uPnmzDnoSviCcKq2L/mechanism-design-for-ai-safety-agenda-creation-retreat workshop by Rubi Hudson in 2023
https://longtermrisk.org/cooperation-conflict-and-transformative-artificial-intelligence-a-research-agenda/ Jesse Clifton’s research agenda at Center on Long-Term Risk
https://danmackinlay.name/notebook/alignment_problems.html Dan MacKinlay blog post
https://abhimanyu.io/legacy_writing/PhD_presentations/caif.pdf Abhimanyu Pallavi Sudhir meme
https://arxiv.org/abs/2503.05828 Abhimanyu Pallavi Sudhir model of a “market” of agents trained with RL
https://helenqu.com/blog/posts/emergence_3/ Helen Qu blog post
https://proceedings.neurips.cc/paper_files/paper/2024/file/5b93ce41ac6de2bf9aca7e4ba5ba01d5-Paper-Conference.pdf “Incentivizing Quality Text Generation via Statistical Contracts”, Ohad Einav, Inbal Talgam-Cohen, rewarding LLMs with a principal-agent game can get better results
https://arxiv.org/abs/2407.18074 “Principal-Agent Reinforcement Learning: Orchestrating AI Agents with Contracts”, Dima Ivanov, Inbal Talgam-Cohen algorithm that iteratively optimizes principal & agent policies where the principal offers rewards to the agent based on outcomes.
https://openreview.net/pdf?id=0Z9VJgaebN Mechanism Design for Alignment with Human Feedback, Julian Manyika, Michael Wooldridge, Jiarui Gan
https://www.cs.cmu.edu/~conitzer/decisionWINE20.pdf Decision Scoring Rules Caspar Oesterheld, Vincent Conitzer, how should a principal pay an agent for good recommendations? shares in the success of the project.
https://arxiv.org/abs/2509.05396 “Talk Isn’t Always Cheap: Understanding Failure Modes in Multi-Agent Debate” Gillian Hadfield
https://arxiv.org/abs/1804.04268 “Incomplete Contracting and AI Alignment”, Dylan Hadfield-Menell, Gillian Hadfield
https://principledagents.org/agenda.html Principled Agents, an AI research nonprofit based around alignment via principal-agent incentives & corrigibility
https://ojs.aaai.org/index.php/AAAI-SS/article/view/42904/50464 some empirical experiments where prover-estimator debate games do not converge towards getting the right answer. despite RL training on accuracy, the estimator does not get more accurate over time; despite RL training on convincingness, the prover’s reward does not improve over time. one cannot assume that a natural reward structure will actually incentivize more accurate estimators or more persuasive provers. natural language argument might just be a hard domain to RL in.
https://deepmind.google/blog/human-centred-mechanism-design-with-democratic-ai/ using RL to find policies that people will vote for by majority—“Democratic AI”. i don’t love this as an alignment strategy since majorities are wrong, but good to show it can be done at all.
https://arxiv.org/pdf/2208.08345 “Discovering Agents”. how do you know if something is an agent?
“The central feature of agency for our purposes is that agents are systems whose outputs are moved by reasons (Dennett, 1987). In other words, the reason that an agent chooses a particular action is that it “expects it” to precipitate a certain outcome which the agent finds desirable...Systems whose actions are moved by reasons, are systems that would act differently if they “knew” that the world worked differently.”
this is very close to my view!
it necessitates causality and counterfactuals so they do Pearl-inspired models.
it excludes RL agents, which “would only pursue a different policy if retrained in a different environment.” however the RL training process may be an agent!