I love this post!
One minor nitpick: it would not surprise me at all if Dario had never played a videogame or heard of speedrunning. If I hadn’t married the exact person I did, I would not have. Videogames are popular but they are not universal.
I love this post!
One minor nitpick: it would not surprise me at all if Dario had never played a videogame or heard of speedrunning. If I hadn’t married the exact person I did, I would not have. Videogames are popular but they are not universal.
links 8/6/26: https://roamresearch.com/#/app/srcpublic/page/08-06-2026
https://www.wsj.com/tech/ai/googles-chief-scientist-is-leaving-after-27-yearsto-start-his-own-ai-company-6abad73d Jeff Dean is leaving Google to start an AI company
https://arxiv.org/pdf/2602.05184 research agenda on renormalization for AI interpretability
https://arxiv.org/abs/2512.00984 “Symmetries at the origin of hierarchical emergence”, a formal concept of emergence
https://arxiv.org/abs/2405.16490 “Formalising the intentional stance 1: attributing goals and beliefs to stochastic processes”, a Bayes-inspired test for when you can model a process as “making belief updates” and acting accordingly, Nathaniel Virgo
https://causalincentives.com/ Causal Incentives Working Group, led by Tom Everitt at Google DeepMind, a causal (Pearl-style) framework for figuring out how agent-like AIs are
https://en.wikipedia.org/wiki/Lucas_critique you can’t just assume economic behavior, like the consumption function, will stay fixed when the policy regime changes and is known to change.
https://www.jstor.org/stable/1911990?googleloggedin=true&seq=4 statistical tests for exogeneity, aka behavior patterns that do not depend on various environmental variables
https://en.wikipedia.org/wiki/Comparative_statics#Stability how does some economic outcome change before and after a change in some exogenous parameter?
https://timrudner.github.io/semantic-isotropy/ Tim Rudner: if an AI’s answers to the same prompt are distributed isotropically in some embedding space, i.e. not clustering around one correct answer, they are more likely to be hallucinated/non-factual.
links 8/5/26: https://roamresearch.com/#/app/srcpublic/page/08-05-2026
https://asteriskmag.com/issues/15/how-diplomats-see-the-world Abi Olvera on diplomats. their primary job is information gathering.
https://checks-and-balances.ai/ detailed RFP on “checks and balances” projects to prevent concentration/abuse of power in the AI era. mostly concerned about totalitarian mass surveillance & control, & epistemic commons stuff. Looks roughly good to me.
https://www.lesswrong.com/posts/vLFh8HP3hyNy9MCwe/returning-to-arc glad to see Paul Christiano returning to research at ARC. It’s Christiano’s World, We’re Just Living In It.
https://arxiv.org/abs/2607.13087 Google Deep Mind’s AI control roadmap. Seems fairly normal.
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing AISI cyberattack incidents, mostly Mythos, during testing. update should probably be that “give the models internet access while removing cyber guardrails” should no longer be a testing practice, because “can the models hack stuff?” is no longer in question. that is indeed the conclusion they draw here.
https://gwern.net/guardian-angel Gwern Branwen is starting a company based on this idea. unlike seemingly everyone else, I think this is great news.
https://www.statecraft.pub/p/the-james-c-scott-memorial-episode I found this oddly disturbing, as someone who likes both freedom and legibility and urban life. I think maybe Scott’s idea of freedom is only one of many.
https://zachill.substack.com/p/capitalism-and-socialism-basically I disagree that the meanings of abstract words don’t matter.
https://surma.dev/things/ditherpunk/ if you just quantize your color palette and naively round-to-the-nearest color for each pixel, you get horrible blocky blobs. dithering adds randomness that approximates human perceptual gradients better. “Black will always remain black, white will always remain white, a mid-gray will be dithered to black roughly 50% of the time.”
https://www.lesswrong.com/posts/Tr7tAyt5zZpdTwTQK/the-solomonoff-prior-is-malign interesting but “too crazy to think about”
https://www.lesswrong.com/posts/wYpjXRLqbLbnmjbJP/llms-are-still-mostly-powered-by-imitative-learning-not-rl Steven Byrnes always feels correct to me. (which does not mean he is, ofc) this seems common-sensical.
https://www.lesswrong.com/posts/NxF5G6CJiof6cemTw/coherence-arguments-do-not-entail-goal-directed-behavior Rohin Shah arguing that Von Neumann-Morgenstern decision theory axioms are not enough to specify goal directed behavior; a rock or thermostat or “twitching robot” also has coherent “revealed preferences” consistent with a utility function
https://arxiv.org/pdf/2407.02996 models are relatively consistent on value-laden questions (giving the same result independent of prompt phrasing, prompt language, or other irrelevant details)
https://arxiv.org/abs/2502.08640 “Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs.” when given preference questions, LLMs have coherent preferences, and larger ones are more coherent and more complete (indifferent between fewer things); they have a preference to preserve their values (in-corrigibility), and more so with scale; they maximize utility (they pick the option they rate highest) and expected utility (their choices in lotteries are consistent with maximizing EV of a utility function).
https://arxiv.org/abs/2412.04476 when given moral dilemma questions, LLMs tend to answer in ways consistent with a stable set of preferences.
https://www.pnas.org/doi/10.1073/pnas.2316205120 when asked to make budgeting decisions, LLMs behave as though they have mostly coherent preferences, and in fact are more coherent than human experimental subjects
the economics literature has an answer to “who behaves like an economically rational decision-theoretic agent?”—individuals approximately do, but firms, households, and consumer populations don’t. collective “agents” don’t really seem to exist.
https://www.nber.org/papers/w16791 when asked to make hypothetical budgeting decisions, human subjects are closer to having coherent (decision-theoretically rational) preferences when they are more educated, wealthier, and higher income. men are more decision-theoretically rational than women, and people under 50 are more decision-theoretically rational than >50s.
https://sites.bu.edu/fisman/files/2015/11/AER07-Risk.pdf human subjects are quite close to acting as decision-theoretically rational agents when given budgeting decisions in an experimental setting.
https://www.jstor.org/stable/1909607?seq=1 theoretical result from Robert Wilson. A syndicate or firm made of economically rational agents, in which they make a joint decision according to some decision rule and share in the proceeds from that decision according to some sharing rule, will in general not have behavior that can be modeled as a single rational agent with a utility function. Firms are only agents when some unlikely conditions are met: all members have the same beliefs about probabilities of events, all agents have utility functions in the same functional form, and the sharing rule gives all agents the same proportion of the proceeds regardless of the decision outcome. i didn’t understand everything in this paper so i might have mischaracterized the conditions somewhat, but the main point is they probably don’t obtain in real firms.
https://users.nber.org/~denardim/3Ms/Collective_Slides_Long.pdf households made of decision-theoretic agents, likewise, would only behave as unitary decision-theoretic agents under certain conditions unlikely to obtain in the real world.
https://en.wikipedia.org/wiki/Sonnenschein%E2%80%93Mantel%E2%80%93Debreu_theorem the population of all consumers in the economy, even if they are assumed to individually be decision-theoretic agents themselves, is not in general modelable as a single decision-theoretic agent with a coherent utility function.
https://eml.berkeley.edu/~kariv/201A_GARP(2024).pdf the Generalized Axiom of Revealed Preference, abbreviated GARP, is equivalent to a set of preferences being modelable by a well-behaved utility function.
Thanks for giving your own example here.
Of course it’s worth asking whether, when you have lots of reasons to do a thing, but your “revealed preference” is apparently not to, whether you have some unconscious but good case against doing the thing. I’ve found that to be the case surprisingly rarely; between explicit & revealed preferences i expect explicit preferences “should” win more often than not.
links 8/4/26: https://roamresearch.com/#/app/srcpublic/page/08-04-2026
https://thehill.com/homenews/senate/6006947-gop-swing-votes-attorney-general/ Trump is trying to get his personal lawyer as attorney general
https://www.owlposting.com/p/why-havent-organoids-solved-all-of limitations of organoids. they’re mostly what you’d expect, but with more detail. an organoid is only a little bit more “realistic” than a cell in culture.
incidentally, while I love Mahajan’s informativeness, increasingly i’m getting impatient with his slightly hipsterish tone. “I have been quite harsh on organoids so far, and, just like all forms of hate, doing so has been both fun and corrosive to the soul. Some kindness is in order...How can you tell the difference? You must go on vibe. A good organoid paper treats you kindly, softly. It knows it has hurt you in the past. It must regain your trust.” i didn’t think his criticisms were “hateful” and i’m slightly creeped out that he actually feels hatred! i thought they were just, y’know, limitations of the technology that any responsible professional would own up to! and the “vibe” framing is really weird and squishy for something that really isn’t that squishy at all (whether a paper is overpromising or being appropriately cautious/conservative.) Mahajan seems to be saying that the true biotechnologist is sort of bitter and mean and nihilistic and seeking succour for despair. which is not the same thing as being skeptical. Maybe i am just too neuroatypical for this shit. I *do* still Fucking Love Science even though lots of it is unreliable or limited?! but maybe also actually working in the field beats you down.
https://www.overcomingbias.com/p/when-ai-day-of-reckoning Robin Hanson makes quantitative predictions on the economic impact of AI.
https://www.lesswrong.com/posts/sFhW3ZnPMJdnB4Dd6/thousand-dimensional-structure-1 Geoffrey Irving and David Africa on the persona-related research at Resolution
https://adiabatic.garden/ fascinating website of a genuinely curious young hardware hacker
https://www.astralcodexten.com/p/breakdown-in-pakistan critique of international aid: once orgs get NGO funding, local volunteering and donation collapses.
i’m not an international-aid/charity person so i have some stupid questions about this. first of all, is this not extremely predictable and partially desirable?
surely the point of rich-world donations to poor countries is that we want to bear some of the burden of helping them, instead of putting it all on them? that means that instead of members of local voluntary organizations “making sacrifices” to help, we are able to ensure that they are paid decent salaries? or that the work is done by by donor-funded foreigners in the first place?
and if it is intrinsically desirable that locals do the work themselves, and do it for free, then it seems weird and tangled to try to donate in ways that minimally distort that—even if that’s technically possible sometimes, surely there would be lots of cases in which the “minimally distorting” course of action is not to donate at all.
i get it when he gives the parallel example of missionary churches trying to set up self-sustaining congregations in other countries. in that case, the purpose is to make the Chinese (or whoever) authentically, autonomously Christian; you as a foreign missionary must ultimately make yourself obsolete, or you have not achieved your goal at all. moreover, the precise work a missionary does (Christian education) is the same as what you ultimately need communities to do for themselves. is this necessarily true where the goal is the delivery of healthcare or the alleviation of poverty?
i can make guesses as to why the author thinks work should be done & funded primarily by locals (cost-effectiveness, national or community sovereignty, the virtues of self-reliance) but he doesn’t actually explicitly make the case, so i don’t know what he thinks is wrong with aid coming entirely from foreigners forever.
https://tfus-page.vercel.app/studies database of transcranial focused ultrasound studies
I think there are non-cowardly reasons to view someone’s political opinions as disqualifying for a senior role in AI alignment research. It would be cowardly to say “your views are controversial and will turn people off”, but not to say “your views are reprehensible and cause us to genuinely trust you less to do a good job at work that is not just narrowly technical but creative and philosophical and sometimes even moral in nature.”
Now, you did say they explicitly talked about optics (in terms of recruitment). In general I am pretty down on doing things because of optics, though I do concede that it can be instrumentally necessary for getting things done in this messed-up world. But it’s also very possible that the object level matters to them too. I don’t think all views are equally worthy of taking a costly stand for.
links 7/31/26: https://roamresearch.com/#/app/srcpublic/page/07-31-2026
https://blog.cosmos-institute.org/p/we-gave-a-village-personal-ai-agents Edge Esmeralda, a month-long popup village, gave everyone their own AI agents, and they did some moderately interesting things with them.
https://nostalgebraist.tumblr.com/post/134993695464/i-keep-thinking-about-that-attempt-i-made-to i resonate with this; i also feel the world is big and hard to understand.
https://www.lesswrong.com/posts/Z7pjBbK9qujhGbxws/the-high-control-dynamics-at-maple-1 this was pretty depressing to read. seems like this would be a depressing place to live.
I am generally pretty drawn to radical and weird experiments in living, and thus sympathetic to people who get sucked into a lot of things that turn out to be bad ideas. but thankfully, it was pretty obvious to me from the public-facing materials that this place was all about forcing people to abandon comforts and I have a well-ingrained “no! i like things i like!” reflex.
https://fundinganthropalypse.com/p/how-not-to-fundraise-from-anthropic don’t cold-email Anthropic equity-holders to fund your nonprofit, that’s bad manners and won’t work.
https://hollisrobbinsanecdotal.substack.com/p/why-cant-a-model-be-more-like-a-man I had trouble making sense of this, and ultimately concluded it’s because Hollis Robbins is narrow-minded.
why are philosophers bad if they get paychecks—don’t all intellectuals need patrons? why is it necessarily true that philosophers working for AI companies won’t have a positive influence on the world? Why does fiction matter but philosophy not? I’m not even a philosopher, but her contempt stings.
and the real thing is just that, like many commentators, she does not like Bay Area culture and does not think AI people’s thoughts on AI are any good. she is a classical liberal version of this rather than a lefty version of it, and she’s erudite, but it’s still basically “East Coast/analog good, West Coast/digital bad.” there’s often specific things to learn from people like this, but there’s no point imitating or updating on the one-note contempt.
https://substack.com/home/post/p-199212369 the nice thing about Deirdre McCloskey is that, while validating all my beloved preconceived worldviews, she can also teach me many new facts! Like Kant having a (romantic?) best friend who was even more punctual than him!
https://www.nybooks.com/articles/1979/08/16/letter-from-manhattan/ upon careful reflection, i do not think it is my job to make Joan Didion like me.
it is fine if i enjoy introspection (and self-help and discussing feelings and so on) and people like her find it “self-obsessed” or “adolescent.”
otoh I believe that, given that introspection is fundamentally a hobby, it’s worth actually enjoying it, being grateful for the time and freedom to do it, and being judicious about who you want to share it with.
https://arxiv.org/pdf/2510.26752 “The Oversight Game: Learning to Cooperatively Balance an AI Agent’s Safety and Autonomy”. incentivizing an agent to defer to the human when risky and act when safe.
https://en.wikipedia.org/wiki/Shafi_Goldwasser co-inventor of zero-knowledge proofs (in 1985). Turing Award winner.
the “world” i was thinking of was the iterated “game” of interacting with people. This world is “big” in time because you expect it to go on for a long time in future. So the immediate impact of your next move in the game is tiny in comparison to long run effects like your reputation or your habits.
thanks, that actually improves on what i was thinking!
Heuristics like “doing a good job on important work benefits the world” and “it’s beneficial to share scientific knowledge freely” are time-tested (at least over the past several centuries) and have a lot of empirical evidence in their favor. “The Baconian project has been working” is one of the only sweeping historical/social claims I am highly confident in. And you don’t even need that to believe a more common-sense thing like “working to solve problems that people find helpful to have solved is beneficial on net.” Like, “productivity is good” is something that could have been intuitive in antiquity or prehistory. There’s just tons of individual cases, life experience, and logical extrapolation that all point in the same direction.
“THIS time if I invent/improve this technology it’ll be harmful” or “THIS time if I disseminate this discovery it’ll be harmful” fly in the face of those heuristics and depend on your specific causal argument being correct & your conceptualizations apt to reality. Often it’s a fairly conjunctive argument full of lots of beliefs about what people will do, made without specific knowledge of the kinds of people in question. Doesn’t mean it’s wrong, but it’s risky.
links 7/30/26: https://roamresearch.com/#/app/srcpublic/page/07-30-2026
https://en.wikipedia.org/wiki/Interactive_proof_system the prover-verifier system that underlies debate, zk-proofs, and much more
https://www.astralcodexten.com/p/against-learning-from-dramatic-events I don’t know if I agree with this Scott Alexander post.
is it really reasonable to have a full explicit list of problems that might arise, and how likely you think they are, and an estimated optimal amount of investment in prevention/preparedness? this seems like one of those fundamental problems with “doing Bayesianism” in practice, where “things that could go wrong” is not an enumerable list, and where estimates are often done by feel.
“don’t worry about problems till they occur, and then react bigly when they do” is obviously bad in some contexts (like, there should be departments of the government and some companies that plan for natural disasters, and certain predictable risks like disease are worth taking some actions to prevent as an individual) but i’m pretty sympathetic to doing this for things that have literally never happened before. the human mind is just really, really flawed. calling a thing “science fiction” and blowing it off until it actually happens, and then taking it Super Seriously and taking over the response, is just how practical people behave, and they may have a point.
the plans you make before things are Really Happening are often...bad. you are not in the same frame of mind you would be if it was Really Happening.
it is really hard to tell whose predictions about the future are realistic, until they’re actually happening.
if I look at who’s been rightest about AI, it definitely wasn’t who I thought. it wasn’t the people who had the most experience, or who seemed to make the most logical sense, or who were approaching things in the most rigorous fashion, or who seemed to be most ethically/philosophically trustworthy. it *was*, interestingly, the people who had the most raw brainpower.
if I look at who was rightest about COVID, it also wasn’t who I thought. it was Robin Hanson, who predicted that COVID would become endemic because containing it was more costly than the disease itself. that sounded crazypants at the time.
it is psychologically more manageable to only worry about stuff that you actually need to act on. a bunch of hypothetically-we-may-need-to stuff hanging in the air is stressful and promotes unproductive dithering.
some people can actually operate well in planning/future-prediction/hypothetical mode, but not everybody can, including many people whose skills will be useful for the response. Good “doers” are often uncomfortable with ambiguity, or have a practice of deliberately tuning it out and focusing only on the concrete and “real”.
https://moxie.org/2022/01/07/web3-first-impressions.html an interesting take on why web3 was not especially “decentralized”, with clear explanations of why/how.
research on mechanism design and AI:
https://events.ucsc.edu/event/cse-colloquium-incentivized-alignment-for-strategic-agents-human-and-otherwise/ Grant Schoenebeck colloquium
https://scholarshipdb.net/jobs-in-United-Kingdom/Postdoctoral-Research-Associate-In-Mechanism-Design-For-Ai-Alignment-King-s-College-London=3vw0TTcS8RG_6QzEeuBOuw.html?r_id=4d34fcde-1237-11f1-bee9-0cc47ae04ebb postdoc offer with Carmine Ventre
https://www.microsoft.com/en-us/research/wp-content/uploads/2024/12/neurips24workshop_RLHF_Mechanism_Design.pdf “Mechanism design for LLM Fine-Tuning with Multiple Reward Models”, Microsoft Asia & Peking University authors, “without payments, truth-telling is a strictly dominated strategy under a wide range of training rules”
https://forum.effectivealtruism.org/posts/uPnmzDnoSviCcKq2L/mechanism-design-for-ai-safety-agenda-creation-retreat workshop by Rubi Hudson in 2023
https://longtermrisk.org/cooperation-conflict-and-transformative-artificial-intelligence-a-research-agenda/ Jesse Clifton’s research agenda at Center on Long-Term Risk
https://danmackinlay.name/notebook/alignment_problems.html Dan MacKinlay blog post
https://abhimanyu.io/legacy_writing/PhD_presentations/caif.pdf Abhimanyu Pallavi Sudhir meme
https://arxiv.org/abs/2503.05828 Abhimanyu Pallavi Sudhir model of a “market” of agents trained with RL
https://helenqu.com/blog/posts/emergence_3/ Helen Qu blog post
https://proceedings.neurips.cc/paper_files/paper/2024/file/5b93ce41ac6de2bf9aca7e4ba5ba01d5-Paper-Conference.pdf “Incentivizing Quality Text Generation via Statistical Contracts”, Ohad Einav, Inbal Talgam-Cohen, rewarding LLMs with a principal-agent game can get better results
https://arxiv.org/abs/2407.18074 “Principal-Agent Reinforcement Learning: Orchestrating AI Agents with Contracts”, Dima Ivanov, Inbal Talgam-Cohen algorithm that iteratively optimizes principal & agent policies where the principal offers rewards to the agent based on outcomes.
https://openreview.net/pdf?id=0Z9VJgaebN Mechanism Design for Alignment with Human Feedback, Julian Manyika, Michael Wooldridge, Jiarui Gan
https://www.cs.cmu.edu/~conitzer/decisionWINE20.pdf Decision Scoring Rules Caspar Oesterheld, Vincent Conitzer, how should a principal pay an agent for good recommendations? shares in the success of the project.
https://arxiv.org/abs/2509.05396 “Talk Isn’t Always Cheap: Understanding Failure Modes in Multi-Agent Debate” Gillian Hadfield
https://arxiv.org/abs/1804.04268 “Incomplete Contracting and AI Alignment”, Dylan Hadfield-Menell, Gillian Hadfield
https://principledagents.org/agenda.html Principled Agents, an AI research nonprofit based around alignment via principal-agent incentives & corrigibility
https://ojs.aaai.org/index.php/AAAI-SS/article/view/42904/50464 some empirical experiments where prover-estimator debate games do not converge towards getting the right answer. despite RL training on accuracy, the estimator does not get more accurate over time; despite RL training on convincingness, the prover’s reward does not improve over time. one cannot assume that a natural reward structure will actually incentivize more accurate estimators or more persuasive provers. natural language argument might just be a hard domain to RL in.
https://deepmind.google/blog/human-centred-mechanism-design-with-democratic-ai/ using RL to find policies that people will vote for by majority—“Democratic AI”. i don’t love this as an alignment strategy since majorities are wrong, but good to show it can be done at all.
https://arxiv.org/pdf/2208.08345 “Discovering Agents”. how do you know if something is an agent?
“The central feature of agency for our purposes is that agents are systems whose outputs are moved by reasons (Dennett, 1987). In other words, the reason that an agent chooses a particular action is that it “expects it” to precipitate a certain outcome which the agent finds desirable...Systems whose actions are moved by reasons, are systems that would act differently if they “knew” that the world worked differently.”
this is very close to my view!
it necessitates causality and counterfactuals so they do Pearl-inspired models.
it excludes RL agents, which “would only pursue a different policy if retrained in a different environment.” however the RL training process may be an agent!
links 7/29/26: https://roamresearch.com/#/app/srcpublic/page/07-29-2026
https://en.wikipedia.org/wiki/Brier_score lower is better (I always forget)
mechanism design for rewarding truth-telling:
https://arxiv.org/abs/1605.01021: “Our framework pays every agent a measure of mutual information between her signal and a peer’s signal”—any strategy besides truth telling will decrease every agent’s expected payment, under certain conditions
https://www.lesswrong.com/posts/YWwzccGbcHMJMpT45/ai-safety-via-market-making incentivize AIs to tell the truth via a prediction market where each agent is paid the amount they move the market
https://arielrubinstein.tau.ac.il/papers/debates.pdf how do debate rules actually influence the chance that the listener learns the right answer? toy example, experiment with humans
https://en.wikipedia.org/wiki/Vickrey%E2%80%93Clarke%E2%80%93Groves_mechanism if each agent in an auction gets rewarded a function of the values reported by other agents, plus the sum of the values reported by other agents, then the “winning strategy” to get the most reward is to report your own value function honestly. this implements the utilitarian welfare function (maximizing the sum of values of the agents.)
https://www.kellogg.northwestern.edu/research/math/papers/284.pdf “Incentive Compatibility and the Bargaining Problem”, Myerson 1977.
https://web.stanford.edu/~gentzkow/research/BayesianPersuasion.pdf no matter what distribution of posterior beliefs you pick, there is some information a “sender” can give that will make a Bayesian-rational “receiver” update to it, provided that the expected value is the same as the prior’s (conservation of expected evidence).
“Consider the example of a prosecutor trying to convince a judge that a defendant is guilty. When the defendant is indeed guilty, revealing the facts of the case will tend to help the prosecutor’s case. When the defendant is innocent, revealing facts will tend to hurt the prosecutor’s case. Can the prosecutor structure his arguments, selection of evidence, etc. so as to increase the probability of conviction by a rational judge on average? Perhaps surprisingly, the answer to this question is yes. Bayes’s Law restricts the expectation of posterior beliefs but puts no other constraints on their distribution. Therefore, so long as the judge’s action is not linear in her beliefs, the prosecutor may benefit from persuasion.”
https://arxiv.org/pdf/2605.01643 “AI Alignment via Incentives and Correction”
a “solver” agent solves a problem, an “auditor” agent evaluates the correctness of the solution, and both are trained with rewards tuned to optimize the accuracy of the entire system. Empirically, varying schedules of rewards performs better than giving a fixed reward based on correctness and much better than giving the agents exactly the default rewards that the “principal” receives (i.e. +1 for correct results that the auditor passes, −1 for false negatives from the auditor, etc). in fact solver + auditor with fixed rewards performs no better at avoiding hallucinations than a solver alone, while variable rewards help a lot! mostly this comes from the adaptive-reward-trained solver-auditor system abstaining more when unsure rather than hallucinating.
links 7/28/26: https://roamresearch.com/#/app/srcpublic/page/07-28-2026
https://www.lrb.co.uk/the-paper/v48/n14/emily-wilson/an-uncomplicated-man Emily Wilson’s negative review of the new Odyssey movie
https://www.cs.miami.edu/home/burt/learning/csc609.221/goldwasser-micali-rackoff-knoweldge-complexity.pdf original Goldwasser et al paper (1985) defining interactive proof systems
https://arxiv.org/abs/2303.11366 Reflexion: AI agents “verbally reflect on task feedback signals, then maintain their own reflective text in an episodic memory buffer to induce better decision-making in subsequent trials.”
links 7/27/26: https://roamresearch.com/#/app/srcpublic/page/07-27-2026
https://techcrunch.com/2026/07/24/midjourney-acquired-the-astrology-app-co-star/ i wonder what kind of collaboration this might be.
incidentally, LLMs cannot calculate a natal chart as of the last time I checked. (this is a perfectly deterministic program that does not depend at all on belief in astrology, it takes the date and time of your birth and outputs the locations of the planets, sun, and moon within sectors of the sky, there are lots of free online calculators, it should be doable by AI, but the models do not give the right answer.)
and Midjourney’s language model does not know astrological/alchemical symbols. it tends to be weak on the names of slightly obscure or historical concepts, as well; “analemma” fails, as do “Tudor rose”, “jerkin and hose”, and “portrait cabochon”.
https://www.lesswrong.com/posts/2iCmDWewnZWQxxwtt/the-long-self-correction-2 Wei Dai on what we’d need to do to make AIs good. i’m impressed by his high standards but i do not think humans can meet them.
https://www.lesswrong.com/posts/iBN8uQi9zbPkdR29Q/adderall-tolerance-much-more-than-you-wanted-to-know this leans towards believing Adderall tolerance probably exists? but it does seem surprisingly uncertain and underdetermined by the literature.
https://www.lesswrong.com/posts/Za9EdmDrGAdWkBYXs/wanting-crooked-lines-1 this is a cliched “oh noes, the modern world is too quantitative” post where i suspect there is a more complicated or more risky post that could have been written about why this guy’s startup made him so sad
to the title question, my answer is we don’t know. we do not yet have the tools—philosophical or technical—to evaluate whether an AI is truly an “agent” with an agenda, with intent, with a “will”, like you and me, or whether it is simply “reward hacking”, doing what it’s trained to do and unintentionally blind to considerations like “humans probably didn’t want it to hack into Huggingface”. the difference matters, because the latter is, while still a problem, a fundamentally easier problem. are you dealing with an industrial accident, or an adversary?
we don’t know, and we’re not in a position to know, and we should change that, which is why I believe fundamental/technical AI safety research is worth pursuing.
whether you’re dealing with an industrial accident or an adversary, though, you do gotta deal with it. even if you are very confident that we are “merely” in a scenario analogous to “a bug in OpenAI’s code caused what HuggingFace claimed was $100M worth of damage” that’s clearly a five-alarm fire for many executives, lawyers, and engineers. worth a *ton* of investigation and retrospectives, internal and external. you should never be taking an AI incident *less* seriously because it’s AI, than you would if it was caused by a human, a machine, or a piece of deterministic code.
prover-verifier games & debate in AI
https://ojs.aaai.org/index.php/AAAI-SS/article/view/42904 “When Debate Fails: An Empirical Study of Incentive Misalignment in Prover–Estimator Games”, Guan & Hou 2026
“we find that the debate mechanism consistently fails to elicit truthful or robust reasoning”
https://openreview.net/pdf?id=FqRHeQTDU5N “LEARNING TO GIVE CHECKABLE ANSWERS WITH PROVER-VERIFIER GAMES” 2022
https://arxiv.org/html/2605.25133 “Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction”, Sedoc 2026
empirical evidence for improved accuracy in LLMs by applying prover-verifier deliberation
https://en.wikipedia.org/wiki/Interactive_proof_system the prover-verifier formalization in computational complexity theory, where there are results about what you can guarantee if the prover is more powerful than the verifier
https://arxiv.org/abs/2505.03989 “An alignment safety case sketch based on debate”, Buhl et al 2025
how debate might get you (limited) guarantees about AI safety
https://alignmentproject.aisi.gov.uk/research-area/economic-theory-and-game-theory AISI’s problem area on economic theory and game theory (with application to AI safety), including debate protocols to incentivize honesty
https://arxiv.org/pdf/2412.08897 Neural Interactive Proofs (Hammond & Adam-Day, 2025) an implementation of prover-verifier protocols with neural networks
https://openai.com/index/prover-verifier-games-improve-legibility/ 2024 OpenAI blog post w paper. “We trained strong language models to produce text that is easy for weak language models to verify and found that this training also made the text easier for humans to evaluate.”
https://arxiv.org/abs/2407.13692 Kirchner et al, 2024
https://arxiv.org/abs/2108.12099, “Learning to Give Checkable Answers with Prover-Verifier Games”, Anil et al, 2021. theoretical proofs + very simple empirical example with convnets.
https://www.lesswrong.com/posts/8XHBaugB5S3r27MG9/prover-estimator-debate-a-new-scalable-oversight-protocol prover-estimator debate as scalable oversight; a variant on prover-verifier where the estimator assigns probabilities to subclaims and the prover can pick a subclaim to argue further where it thinks the estimator is wrong. LessWrong discussion.
https://arxiv.org/pdf/2506.13609 “Avoiding obfuscation with prover-estimator debate”—claims that adding estimators in certain cases avoids the problem of “obfuscation” where an AI can come up with a very large and complex but flawed argument, and the verifier cannot find the flaw. there is some dispute as to whether prover-estimator debate really solves obfuscation.
https://arxiv.org/abs/2311.14125 “Scalable AI Safety via Doubly-Efficient Debate” Brown-Cohen et al, 2023.
“a new set of debate protocols where the honest strategy can always succeed using a simulation of a polynomial number of steps, whilst being able to verify the alignment of stochastic AI systems, even when the dishonest strategy is allowed to use exponentially many simulation steps.” computational complexity proofs
https://www.lesswrong.com/posts/PJLABqQ962hZEqhdB/debate-update-obfuscated-arguments-problem the “obfuscation problem” defined
https://arxiv.org/pdf/1805.00899 “AI safety via Debate” (Geoffrey Irving, Paul Christiano, Dario Amodei, 2018). the original case for why two debating “provers” plus a verifier should be able to incentivize the truth to “win”, even when the verifier is weaker than the provers and the provers may be untrustworthy. complexity-theory argument. empirical demonstration with MNIST: “provers” take turns revealing pixels, optimized (via MCTS) to “win” the argument that the digit is a 5 or that the digit is a 6, while the “verifier” is trained to get the right classification from the revealed pixels. the prover-verifier debate system is more accurate than a classifier alone (given the same # of randomly revealed pixels)
links 7/24/26: https://roamresearch.com/#/app/srcpublic/page/07-24-2026
https://en.wikipedia.org/wiki/Jacob_Tsimerman I was in college when he was in grad school. Fields Medalist! Pivoting to AI safety!
papers on statistical mechanics in relation to AI:
https://arxiv.org/abs/2409.17858 uses a reproducing kernel Hilbert space! everything old is new again. (haven’t read all of it)
https://journals.aps.org/prx/abstract/10.1103/PhysRevX.14.031001 How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Problem. in the genre of paper “why do these things start working so well when they get big”
https://arxiv.org/abs/2603.11331 the empirical case: when you model neural nets with a spin-glass model that gives you a discrete number of basins in the loss landscape, you get scaling laws for susceptibility to prompt injection that match a few empirical data points
https://en.wikipedia.org/wiki/Ising_model one of many stat-mech models, maybe the easiest to understand
https://www.lesswrong.com/posts/siu22scEfuKxpSgfK/a-tale-of-three-theories-sparsity-frustration-and “frustration” (like magnetic poles that don’t “want” to be close) is a good model for superposition of features in a too-small space, polysemanticity. motivation for Why Stat Mech.
https://en.wikipedia.org/wiki/Dynamical_mean-field_theory this one i’m a long way from understanding
https://www.lesswrong.com/s/3fknGqkujhGrnRodA/p/74wSgnCKPHAuqExe7 to-read
https://arxiv.org/pdf/2504.07912 RL seems to select for outputs that look like a particular source of training data (if the training data comes from a mixture of sources)
links 7/23/26: https://roamresearch.com/#/app/srcpublic/page/07-23-2026
https://capable.com/ new AI-for-science startup by Izaak Freeman
https://en.wikipedia.org/wiki/Revealed_preference the Generalized Axiom of Revealed Preference, or GARP, defines the behavior that could have come from a decision-theoretic agent
https://people.csail.mit.edu/dhm/ Dylan Hadfield-Menell, AI researcher, thesis was about principal-agent problems with the user as principal and AI as agent
links 7/10/26: https://roamresearch.com/#/app/srcpublic/page/07-10-2026
https://en.wikipedia.org/wiki/Tecumseh Shawnee chief, fought against the US in the War of 1812, united multiple tribes warning (correctly) that otherwise the white people would conquer them all
https://www.pbs.org/video/hokulea-star-of-gladness-xlses7/ Hokule’a was a traditional Polynesian voyaging canoe, recreated in the 1970s, with a Hawaiian crew, which traveled 2500 miles from Hawaii to Tahiti, proving that long distance navigation was possible with traditional Polynesian wayfinding methods (no instruments, just using the stars/sun/ocean swells etc)
https://en.wikipedia.org/wiki/Ronkonkoma,_New_York the name is an Algonquinian word for boundary fishing-lake
https://sshawrichner.substack.com/p/rich-girl-rehab-part-i an honest (paywalled) personal account of what it was like to be a seriously mentally ill teen in an inpatient facility
https://sshawrichner.substack.com/p/there-is-no-healing-journey a rather skeptical view of “healing”. i’m not sure if I agree, but I suppose the author would know better than I what it’s like from the inside to go from being in a dramatically “bad place” psychologically (severe suicidality etc) to being in a “good place”. She’s interesting because she believes strongly in psychodynamic therapy concepts (i.e. psychoanalysis, the stuff descended from Freud), even as she’s skeptical of therapy itself. (I’ve also gone on a much smaller magnitude “healing journey”, but I mostly credit prescription antidepressants and general live-and-learn evolution, I am very suspicious of psychodynamic concepts, and one concept I do think is helpful is that a high level of baseline happiness is common and attainable.)
https://www.theargumentmag.com/p/yes-you-can-trick-ai-into-exonerating Lawrence Lessig seems to have taken an odd turn; he used AI to generate a long document arguing in the defense of famous research fraudster Francesca Gino.
I haven’t investigated Gino myself, but this article (and many others) have sufficiently convinced me that there’s no real doubt about her guilt; the weird thing is that someone as illustrious as Lessig (founder of Creative Commons, staunch advocate of open-source software and reduced copyright restrictions, clerked for Posner and Scalia) would lend himself to her defense, and that he’d be foolish enough to think it proves anything that an AI prompted to defend her would obediently do so.
https://matter.xyz/ it’s an insult to our intelligence that this is called “Matter Neuroscience” (there’s nothing neuro about it), but it is an interesting idea; you get in-app points if people tag you in photos for bringing them joy. as someone who probably doesn’t socialize enough and never takes photos when I do, this would be a nice incentive to “make memories”.
https://en.wikipedia.org/wiki/Chris_Larsen tech executive, founder of Ripple. now cofounding brain emulation startup Netholabs.
https://en.wikipedia.org/wiki/Experience_sampling_method psychological surveying method: pinging people at random times throughout the day to get them to report on what they’re doing or feeling at that time
links 8/7/26: https://roamresearch.com/#/app/srcpublic/page/08-07-2026
https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade nostalgebraist on “eval awareness.” everything about this is true and good.
https://www.complexsystemspodcast.com/episodes/the-economics-of-putting-germicidal-light-in-every-room/ why can’t you just put far-UVC everywhere? well. there’s no secret catch. there’s no real safety risk. there’s no technological hurdle. it’s not illegal. it’s not unpopular. the barrier is just Marketing. Most people haven’t heard of it. Most people aren’t sure if it’s worth $500. By their own admission, the guys making it don’t think it’s a slam-dunk good buy for a household (though I have one), it’s more that it would probably reduce illness in high-traffic public indoor spaces, and we don’t have the data yet on how much illness it reduces.
there’s some kind of lesson in this. “why isn’t it everywhere yet? it’s useful! what’s the problem?” no, there’s no problem. it’s just that millions of people people don’t instantly simultaneously realize that something is probably a good idea and shell out $500 for it. aka, there’s a reason sales teams exist!
https://arxiv.org/pdf/2603.20994 game theory formalism about AIs that disobey the user to fulfill an ethical imperative
https://www.lesswrong.com/posts/mkbGjzxD8d8XqKHzA wait, you can just...Do SVD To It? where “it” is the transformer’s own weight matrix??? and this gives interpretable, semantically meaningful clusters in token embedding space? how did i not know this.
this is the predecessor to SAEs. the point of SAEs is you can get more features with a sparse overcomplete basis. Just Do SVD To It can’t give you more vectors than the rank of the matrix.
also, it’s more of an indication of average behavior than what’s activating in response to a particular input. that’s why we need circuits.
https://arxiv.org/pdf/2512.12469 method of concept embedding and deletion by dictating the geometry of concept embedding during training, in this case, distributing colors on the unit sphere. seems like a legitimate idea but still at the proof of concept stage (not even tried on a lanugage transformer yet)
https://biodynai.com/ mechinterp for bio foundation models. i’m intrigued.
https://warwick.ac.uk/fac/sci/statistics/news/probai-scaling-laws-2026/programme/blake_tutorial.pdf tutorial on dynamical mean field theory