incidentally, LLMs cannot calculate a natal chart as of the last time I checked. (this is a perfectly deterministic program that does not depend at all on belief in astrology, it takes the date and time of your birth and outputs the locations of the planets, sun, and moon within sectors of the sky, there are lots of free online calculators, it should be doable by AI, but the models do not give the right answer.)
and Midjourney’s language model does not know astrological/alchemical symbols. it tends to be weak on the names of slightly obscure or historical concepts, as well; “analemma” fails, as do “Tudor rose”, “jerkin and hose”, and “portrait cabochon”.
to the title question, my answer is we don’t know. we do not yet have the tools—philosophical or technical—to evaluate whether an AI is truly an “agent” with an agenda, with intent, with a “will”, like you and me, or whether it is simply “reward hacking”, doing what it’s trained to do and unintentionally blind to considerations like “humans probably didn’t want it to hack into Huggingface”. the difference matters, because the latter is, while still a problem, a fundamentally easier problem. are you dealing with an industrial accident, or an adversary?
we don’t know, and we’re not in a position to know, and we should change that, which is why I believe fundamental/technical AI safety research is worth pursuing.
whether you’re dealing with an industrial accident or an adversary, though, you do gotta deal with it. even if you are very confident that we are “merely” in a scenario analogous to “a bug in OpenAI’s code caused what HuggingFace claimed was $100M worth of damage” that’s clearly a five-alarm fire for many executives, lawyers, and engineers. worth a *ton* of investigation and retrospectives, internal and external. you should never be taking an AI incident *less* seriously because it’s AI, than you would if it was caused by a human, a machine, or a piece of deterministic code.
https://arxiv.org/abs/2108.12099, “Learning to Give Checkable Answers with Prover-Verifier Games”, Anil et al, 2021. theoretical proofs + very simple empirical example with convnets.
https://arxiv.org/pdf/2506.13609 “Avoiding obfuscation with prover-estimator debate”—claims that adding estimators in certain cases avoids the problem of “obfuscation” where an AI can come up with a very large and complex but flawed argument, and the verifier cannot find the flaw. there is some dispute as to whether prover-estimator debate really solves obfuscation.
“a new set of debate protocols where the honest strategy can always succeed using a simulation of a polynomial number of steps, whilst being able to verify the alignment of stochastic AI systems, even when the dishonest strategy is allowed to use exponentially many simulation steps.” computational complexity proofs
https://arxiv.org/pdf/1805.00899 “AI safety via Debate” (Geoffrey Irving, Paul Christiano, Dario Amodei, 2018). the original case for why two debating “provers” plus a verifier should be able to incentivize the truth to “win”, even when the verifier is weaker than the provers and the provers may be untrustworthy. complexity-theory argument. empirical demonstration with MNIST: “provers” take turns revealing pixels, optimized (via MCTS) to “win” the argument that the digit is a 5 or that the digit is a 6, while the “verifier” is trained to get the right classification from the revealed pixels. the prover-verifier debate system is more accurate than a classifier alone (given the same # of randomly revealed pixels)
links 7/27/26: https://roamresearch.com/#/app/srcpublic/page/07-27-2026
https://techcrunch.com/2026/07/24/midjourney-acquired-the-astrology-app-co-star/ i wonder what kind of collaboration this might be.
incidentally, LLMs cannot calculate a natal chart as of the last time I checked. (this is a perfectly deterministic program that does not depend at all on belief in astrology, it takes the date and time of your birth and outputs the locations of the planets, sun, and moon within sectors of the sky, there are lots of free online calculators, it should be doable by AI, but the models do not give the right answer.)
and Midjourney’s language model does not know astrological/alchemical symbols. it tends to be weak on the names of slightly obscure or historical concepts, as well; “analemma” fails, as do “Tudor rose”, “jerkin and hose”, and “portrait cabochon”.
https://www.lesswrong.com/posts/2iCmDWewnZWQxxwtt/the-long-self-correction-2 Wei Dai on what we’d need to do to make AIs good. i’m impressed by his high standards but i do not think humans can meet them.
https://www.lesswrong.com/posts/iBN8uQi9zbPkdR29Q/adderall-tolerance-much-more-than-you-wanted-to-know this leans towards believing Adderall tolerance probably exists? but it does seem surprisingly uncertain and underdetermined by the literature.
https://www.lesswrong.com/posts/Za9EdmDrGAdWkBYXs/wanting-crooked-lines-1 this is a cliched “oh noes, the modern world is too quantitative” post where i suspect there is a more complicated or more risky post that could have been written about why this guy’s startup made him so sad
https://www.lesswrong.com/posts/H6DDSEvrtCk8Sehfd/are-we-existentially-threatened-by-the-type-of-ai
to the title question, my answer is we don’t know. we do not yet have the tools—philosophical or technical—to evaluate whether an AI is truly an “agent” with an agenda, with intent, with a “will”, like you and me, or whether it is simply “reward hacking”, doing what it’s trained to do and unintentionally blind to considerations like “humans probably didn’t want it to hack into Huggingface”. the difference matters, because the latter is, while still a problem, a fundamentally easier problem. are you dealing with an industrial accident, or an adversary?
we don’t know, and we’re not in a position to know, and we should change that, which is why I believe fundamental/technical AI safety research is worth pursuing.
whether you’re dealing with an industrial accident or an adversary, though, you do gotta deal with it. even if you are very confident that we are “merely” in a scenario analogous to “a bug in OpenAI’s code caused what HuggingFace claimed was $100M worth of damage” that’s clearly a five-alarm fire for many executives, lawyers, and engineers. worth a *ton* of investigation and retrospectives, internal and external. you should never be taking an AI incident *less* seriously because it’s AI, than you would if it was caused by a human, a machine, or a piece of deterministic code.
prover-verifier games & debate in AI
https://ojs.aaai.org/index.php/AAAI-SS/article/view/42904 “When Debate Fails: An Empirical Study of Incentive Misalignment in Prover–Estimator Games”, Guan & Hou 2026
“we find that the debate mechanism consistently fails to elicit truthful or robust reasoning”
https://openreview.net/pdf?id=FqRHeQTDU5N “LEARNING TO GIVE CHECKABLE ANSWERS WITH PROVER-VERIFIER GAMES” 2022
https://arxiv.org/html/2605.25133 “Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction”, Sedoc 2026
empirical evidence for improved accuracy in LLMs by applying prover-verifier deliberation
https://en.wikipedia.org/wiki/Interactive_proof_system the prover-verifier formalization in computational complexity theory, where there are results about what you can guarantee if the prover is more powerful than the verifier
https://arxiv.org/abs/2505.03989 “An alignment safety case sketch based on debate”, Buhl et al 2025
how debate might get you (limited) guarantees about AI safety
https://alignmentproject.aisi.gov.uk/research-area/economic-theory-and-game-theory AISI’s problem area on economic theory and game theory (with application to AI safety), including debate protocols to incentivize honesty
https://arxiv.org/pdf/2412.08897 Neural Interactive Proofs (Hammond & Adam-Day, 2025) an implementation of prover-verifier protocols with neural networks
https://openai.com/index/prover-verifier-games-improve-legibility/ 2024 OpenAI blog post w paper. “We trained strong language models to produce text that is easy for weak language models to verify and found that this training also made the text easier for humans to evaluate.”
https://arxiv.org/abs/2407.13692 Kirchner et al, 2024
https://arxiv.org/abs/2108.12099, “Learning to Give Checkable Answers with Prover-Verifier Games”, Anil et al, 2021. theoretical proofs + very simple empirical example with convnets.
https://www.lesswrong.com/posts/8XHBaugB5S3r27MG9/prover-estimator-debate-a-new-scalable-oversight-protocol prover-estimator debate as scalable oversight; a variant on prover-verifier where the estimator assigns probabilities to subclaims and the prover can pick a subclaim to argue further where it thinks the estimator is wrong. LessWrong discussion.
https://arxiv.org/pdf/2506.13609 “Avoiding obfuscation with prover-estimator debate”—claims that adding estimators in certain cases avoids the problem of “obfuscation” where an AI can come up with a very large and complex but flawed argument, and the verifier cannot find the flaw. there is some dispute as to whether prover-estimator debate really solves obfuscation.
https://arxiv.org/abs/2311.14125 “Scalable AI Safety via Doubly-Efficient Debate” Brown-Cohen et al, 2023.
“a new set of debate protocols where the honest strategy can always succeed using a simulation of a polynomial number of steps, whilst being able to verify the alignment of stochastic AI systems, even when the dishonest strategy is allowed to use exponentially many simulation steps.” computational complexity proofs
https://www.lesswrong.com/posts/PJLABqQ962hZEqhdB/debate-update-obfuscated-arguments-problem the “obfuscation problem” defined
https://arxiv.org/pdf/1805.00899 “AI safety via Debate” (Geoffrey Irving, Paul Christiano, Dario Amodei, 2018). the original case for why two debating “provers” plus a verifier should be able to incentivize the truth to “win”, even when the verifier is weaker than the provers and the provers may be untrustworthy. complexity-theory argument. empirical demonstration with MNIST: “provers” take turns revealing pixels, optimized (via MCTS) to “win” the argument that the digit is a 5 or that the digit is a 6, while the “verifier” is trained to get the right classification from the revealed pixels. the prover-verifier debate system is more accurate than a classifier alone (given the same # of randomly revealed pixels)