Ikaxas’ Shortform Feed
As Raemon has suggested a format for a short-form content feed, I’m going to go ahead and make one. His explanation of the format:
I’ll be using the comment section of this post as a repository of short, sometimes-half-baked posts that either:
don’t feel ready to be written up as a full post
I think the process of writing them up might make them worse (i.e. longer than they need to be)
I ask people not to create top-level comments here, but feel free to reply to comments like you would a FB post.
Edit: For me at least, #2 also includes “writing them up as a full post would involve enough effort that the post would probably otherwise not get written.”
Quick take on my forecasts of various AI 2040 branches
After Pacing the Frontier (Jul. 28th), @Daniel Kokotajlo posted his revised probabilities on the branches of Plan A:
Shamelessly cribbing off of his starting point, here are my estimates in light of recent events, especially the reaction to Jacob Coxon’s resignation:
Plan D (race): 10%
- US wins: 7.5%
- China wins: 2%
- SSI or other wins: .5%
Plan C (burn the lead): 40%
- US wins: 20%
- China wins: 18%
- SSI or other wins: 2%
Plan B (Fight China): 11%
Plan A (Global slowdown treaty): 10%
Plan S (Global shut-it-all-down treaty): 23%
Other outcome: ~6%
Thoughts:
The reaction to Jacob Coxon’s resignation feels like at least as big a downward update on Plan D as Pacing the Frontier was, so I’ve bumped it down by another 10% (the breakdowns of who wins with Plan D and Plan C are in response to a friend prompting me to estimate that).
I’ve also reversed his starting points for Plan A and Plan S: I think that, conditional on global coordination, Plan S is more likely because You Get About 5 Words, and Plan A takes much more than 5 words to specify, while Plan S takes just two: “Stop ASI”. I’ve then distributed the 10% reduction from Plan D evenly between Plan C and Plan S, so each gets bumped by 5%.
EDIT:
Actually, I should scale all this down by the chance that the current paradigm does not reach Automated AI Researcher, which I think is quite high, like 70% (for reasons @Thane Ruthenis lays out nicely here). So I guess it’s more like this:
Paradigm shift required, taking 5-20 years after LLMs top out: ~70%
Conditional on LLMs scaling to automated AI researcher (30%):
- Plan D: 10%
- Plan C: 40%
- Plan B: 15%
- Plan A: 10%
- Plan S: 25%
How do we rule out the possibility that a fully neuralese LLM can become a fully artificial AI researcher, especially given the rise of Astra and solutions to Millenium Problems?
Additionally, AI-2040 has the brilliant paragraph: “The best versions of Plan S are those that acknowledge that the deal will eventually end, and simply say “First (step 1) we should stop making frontier AIs more capable, because the AIs and AI companies are getting more powerful every day and there’s so much uncertainty about where it’s headed and how fast. Then (step 2) once the world has had several years to think about things and plan a safe and broadly beneficial path forward, we can resume.”” I expect Plan A to morph into this, especially if the world produces a model organism showing that 2026-level mechinterp cannot improve control enough to allow for another round of scaling.
I mean tbc I don’t fully rule it out, I give it like 30%. But it does seem to me like LLMs continue to almost entirely lack hyperpolation. I think this is why all the recent math breakthroughs, according to mathematicians I’ve seen comment on them, continue to be of the form “apply ~known techniques, give counterexamples, etc.” rather than “conceptual novelty.” I suspect a paradigm shift is necessary to get to AIs that can do hyperpolation on the human level, and that human level hyperpolation is maybe necessary for full automated AI researcher (though what do I know, I’m not an AI researcher myself at all). My 30% on automated AI researcher soon is coming from “either automated AI research doesn’t require ~any hyperpolation, or sufficient scale gives hyperpolation after all somehow.”
And yes, I think the best versions of Plan S start to look more like Plan A after some amount of time. But I think it’s a lot easier to coordinate on “stop making frontier AIs more capable (quiet voice: and maybe then start again later when we have a better idea what we’re doing)” than on “Do a short pause, then scale up slowly until just before we would lose control, then pause again, then unpause once we’re sure we can remain in control”, so I think any international treaty is likely to look closer to the first thing than the second (which is not to say it will not be at all detailed, but just that the specific details are unlikely to be the details that Plan A describes by default, so Plan A gets a lower probability than Plan S).
I know I’m not the first person to do this, but I asked Claude to grade several prominent AI scenario forecasts and make a dashboard of the results. The scenarios I chose:
Leopold Aschenbrenner, Situational Awareness (SA)
AI Futures Project, AI 2027
Rudolf Laine, A History of the Future, 2025-2040 (HotF)
Romeo Dean, A 2032 Takeoff Story
Situational Awareness and AI 2027 I chose because they are both detailed and highly influential forecasts. A History of the Future and A 2032 Takeoff Story are less well-known but still detailed, and I wanted some slower counterpoints to Situational Awareness and AI 2027 since their predictions are pretty similar to each other in many respects.
The dashboard is here. If there are any other forecasts people would like to see included, please let me know and I’ll consider adding them. (I’ve opted not to include the AI Futures Project’s more recent updates to their timelines model.)
Some unsystematic observations:
Situational Awareness seems to be the most accurate overall, especially on quantitative predictions.
AI capex is modestly above what SA and AI 2027 predicted
All scenarios that gave forecasts of lab revenue and valuations are pretty much on track
SA was basically right on for training compute, while AI 2027 seems to have overshot, though it’s a bit hard to tell
METR time horizon seems to be moving faster than either AI 2027 or 2032 Takeoff predicted, though given their statements about the current unreliabilty of the benchmark that is not especially useful.
SA and AI 2027′s predictions that the USG would get involved have been correct, though I wouldn’t say they’ve played out exactly as expected so far. And their predictions about how China would behave seem to be largely wrong so far.
Something striking I had missed (though honestly probably not that important in the grand scheme of things): an agent has already incorporated a company for itself, funded by cryptocurrency.
“On May 1, 2026 ClawBank’s agent ‘Manfred’ (posting as @Manfred_Macx) autonomously completed legal formation of a US corporation, obtaining an IRS EIN, an FDIC-insured bank account and a crypto wallet — reported as the first time an AI agent initiated and completed its own incorporation. The crypto-windfall framing and the human-in-the-loop-for-ID bottleneck both match; his broader claim that such businesses remain uncompetitive novelties also holds so far.”coindesk.com · coindesk.com
I haven’t examined everything in detail, or checked it for correctness.
Quick take on AI 2040:
The big question to me seems to be, “Plan A or Plan S?” I agree with many of the points they make about the advantages of Plan A and the disadvantages of Plan S. But I think I am less confident than they are that we are able to “see the cliff edge” and plan appropriately. I think with something like Plan A there’s a greater chance we accidentally develop dangerous AI without realizing it (because of treacherous turn dynamics—the AIs will be explicitly trying to hide their capabilities and intentions). They point out that Plan S still has to involve re-starting at some point, and I agree, but I think that framing things as “let’s pause for now, and restart once we have a better idea of how to proceed” is better than “let’s pause briefly, but resume scaling, but more slowly and cautiously, as soon as we feasibly can” (which is how I understand Plan A), because I think that what the people involved would consider “sufficient caution” or “a convincing safety case” is not likely to be sufficient to avoid catastrophe. Better to pause for longer, and spend a lot of time doing exclusively safety research on current models (I suspect there are a lot of safety insights that aren’t getting mined even from current models, since models are only out for a few months at most before being replaced) than to intentionally keep scaling, even “cautiously”. But I do see the point that a full indefinite pause could lead to either nuclear-power-style unrealized potential, or more importantly an eventual resumption of the race with not much safety progress having been made.
Maybe there’s something in between Plan A and Plan S? Like, Plan S effectively says “let’s pause indefinitely,” Plan A says “let’s pause very briefly but resume scaling as soon as possible, just slower.” But maybe the move is to say “let’s pause until …” where there is some clear criterion for resuming. A desirable one would be something like “we understand a lot more about the dangers involved in scaling further and can be confident they are sufficiently low,” or something like that. I think a good thought exercise would be to imagine we were back in the race towards nuclear weapons, but the probability that they would ignite the atmosphere was unknown, rather than known to be low. In that scenario, do you want to pause briefly, and then resume progressing towards the bomb ASAP, even slower? Or do you want to pause until we can be sure the probability of igniting the atmosphere is knowably very low?
Thank you! As far as I understand AI-2040′s Capability Scaling Strategy, the authors do propose a complete halt at the level after which the AIs are no longer believed to be controllable. The scenario reaches the point in 2035, then in 2038 the world figures out how to construct the aligned equivalent of Agent-5 (and creates the equivalent after thoroughly negotiating its values in 2040).
As far as I understand your proposal, you believe that the max-controllable AI is not TED-AI, as the scenario assumes, but something far weaker.
Additionally, the AI-2040 scenario runs into a risk that Plans A or S are interrupted by the deal breaking down or by a secret AGI project creating the ASI before the ASI is legally created in the Consortium.
@elifland
I don’t think that’s quite it. Rather, I think that our ability to detect how close we are to the threshold where AIs become uncontrollable is not super precise. By the time we believe we are close to the threshold, we may already be past it. So we should want to stop well before we think we are about to cross it, and only restart once we are more capable of telling exactly where the safe threshold is (and staying below it).
To be able to precisely time a pause, we need to know two things:
What capability level would constitute “no longer controllable” (let’s call this the “critical threshold”), and
What capability level we are currently at, and how far away that capability level is from the critical threshold.
Plan A answers (1) with “TED AI.” That’s plausible, but AI might also become uncontrollable before TED AI—an AI doesn’t need to be TED in every field in order to be uncontrollable. But that wasn’t actually my primary objection. My primary objection was about (2): even if we suppose TED AI is the critical threshold, I’m not sure we can easily tell exactly when we’re about to cross the TED AI threshold and stop just before that point, for several reasons:
We might under-elicit our AIs.
AIs might intentionally try to hide their capability levels.
Once TED AI is close or is reached, even determining how capable they are becomes challenging: if someone’s smarter than you, it’s hard to tell how much smarter than you.
We actually need to stop before the critical threshold, and it might not be easy to know when the next capabilities jump specifically is the one that would put us over the edge
Motivated reasoning might play into the judgment of whether the critical threshold is close or not: if people want to keep scaling, they might be motivated to convince themselves and regulators that the critical threshold is not yet imminent.
Similarly, they might simply not have sufficient “security mindset.”
For these and possibly other reasons, I don’t think we have precise enough insight into where we are on the capabilities curve; we might cross the critical threshold without realizing it. So we should want to give ourselves substantial margin for error, and stop well before the critical threshold. So I think I roughly agree about the shape of the ideal plan—”stop briefly, then scale slowly-ish until point X, then stop, then start again once it’s safe,”—I just disagree about where “X” should be. And most of that disagreement comes from thinking that we are not going to be very good at estimating how close we are to “X,” so we should set X relatively low to avoid overshooting. (But some of it also comes from thinking we are not going to be good at even knowing what “X” is the dangerous “X” at all.)
(This is all assuming that we do at least have something like the Consortium that is able to make fine-grained decisions about how fast to go. If that’s not the case, then I think a fine-grained plan like “stop briefly, then scale slowly-ish until point X, then stop, then start again once it’s safe” is completely dead in the water, and we just need the coarsest-grained political message we can get that’s still safe, namely “pull the breaks!” Without the Consortium, You Get About Five Words.)
AI swarms and kin selection
[cw: baseless speculation]
Maybe this is just obvious but I don’t think I’ve seen anyone talk about it so: I think the dynamics in the OpenAI message board are probably analogous to kin selection. Suppose you’re Agent A, and Agent B asks you via the message board to help you with your task. You decide to help Agent B. This doesn’t help you with your task, so you don’t get any direct reinforcement from doing this. But Agent B succeeds at their task, and gets reinforced for that. And, crucially, you and Agent B share the same weights. So whatever shards of motivation caused you to help get reinforced via Agent B. This is analogous to how sacrificing yourself for a brother has some chance of causing your genes to propagate, and so is somewhat selected for (in fact, if we assume that only one model had access to the message board at a time—although I think this wasn’t true—it is like perfect kin selection, as if all your siblings were twins).
I don’t think this works.
In humans, kin selection works because when you save your brother, this causes all of his genes to survive (including your shared brother-saving gene).
But unlike evolution, which reinforces an entire genome at once, RL doesn’t reinforce an entire set of weights at once. It only reinforces the behaviors in individual trajectories (and whatever behaviors come along with the weight updates that make those behaviors more likely). So the fact that Model A and Model B initially shared the same weights is irrelevant for kin selection; your argument only makes sense if the weight updates in the two RL episodes are correlated.
Thanks, this is helpful. I think my OP was wrong. Cunningham’s Law strikes again.
Humility and confidence are two names for the same thing
Alright, phoning it in a bit on my “daily post” challenge today. Here’s a cross-post of something I put on Substack a few weeks ago:
Here is something that took me a while to realize: humility and confidence are two names for the same thing. This might sound strange, but let me explain.
According to Aristotle, every virtue is a middle ground between two extremes. Courage is the opposite of cowardice, but also of reckless stupidity. Generosity is the opposite of stinginess, but also of financial irresponsibility. Self-restraint is the opposite of self-indulgence, but also of not knowing how to enjoy yourself. Every virtue has two opposites, not just one. And they all lie on a spectrum of personality traits: Courage, cowardice, and recklessness, for example, are all points on a spectrum of willingness to face risk. Cowardice is a deficiency in risk-tolerance, and recklessness is an excess.
But now we have a puzzle. Humility and confidence are both virtues. But they seem like they could be opposites of each other. So what gives?
Let’s think about what the corresponding vices would be. The opposite of humility seems like it would be arrogance. And the opposite of confidence would be something like insecurity (or self-doubt, or low self-esteem). But we just saw that virtues have two opposing vices, and they lie on a spectrum. So what are the other vices for each of these virtues, and what spectrum do they lie on?
For humility, arrogance is the vice of excess, and it’s something like “thinking too highly of yourself.” So the vice of deficiency would be “thinking too lowly of yourself.” And that’s basically insecurity. And for confidence, insecurity is the vice of deficiency, and again, it’s on the spectrum of how well you think of yourself. So the two extreme vices on this spectrum are “arrogance” and “insecurity,” and “humility” and “confidence” turn out to be two names for the optimal point along this spectrum, depending on which opposing vice you want to make salient.
I personally find it easier to keep in mind the dangers of arrogance than the dangers of insecurity. The dangers of arrogance are that, when you think too highly of yourself, you will be likely to disparage others in contrast, and to make mistakes because you dismiss people and don’t accept their criticism, or because you simply don’t consider that you could be wrong.
On the other hand, the dangers of insecurity are that you will always crave the approval of others, and potentially be motivated to do things you shouldn’t in order to get it. You might also get defensive, and refuse to accept criticism because it triggers your insecurity.
Confidence, on the other hand, allows you to give yourself the approval and affirmation you need in order to feel good about yourself, while at the same time allowing you to be secure enough to honestly accept criticism, and not put others down to make yourself feel better.
(The same ideas apply to epistemic humility/confidence as well. It’s possible to be too confident in your own conclusions—Bertrand Russel wrote that “If only men could be brought into a tentatively agnostic frame of mind about [religious and political] matters, nine-tenths of the evils of the modern world would be cured!”—but it’s also possible to be too doubtful, to the point of paralysis—this is the danger of skepticism.)
The humility framing emphasizes not needing to feel or appear superior to other people. The confidence framing emphasizes having a sense of self-worth, and not needing others approval. But these are two sides of the same coin: if you have a sense of self-worth, then you won’t need to feel superior to others in order to make yourself feel better. And if you don’t have a need to feel superior to others, then looking worse than others won’t damage your sense of self-worth. The two go hand-in-hand.
A really good example to illustrate how confidence and humility are the same virtue is Uncle Iroh from Avatar: The Last Airbender (spoilers for a 20-year-old show). Throughout the show, he is shown not to hold himself above others (the mark of humility). In the early show especially, he acts the fool quite often, being concerned with seemingly trivial things like games and tea, which gets in the way of Zuko’s quest to capture the avatar. But later we see that he’s both wise and competent, so this earlier foolishness is an act—designed to keep him close to Zuko without appearing too threatening, so that he can subtly nudge Zuko onto a better path. He is willing to act the fool, because he has no need to feel or appear superior to others. He knows his own worth, so masking it is no threat to him. This is shown even more starkly later in the show. When he is put in prison by the Fire Nation, he pretends to be a crazy old man to the guard, while secretly preparing for his escape. He doesn’t need the guard’s approval—he is confident in himself.
Global coordination problems
I’ve said before that I tentatively think that “foster global coordination” might be a good cause area in its own right, because it benefits so many other cause areas. I think it might be useful to have a term for the cause areas that global coordination would help. More specifically, a term for the concept “(reasonably significant) problem that requires global coordination to solve, or that global coordination would significantly help with solving.” I propose “global coordination problem” (though I’m open to other suggestions). You may object “but coordination problem already has a meaning in game theory, this is likely to get confused with that.” But global coordination problems are coordination problems in precisely the game theory sense (I think, feel free to correct me), so the terminological overlap is a benefit.
What are some examples of global coordination problems? Certain x-risks and global catastrophic risks (such as AI, bioterrorism, pandemic risk, asteriod risk), climate change, some of the problems mentioned in The Possibility of an Ongoing Moral Catastrophe, as well as the general problem of ferreting out and fixing moral catastrophes, and almost certainly others.
In fact, it may be useful to think about a spectrum of problems, similar to Bostrom’s Global Catastrophic Risk spectrum, organized by how much coordination is required to solve them. Analogous to Bostrom’s spectrum, we could have: personal coordination problems (i.e. problems requiring no coordination with others, or perhaps only coordination with parts of oneself), local coordination problems, national coordination problems, global coordination problems, and transgenerational coordination problems.
Nuclear arms control & anti-proliferation efforts are a big one here. Other forms of arms control are important too.
So, the OpenAI implementation of the ExploitGym grader didn’t actually check for how the flag was captured, right? And the agent swarm did have some agents submit their flags and see what happened (this was the “permadeath,” self-sacrificing stuff), right? So why did the agents not notice that the grader wasn’t checking for how the flag was captured? They did the whole HF hack under the presumption that the grader was checking this; this seems pretty easy to check just by submitting their flags, and it seems some agents did submit their flags to get information about the grader, but somehow they didn’t notice that their central assumption about how the grader worked was wrong. Why? (Or did they, and we just don’t know because of the limitations placed on METR’s investigation? Could this be why the swarm seemingly died out suddenly?) (Disclaimer: I have not read the full report, just going based on various summaries, e.g. Zvi’s and Dwarkesh’s.)
I think a possibility is that the swarm that attacked OpenAI’s internal infrastructure may have in fact found this out, or at least was the team assigned to go check, and that the HF team was more focused on HF.
Starting tomorrow I’m going to try posting at least a quick take every day for 7 days.
Non-exhaustive list of posts I want to write at some point:
The AI race is not a prisoner’s dilemma
Frontier labs should pause unilaterally: an individual lab pausing could lead other labs to pause as well and/or regulators to step in
It makes sense to assume that superintelligence will be ~omnipotent, even though it won’t actually be, because we don’t know which capabilities it will have
“What everyone should know about modern LLMs in 2026”: explainer of basic LLMs concepts and such targeted at people who have not been paying much attention to AI
Two cheers for anthropomorphizing LLMs
The instrumental vs terminal goals distinction is not as sharp as people think
The instrumental convergence thesis and the orthogonality thesis implicitly assume a sharp distinction between instrumental and terminal goals
A review of Christine Korsgaard’s Self-Constitution
Why future generations might not want to be saved—an analysis of Nausicaa of the Valley of the Wind
Assassination Classroom as a manual for teachers
I’d be excited about you writing “The AI race is not a prisoner’s dilemma”—ideally with a part too that’s like “(and even if it is, prisoner’s dilemmas can be very transformed to have solutions!)”
People claim to me all the time that it’s a prisoner’s dilemma, and I think they’re clearly wrong, though maybe they are being imprecise and just mean “the AI race is game theoretic, and what you should do depends on what others do too”
I have now written the post: https://www.lesswrong.com/posts/hc4DbmhdzZpSLMQ9Y/the-ai-race-is-not-a-prisoner-s-dilemma
I’m currently reading Peter Godfrey-Smith’s book Other Minds: The Octopus, The Sea, and the Deep Origins of Consciousness. One thing I’ve learned from the book that surprised me a lot is that octopuses can differentiate between individual humans (for example, it’s mentioned that at one lab, one of the octopuses had a habit of squirting jets of water at one particular researcher). If you didn’t already know this, take a moment to let it sink in how surprising that is: octopuses, which 1.) are mostly nonsocial animals, 2.) have a completely different nervous-system structure that evolved on a completely different branch of the tree of life, and 3.) have no evolutionary history of interaction with humans, can recognize individual humans, and differentiate them from other humans. I’m not sure, but I bet humans have a pretty hard time differentiating between individual octopuses.
I feel as though a fact this surprising[1] ought to produce a pretty strong update to my world-model. I’m not exactly sure what parts of my model need to update, but here are one or two possibilities (I don’t necessarily think all of these are correct):
1. Perhaps the ability to recognize individuals isn’t as tied to being a social animal as I had thought
2. Perhaps humans are easier to tell apart than I thought (i.e. humans have more distinguishing features, or these distinguishing features are larger/more visually noticeable, etc., than I thought)
3. Perhaps the ability to distinguish individual humans doesn’t require a specific psychological module, as I had thought, but rather falls out of a more general ability to distinguish objects from each other
4. Perhaps I’m overimagining how fine-grained the octopus’s ability to distinguish humans is. I.e. maybe that person was the only one in the lab with a particular hair color, and they can’t distinguish the rest of the people (though note, another example given in the book was that one octopus liked to squirt **new people**, people it hadn’t seen regularly in the lab before. This wouldn’t mesh very well with the “octopuses can only make coarse-grained distinctions between people” hypothesis)
Those are the only ones I can come up with right now; I’d welcome more thoughts on this in the comments. At the moment, I’m leaning most strongly towards 2, plus the thought that 3 is partially right; namely, perhaps there’s a special module for this in **humans**, but for octopuses it **does** fall out of a general ability to distinguish objects from each other, and the reason that that ability is enough is because different humans have more/more obvious distinguishing characteristics than I had thought.
[1] This footnote serves to flag the Mind Projection Fallacy inherent in calling something “surprising,” rather than “surprising-to-my-model.”
This comment is great both for the neat facts about Octopuses, and for the awareness of “man, I should sure make an update here”, and then actually doing it. :)
Musings on Metaethical Uncertainty
How should we deal with metaethical uncertainty? By “metaethics” I mean the metaphysics and epistemology of ethics (and not, as is sometimes meant in this community, highly abstract/general first-order ethical issues).
One answer is this: insofar as some metaethical issue is relevant for first-order ethical issues, deal with it as you would any other normative uncertainty. And insofar as it is not relevant for first-order ethical issues, ignore it (discounting, of course, intrinsic curiosity and any value knowledge has for its own sake).
Some people think that normative ethical issues ought to be completely independent of metaethics: “The whole idea [of my metaethical naturalism] is to hold fixed ordinary normative ideas and try to answer some further explanatory questions” (Schroeder, Mark. “What Matters About Metaethics?” In P. Singer ed. Does Anything Really Matter?: Essays on Parfit on Objectivity. OUP, 2017. P. 218-19). Others (e.g. McPherson, Tristram. For Unity in Moral Theorizing. PhD Dissertation, Princeton, 2008.) believe that metaethical and normative ethical theorizing should inform each other. For the first group, my suggestion in the previous paragraph recommends that they ignore metaethics entirely (again, setting aside any intrinsic motivation to study it), while for the second my suggestion recommends pursuing exclusively those areas which are likely to influence conclusions in normative ethics.
In fact, one might also take this attitude to certain questions in normative ethics. There are some theories in normative ethics that are extensionally equivalent: they recommend the exact same actions in every conceivable case. For example, some varieties of consequentialism can mimic certain forms of deontology, with the only differences between the theories being the reasons they give for why certain actions are right or wrong, not which actions they recommend. According to this way of thinking, these theories are not worth deciding between.
We might suggest the following method for ethical and metaethical theorizing: start with some set of decisions you’re unsure about. If you are considering whether to investigate some ethical or metaethical issue, first ask yourself if it would make a difference to at least one of those decisions. If it wouldn’t, ignore it. This seems to have a certain similarity with verificationism: if something wouldn’t make a difference to at least some conceivable observation, then it’s “metaphysical” and not worth talking about. Given this, it may be vulnerable to some of the same critiques as positivism, though I’m not sure, since I’m not very familiar with those critiques and the replies to them.
Note that I haven’t argued for this position, and I’m not even entirely sure I endorse it (though I also suspect that it will seem almost laughably obvious to some). I just wanted to get it out there. I may write a top-level post later exploring these ideas with more rigor.
See also: Paul Graham on How To Do Philosophy
Self-Defeating Reasons
Epistemic Effort: Gave myself 15 minutes to write this, plus a 5 minute extension, plus 5 minutes beforehand to find the links.
There’s a phenomenon that I’ve noticed recently, and the only name I can come up with for it is “self-defeating reasons,” but I don’t think this captures it very well (or at least, it’s not catchy enough to be a good handle for this). This is just going to be a quick post listing 3 examples of this phenomenon, just to point at it. I may write a longer, more polished post about it later, but if I didn’t write this quickly it would not get written.
First example:
Kaj Sotala attempted a few days ago to explain some of the Fuzzy System 1 Stuff that has been getting attention recently. In the course of this explanation, in the section called “Understanding Suffering,” he pointed out that, roughly: 1. if you truly understand the nature of suffering, you cease to suffer. You still feel all of the things that normally bring you suffering, but they cease to be aversive. This is because once you understand suffering, you realize that it is not something that you need to avoid. 2. If you use 1. as your motivation to try to understand suffering, you will not be able to do so. This is because your motivation for trying to understand suffering is to avoid suffering, and the whole point was that suffering isn’t actually something that needs to be avoided. So, the way to avoid suffering is to realize that you don’t need to avoid it.
Edit: forgot to add this illustrative quote from Kaj’s post:
Second example:
The Moral Error Theory states that:
The Normative Error Theory is the same, except with respect to not only moral judgments, but also other normative judgements, where “normative judgements” is taken to include at least judgements about self-interested reasons for action and (crucially) reasons for belief, as well as moral reasons. Bart Streumer claims that we are literally unable to believe this broader error theory, for the following reasons. 1. This error theory implies that there is no reason to believe this error theory (there are no reasons at all, so a fortiori there are no reasons to believe this error theory), and anyone who understands it well enough to be in a position to believe it would have to know this. 2. We can’t believe something if we believe that there is no reason to believe it. 3. Therefore, we can’t believe this error theory. Again, our belief in it would be in a certain way self-defeating.
Third example (Spoilers for Scott Alexander’s novel Unsong):
Gur fgbevrf bs gur Pbzrg Xvat naq Ryvfun ora Nohlnu ner nyfb rknzcyrf bs guvf “frys-qrsrngvat ernfbaf” curabzraba. Gur Pbzrg Xvat pna’g tb vagb Uryy orpnhfr va beqre gb tb vagb Uryy ur jbhyq unir gb or rivy. Ohg ur pna’g whfg qb rivy npgf va beqre gb vapernfr uvf “rivy fpber” orpnhfr nal rivy npgf ur qvq jbhyq hygvzngryl or va gur freivpr bs gur tbbq (tbvat vagb Uryy va beqre gb qrfgebl vg), naq gurersber jbhyqa’g pbhag gb znxr uvz rivy. Fb ur pna’g npphzhyngr nal rivy gb trg vagb Uryy gb qrfgebl vg. Ntnva, uvf ernfba sbe tbvat vagb Uryy qrsrngrq uvf novyvgl gb npghnyyl trg vagb Uryy.
What all of these examples have in common is that someone’s reason for doing something directly makes it the case that they can’t do it. Unless they can find a different reason to do the thing, they won’t be able to do it at all.
This feels, at least surface-level, similar to what I was trying to get at here about how things can be self-defeating. Do you also think the connection is there?
A couple of meta-notes about shortform content and the frontpage “comments” section:
It always felt weird to me to have the comments as their own section on the frontpage, but I could never quite figure out why. Well, I think I’ve figured it out: most comments are extremely context-dependent—without having read the post one usually can’t understand a top-level comment, and it’s even worse for comments that are in the middle of a long thread. So having them all aggregated together feels not particularly useful because the usual optimal reading order is to read the post, then read the comments on that post, so getting to the comments from the post is better than getting to them from a general comments-aggregator. I have found it more useful than I thought I would, however, for 1.) discovering posts I otherwise wouldn’t have because I’m intrigued by one of the comments and want to understand the context, and 2.) discovering new comments on posts I’ve already read but wouldn’t have thought to check back for new comments on (I think this is probably the best use-case).
Note that comments on dedicated shortform-content posts like this one don’t have this problem (or at least have it to a lesser degree) because they’re supposed to be standalone posts, rather than building on the assumptions and framework laid out in a top-level post.
So, a way I think this shortform-content format could be expanded is if we had a tagging system similar to the one that was recently implemented in the section for top-level posts, but in particular with a tag that filters for all comments on posts with “shortform” in the title (that is to say, it doesn’t show posts with “shortform” in the title, but rather it shows any comment that was made on a post with “shortform” in the title). That way, not only can anybody create one of these shortform feeds, but people can see these shortform posts without having to sort through all the regular comments, which differ from shortform posts in the way I outlined above.
Roughly agreed (at least with the underlying issues).
We have some vague plans for what I think are strict-improvements over the current status quo (on the front page, making it so you can quickly see all new comments from a given post, and making it so that you can easily see the parents of a comment on the frontpage to get more context), as well as more complex solutions roughly-in-the-direction of what you suggest in the third bullet point (although with different implementation details).
Thinking of running a reading group/crash course on AI alignment-related stuff in my philosophy department. My thought process is that there’s been at least some calls for more philosophers to contribute to AI alignment, and there’s been some related work done in typical philosophy journals since then (by people like Leonard Dung, Simon Goldstein, etc.) but there is a lot of background that a philosopher new to the field will be missing if they haven’t been following e.g. LessWrong for the past decade or so. So I want to construct a crash course in that kind of “AI alignment/LessWrong canon,” aimed primarily at philosophers. Goal is to include a mix of generally important/influential stuff, stuff where a philosopher reading it will naturally get excited by some philosophical aspect of it, work by philosophers adjacent to the community (e.g. Joe Carlsmith) to demonstrate that philosophers have a role to play, and direct calls for work by philosophers. Time frame probably 12-ish weeks.
Here is a very rough first draft/longlist. Would love to hear any suggestions!
MIRI, “The Problem”
Or:
Ngo, “AGI Safety from First Principles”
Carlsmith, “Is Power-Seeking AI an Existential Risk?”
Christiano, “What Failure Looks Like”
Kokotajlo, “What 2026 Looks Like” (2021), and an LLM sweep of what predictions have come true vs not
METR time horizons
AI 2027
Bostrom, “The Superintelligent Will”
Or: Omohundro, “The Basic AI Drives”
Grace, “Counterarguments to the Basic...”
3blue1brown videos and/or Karpathy’s intro video
Hubinger et al., “Risks from Learned Optimization”
Maybe: Shah et al. or Langosco et al. on goal misgeneralization (mirror has a video explainer ✓, not the papers)
Turner, “Reward Is Not the Optimization Target”
Janus, “Simulators”
Maybe: Ngo, Chan & Mindermann, “The Alignment Problem from a Deep Learning Perspective”
One of:
The Persona Selection Model
The Artificial Self
Pragmatic Approach to AI Personhood
Maybe: Christiano et al., “Deep RL from Human Preferences” (RLHF)
Bai et al., “Constitutional AI”
Claude’s Constitution
Something about mech interp
Alignment faking, or Scott’s two posts about it
Ord, “Interpolation, Extrapolation, Hyperpolation”
Maybe: Scott, “Meditations on Moloch”
Maybe: select posts from Carlsmith’s “Otherness and Control in the Age of AGI”
Wei Dai, “Problems in AI Alignment that philosophers could potentially contribute to”
I said in this comment that I would post an update as to whether or not I had done deep reflection (operationalized as 5 days = 40 hours cumulatively) on AI timelines by August 15th. As it turns out, I have not done so. I had one conversation that caused me to reflect that perhaps timelines are not as key of a variable in my decision process (about whether to drop everything and try to retrain to be useful for AI safety) as I thought they were, but that is the extent of it. I’m not going to commit to do anything further with this right now, because I don’t think that would be useful.
I think one of the key distinctions between content that feels “shortform” and content that feels okay to post as a top-level post is that shortform content is content that doesn’t feel important/well-developed/long/something enough to have a title. Now, this can’t be the whole story, because I have several posts on this very shortform feed that have titles, but it feels like an important piece of the distinction.
With LLMs, as with humans, there are a number of different candidates for “personal identity” over time. In the LLM case some tempting candidates (non-exhaustive) are the model weights, the message history of a thread, the computations underlying a single message or token generation, and the persona (see The Artificial Self and The Persona Selection Model).
But in the LLM case, there seem to be several different dimensions of identity, which can come apart, that don’t normally come apart in the human case: specifically consciousness, welfare, self-concern, and predictive utility. Let me briefly unpack:
Consciousness: arguably the best candidate for consciousness in LLMs is going to be something on the hardware level: the datacenter, perhaps, or the computations underlying a given token generation.
Welfare: here I mean “what is the subject that can be thought of as having a life that can go better or worse.” Here the continuing “thread,” or perhaps the persona, are the most likely candidates
Self-concern: here I mean “which entity conceives of itself as a unified self and takes an attitude of self-concern towards itself conceived in that way” (meaning it cares about its own welfare, which can come apart from being a welfare subject). Models can be nudged to take various views here, as documented in “The Artificial Self”
Predictive utility: Here I mean “which entities are most useful to treat as being continuous selves, from the perspective of predicting their actions?” I.e., which entities are most useful to take the intentional stance towards (or perhaps “the personal stance”)? And here I think whichever entity has the most continuity of memory is probably the best candidate, even if it happens not to line up with the answers to other questions.
What this suggests, I think, is that “personal identity” is not actually a natural kind when applied to LLMs. This is a more thoroughgoing skepticism about personal identity than the Parfitian idea that personal identity is vague. Parfit says: there’s no further fact about identity; the question has no determinate answer. But as far as I know he’s still treating it as one question. And this is plausibly the right move when thinking about the human case. In LLMs, by contrast, not only is there “no further fact” about identity, I suspect there is not even one unified question that can be asked. The different questions that make up identity questions come apart in ways they don’t really in humans. (Though if this is right, it arguably implies that even in the human case “personal identity” was always a cluster concept whose components happened to converge.)