English teacher for adults and teacher trainer, a lover of many things (math, truth, literature(s), art, physics, history, languages) and people.
manueldelrio
Fair enough. I’ve given a quick browse to the Yud post, but it is somewhat dense and requires more time than I have now (I’ll take a deeper look in the evening). I think it is fair to say that ultimately, ethical views rely on intuitions. Problem with this is that I would dispute we share them ‘even after subjecting them to reflective action’. I’ve been thinking much about ethics in the last, say, 6 or so years. I don’t think I had any intuition about moral realism to start with. I do remember liking then Kantian deontology, but the thought-process that led there was something like this:
1) I am secular, and can only accept rational, secular justifications for ethics
2) If ethical realism isn’t true, it doesn’t make sense at all to speak about ethics. It is just subjective preferences and/or useful evolutionary tricks for individual/group survival and multiplication.
Back then, I just disliked the consequences of 2, but the more I kept reading about ethics, as well as other stuff which I’ll bunch together as ‘evolutionary biology and psychology’, the more I felt that wanting ethical realism to be true was no different from wanting religion to be true. They were exactly the same sort of reference class, they had roughly the same sort of explanation for why they appear, for why they are individually appealing to us and for why they are useful tools for group coordination. I still feel like 0 intuition for ethical realism: I can have strong feelings about some stuff, but I can easily explain them away as a mere result of historically fit natural-and-social evolutionarily useful strategies that get coded in our genes and which we are indoctrinated and ‘aligned’ to interiorize since children. I grant the same truth-status to religion and ethical realism, which is none.
So these kinds of debates do feel rather circular if they just end up with ‘okay, I have these intuitions and you have yours. Have you reflectively examined yours? Then you should agree with me’. I do get one thing Yud says in the article, i.e., that it is probably more pragmatically useful just to focus on the object-level (“what do I do in this regard?”), but I feel it is going to just make me very unreceptive to any sort of ethical argumentation to change any of my current preferences or my anti-realist, contractualist view of stuff.
Yes, I employed the ice-cream for two different metaphors. In the first case, the point I was making that , if you don’t believe in moral realism, any claims to ‘shoulds’ have no truth-value: they are either purely meaningless and have no referent (which is what I think an Error Theorist would say) or they are acting as mere preferences that you acquired for whatever subjective set of reasons and life experiences, and none more authoritative than another. In the second example, I was using ice-cream as providing a good example of why we can’t trust intuitions to be truth-bearing, but rather to being (for a give time and place) evolutionarily adaptive, which is kind of what I believe is the reason why people and societies develop strong moral intuitions and feelings.
I don’t think preferences, in the way I am taking about them, can be anything other than tier 3: at most, there is a psychological truth about them (they condense what you as an individual and messy, evolutionary rube-goldberg machine with a socio-historical context and impossible to recreate individual nuances of experience feel you value). Preferences don’t tell you anything about the external, physical world; as for the socially constructed social world and agreed-upon beliefs, preferences are generally causally dependent on these, but not completely so.
So yeah, we don’t agree on the plausibility of moral realism. I think I understand what the consequences of believing in it would be, but that doesn’t affect in any way how credible I find the concept to begin with. That leads me back to what start the chat from the notes I made: “how incredibly baffled I get when trying to understand moral realists”. My lack of understanding is not that their arguments are inconsistent given their premises, but rather that the premise seems to me plainly absurd, wishful thinking and an attempt, conscious or not, to rescue morality in a world without religion. And when I ask or look for clarifications of why a realist holds the belief, I just get extremely insatisfactory answers, almost all of which are variants of ’well, these are self-evident intuitions that I have”.
I don’t think I’ve understood your ‘I’m saying the map changes but the territory might not’ . Would you care to clarify? If I understand the analogy, you’re saying that in tier 1, there’s a territory (the natural world) that remains stable whatever maps (= human beliefs and interpretations of it) are made of it, whereas in tier 2, the map (also human beliefs and interpretations) changes but the territory (=moral truths?) might not?. I mean, my core claim here is that moral truths are a fictional territory which doesn’t really exist, except as a figment of people’s imaginations and as a useful social technology. So I would dispute that “When many thought slavery ok, that world was in some real way worse for it”. This just feels to me like some flavor of hindsight bias: given who you are, when you are alive and the set of values/beliefs of the society you belong to, you interiorize our moral values as something true and retroactively apply them as a benchmark, while assuming some sort of narrative of moral progress where our current configuration is at the very least a local maximum in some function that is growing to infinity. While I can share a distaste for slavery, I think describing the world with it as ‘worse’ simpliciter is meaningless. At most, you can say ‘Given my and my societies’s beliefs, which include a degree of moral delusions, by our standard a society with slavery can be described as ‘bad’ and ’worse″.
As for the ice-cream example, I fail to see why realists says this type of statement is counterintuitive. What I imagine is that it hints at the fact that we interiorize and naturalize moral beliefs as roughly equivalent to scientific ones. I may feel ethics is an Error Theory without a referent, but like most people, I recoil at violence being exerted against the helpless. But this recoiling and the emotional effects are something I can easily not take serious as truth-describing, in just the same way as I can easily not take as ‘truth-describing’ that fatty, sugary foods are good for me, even if my brain rewards me for eating them. I’d say my base rate for scenarios like the one you depict is something like this:
a) Me, an animal with desires and needs, dislikes intensely being harmed or having my preferences (including not being hurt) violated.
b) Me also, an narturally and socially evolved animal that lives in a society with others, has interiorized and being indoctrinated about a set of rules about how to behave and what to expect with respect to other members of the community. These include notions of ‘Fairness’ which get abstracted and extrapolated from their real origin (in a relatively egalitarian hunter-gatherer band in which no human is more powerful than the others combined, a way of ‘trimming’ the too violent and too pushy by ganging against him and killing him at night / from a distance / with numbers) to be assumed to be true and applicable to other groups of humans and/or to be expected from the Universe at large.
Maybe that is our crux? I am inclined to Error Theory anti-realism. I think social norms are indeed just consensus equilibria, which doesn’t mean that it is just a vote thing. My claim is there is no ‘real’ about social norms (and ethics). If you abide by some axioms x, y, z on what your goals and/or what the goals of a community should be you can say ‘doing a, b, c makes the satisfaction of those goals more probable’. But you can’t say anything about the norms per se as to their ‘correctness’ or ‘incorrectness’ with respect to some base truth independent from the axioms.
I would be very suspicious of taking consensus as evidence for anything ‘real’ . That would have a first issue: at what moment in time and space is the consensus ‘real’ as opposed to other moments when the consensus is different? Is the historically rather strong lack of overall moral homogenousness evidence against its realness? At most, it just shows a community has arrived at some sort of agreement and temporary equilibrium on something, and it is likely to infer that consensus that last long probably track something which doesn’t have to be ‘truth’, but rather ‘usefulness for the survival and expansion of the community with said consensus’.
I think the art example actually works against your argument. What makes the Mona Lisa highly appreciated, much more than very similar paintings in the same reference class by the same author, is a series of historical accidents (like the fame that it accrued after having been famously stolen in 1911, with the consequent media fuss). The value of works of art is that of a consensus, even if it is not strictly that of the wide public, but of the bubble of people involved in ‘the art world’ (critics, artists, curators, patrons, etc...). Of course, there’s a spectrum between a stick figure and the Mona Lisa. Some would say that spectrum is not big when comparing a Pollock (which is basically just randomly dripping paint over a canvas, something that a gorilla or a young kid could do). But even those differences (ot technique, style, originality...) are ‘axioms’ the ‘art world’ changes from time to time.
I just feel that for stuff that’s not in tier one, you can’t really speak about ‘shoulds’. Not even there: there are no should with tier one either: there’s just truths about an external world that are independent of human desires and cognition. You can ignore or reject them, but reality will reassert itself in spite of your wishes, or those of the whole human community. You can ignore reality, at your peril. For tier 2, individuals seldom can go unscathed by ignoring or rejecting the societal shoulds, but they change with times and with movements in favor of changes in some directions. My claim is, with respect to shoulds, everything is of the same reference class as ‘I prefer vanilla ice-cream’ or ‘I prefer blue’.
That’s a good and deep question, for which I don’t have an answer. My intuition about the answer would be no, though, i.e., we have different values/preferences with a complicated history and a degree of inconsistency forming some n-dimensional vector space that cannot be reduced to a simple scalar. But even granting that they could be reduced to a vague ‘desire for the good’ terminal value, it would seem you can develop this in wildly different and opposite ways. The fact that minds desire a common x would, if proven (like, sampling aliens for example) update me towards the existence of some shared, emergent subjective belief that arises in sentient and intelligent creatures, but I don’t see how it would update me towards believing this belief is ‘true’ in the naturalist, correspondence-to-an-external-reality kind of sense. I think I’d agree with your second interpretation, ‘that makes goodness real for minds, but not objectively’. Objectivity is a term whose use I always find extremely confusing. It makes sense to talk about natural phenomena with the lens of “based on real facts, evidence, and logic without letting personal feelings, likes, or unfair biases change your view”, because we’re assuming they are mind-independent and follow laws that are orthogonal to our subjective feelings, likes and biases. When you talk about what minds believe, it becomes messier (because a lot of facts are socially constructed, like money or laws, and not independent of feelings, likes and biases; in fact, they are intrinsic to what it is like to be a human, part of a human community with its intellectual and material limitations. What does ‘real’ mean when we say ‘this piece of paper is worth 100 dollars’ except something like “well, this is some kind of useful and enforced belief/practice which is contingent on the existence and interactions of a particular type of being with properties x, y, z”). There’s like three tiers here: stuff that’s mind-independent (it doesn’t matter what any mind or group of minds believe and think with respect to the underlying truth), stuff that’s intersubjective (which is mind-dependent but not amenable to change by an individual) and purely subjective stuff (which is completely dependent on the subject’s subjective beliefs, feelings, whatever). Morality is for me something that belongs to somewhere between the second and third tier (the ‘agreed social norms’ is 2nd tier, the personal belief that x is wrong independently of what anyone thinks, to 3rd tier).
So on moral realism: I just feel there’s nothing you can really use for updating on the view that there are “stance-independent moral truths whose truth does not depend on what any mind happens to value or desire.” (Theology is more updatable here; like, you just need to have God(s) appear, interact with you and do supernatural stuff repeatedly and measurably). If you view of moral realism is something like “Given creatures with volition, reason, different interests, etc..., who would want to engage in positive-sum social games, there might exist an optimal configuration of the latter that would be reachable by different creatures without any contact with each other”, then while I am still somewhat skeptic, this is something I can imagine to be true, and I can think of ways in which it can be updated.
I wouldn’t dispute there’s some difference, i.e., given a set of normative values, an inner/outer observer would judge (and feel?) the situations differently depending on their values. I just don’t think it is doing any truth-tracking about reality. Your second example (‘the misery-preferer is normatively confused’) reminds me of a strategy I usually find in moral realists which I do not find convincing: it tries to establish a parallelism between logical rationality and normative rationality (there are ends or actions you have reason to choose regardless of your desires). I just disagree with the latter: there are no ends you have ‘rational reason’ to choose except those you desire as part of your preferences, and this is just some kind of quasi-Kantian attempt to smuggle normative ethics into reason. Affirming someone is ‘normatively confused’ seems to imply there is some correct, ‘objective’ set of norms a rational being has to follow lest he/she fall into error. Reason can tell you that your beliefs are inconsistent, or that your chosen means won’t achieve your ends, but I don’t see how reason by itself can supply the terminal ends.
I think I might not have explained myself clearly in the previous comment. What I meant by instantiate was something like ‘validated by experience’, not ‘manifested in experience’. In the 2+2 case, your intuition gets externalized in a way that provides for independent checking (you can count again, somebody else who wasn’t there can count them and tell you the answer, etc...). The world constrains the answer. In the regret example, the regret testifies to a psychological state, which can be shared by others (or not?; definitely less shared than 2+2) but how you or others feel about an event is not constrained by the world. There is no detectable difference in a world were the only thing that changes in that scenario from feeling to not feeling regret. There is a world of difference (pun intended) in a world where adding two objects generates four to a world where adding two objects generates five.
I am generally very skeptical of purported knowledge that is neither gained through the senses nor independently constrained by empirical reality (but then, see my mathematical platonism). But in the case of logic (and math), I feel they are vindicated by contact with reality in a way in which moral beliefs aren’t. Math is indispensible for all our scientific explanations—without it, there would hardly be anything left in science; Logic can’t be proved by experiment, but Reality appears not to instantiate logical contradictions (so it is also indispensible). By contrast, objective moral facts seem explanatorily dispensible. I mean, you need them for social interaction, positive-sum games, but a world A where some moral beliefs are intrinsically, objectively true and a world B where they don’t exist work in exactly the same perceptible, visible, detectable way.
Hey! Thanks for the comment! I really appreciate it!
I mean, for 1, I do think that most things I believe are, reductionistically speaking, ‘just my opinion’, in the sense that they lack a strong, unassailable foundation. I just feel its very easy for to to compartmentalize ‘I have strong (and emotional) opinions about x’ from ‘I have a strong conviction and evidence that x is true’. For some of those beliefs I could perhaps fall back on what could count as acceptable evidence and/or arguments in their support that third parties would recognize. If we center on ethics, yes, there’s lots of stuff I feel strongly about: in most respects, I am your average WEIRD Westerner, and have been aligned since youth to find some stuff horrible, but I just fail to see the epistemic credentials of my moral intuitions. This connects with your second point: while ultimately, you can argue that it’s turtles all the way down and that logic, math, science, induction… ultimately rest on intuitions, I feel you are skipping a significant difference here. All of the latter are intuitions that can be tested with respect to some External World (it doesn’t matter if you don’t want to give metaphysical foundation to an external, nonhuman reality. It suffices to accept that we seem to interact with some outer world that has seems to have laws of its own that can be discovered and tested, have decent predictive power and that seem to have bad effects when you try to ignore them). You might not have certainty, but empirical observation, causality, logical consistency and mathematical modelling can be tested; while the moral intuitions cannot.
Map and Territory, Predictably Wrong (2)
Map and Territory, Predictably Wrong
Starting The Sequences: Some brief notes from the preface and the introduction
Just wanted to say I found this conversation very interesting - so much so I will be taking a look at the previously linked ones. I think I also found it useful. Although I’ve read what I think is a significant amount of Rationality texts, I still feel that for some very basic stuff that is touched upon here, I am still as flabbergasted as much as any layman, and that the talk helped clarify some things.
Santiago de Compostela, Spain—ACX Spring Schelling 2026
I liked your post, and I’m probably the sort of person predisposed to like it (for context: history has been my favorite subject since I was 11; it was my first degree at university; I’ve read widely both in historical scholarship and in the meta-justifications for studying it). I’ve also been deeply steeped in the broader humanities, and that background has shaped how I see the world and what my preferences and the things I care about are. Still, statements like “the humanities exist to improve our minds” and “history improves us” strike me as mostly normative aspirations rather than accurate descriptions of how interaction with those fields typically functions.
The core difficulty with the “history as context” argument is not merely that context is selected, but that the criteria of selection are instrumentally shaped. In actual educational practice, the organizing framework is almost always the nation-state; history becomes a kind of secular civic religion that presents the polity as an inevitable, self-evident, and temporally continuous subject. That isn’t arbitrary so much as ideological: the selections are optimized for legitimacy, cohesion, and identity formation, not for epistemic clarity. The same issue appears in “history as memory,” where decisions about what is worth remembering—and how to interpret what is remembered—follow political, cultural, and institutional imperatives. Textbook committees and curricula do not behave like disinterested archivists; they behave like legitimacy-producing institutions.
I also have to admit that I don’t have strong evidence that history has “improved” me in any of the virtues usually invoked. I’ve enjoyed reading it, and learning about the actions, thoughts, constraints, and delusions of past humans. But I don’t think it has given me superior judgment, foresight, epistemic humility, or civic virtue relative to STEM peers. People often claim that the humanities make one more open, curious, or empathetic toward other societies; that they illuminate institutional dynamics or the grammar of civilizations. Maybe that can happen, but I’m not confident it did in my case. It is extremely easy to adopt a single interpretive lens—in my case for many years it was Marxism—and read everything through that filter, which produces narrative satisfaction without necessarily producing accurate models of reality.
This isn’t an argument against history or the humanities. It’s an argument against a certain idealized story we tell about them. If history improves people, the mechanism seems non-trivial: it requires meta-reflective skills, comparative reasoning, and some capacity for model-building. Those are not reliably taught within the discipline, and are often actively undermined by curriculum structures designed for identity formation rather than truth-seeking.
Hi, Anna. The post was very interesting, and I’d be happy to do some ‘kibitzing’ (had to check the word’s meaning, actually), even if I give low credence that anything I’ll say might be of use.
I am a relatively late arriver to the world of Rationalism (from about 3 years ago or so), having gone through some of the usual suspects (Julia Galef’s Scout Mindset and her podcasts; Chivers’ book on Rationalism; HPMOR and some online posts from Scott Alexander; some CFAR materials online, and some youtube videos; not much Less Wrong or The Sequences though, until now). What attracted me to it was my quasi-religious attachment for truth-seeking and for avoiding self-deception, and my desire to become more rational in my thinking and decision-making. I am still reading and exploring. You might remember that I’ve been one of the assistants to the CFAR test sessions, and would be interested in the future in going a workshop if it doesn’t entail having to go across the Atlantic (I live in rural, NW Spain).
Through my eyes, CFAR seems like a group of people who are engaged in the sort of ‘right thinking’ I aspire to (LW also; perhaps the distinction from the outside is that CFAR seems to be a more educational-oriented project in practice).
A lot of what I see in the test sessions I find confusing, but I don’t think it’s from any fault of yours. They aren’t workshops, and I feel a lot of uncanny valley-ness in that I recognize most of the ideas and terms but haven’t really interiorized their meaning. Also, the only correlate your practice brings to mind is something like talk-therapy which, while pretty common in the States, is really unusual over here (the only, and very few people I know who have engaged in something like it do it with psychologists and for medical reasons). This isn’t a criticism so much as a cultural translation problem: from where I sit, it’s not always clear how to distinguish applied rationality coaching from therapeutic modes of engagement, and that makes it harder to know what norms to bring.
I think it would help me a lot to actually go in detail over some of the material you have online (particularly, the CFAR Handbook). I think I just mostly lack the grammar of how these things are supposed to be done.
‘Hope for something out loud’: I’d hope for a chance in the not-too-remote future to take once of your workshops this side of the ocean. ‘Try to speak to why you care rather than rounding to the nearest conceptual category’: as I said, I care really personally about truth-finding, truth-seeking and being part of an expanding circle of people who share this frame of mind, which makes me see your efforts, whatever they end up producing, as merit-worthy. And that’s why I care enough to watch closely how this iteration of CFAR actually plays out.
Question: wouldn’t how we dealt with Acid Rain and the Ozone layer be counterexamples? In these cases one didn’t have a clear deadline, but we did manage to muster the resources and effort to overcome the issues. I would think the issue is not so much just status quo, but actual, generalized understanding of the magnitude of the risk + degree of certainty + actionability. AI risks seem to have big problems with each of those three.
That sounds very reasonable. In the review, I wasn’t consciously trying to play a blame game with Yudkowsky and Soares (I generally think blame is ineffective at producing good outcomes anyway) but rather to articulate a reader’s uncertainty about what their reference class for relevant expertise actually is.
My own naïve take would be something like what you say: people with substantial hands-on technical experience in contemporary AI systems, combined with people who have thought deeply about the theoretical aspects of alignment. My impression is that even within this relatively restrictive class there remains a wide diversity of views, and that these do not, on average, converge on the positions defended in the book.
I only have a superficial understanding of Yudkowsky’s work over the years, but I am aware that he led MIRI for roughly two decades, and that it was a relatively well-funded, full-time research organization explicitly created to work on what was seen as “the real alignment problem” outside of frontier labs. From an outsider’s perspective, however, it is not obvious that MIRI functioned as a place where deep, hands-on technical understanding of AI systems was systematically acquired, even at a smaller or safer scale.
If avoiding frontier labs is justified on the grounds that they accelerate catastrophic risk, then MIRI would seem to have been the natural alternative pathway for developing compensating expertise, yet it is not clear (at least to a non-insider) what concrete forms of technical or empirical understanding were accumulated there over time, or how this translated into transferable expertise about real AI systems as they actually evolved. In fact, from a superficial impression, it is difficult not to come away with the (possibly mistaken) impression that much of the work remained at the level of highly abstract theorizing rather than engagement with concrete systems..
That gap makes it harder, for me to see the absence of conventional credentials as epistemically “screened off” rather than simply displaced.
I mean… if there’s one thing I learned by studying literature and literary criticism at uni (spoiler: I didn’t learn much, and most of what I did was not very valuable), it is that texts are very seldom completely self-consistent. Still, I think what you say is fair: the authors haven’t just given up (if they had, they wouldn’t have written the book in the first place), but it feels to me that the solutions they propose are wildly impractical, and perceived as such by the authors, and that this perception likely plays a big role in their P(doom). If the bar for “meaningful risk reduction” is set so high that only globally coordinated, near-wartime restrictions count, then the conclusion of extreme doom follows almost automatically from the premises. I’m not convinced the argument sufficiently explores whether there are intermediate, messy, politically imperfect interventions that could still substantially lower risk without meeting that idealized threshold.
Yes, I definitely agree that wanting something is not the same as having an intuition that it is true. But we humans are feeble creatures, and have to be on guard against ourselves, and of believing what we want to believe for no other reason than that it’s comforting. If ethical realism were true, it would grant a measure of certainty that I appreciate.
And yes, reason can be explained away an evolutionary useful, gene encoded, quirk, but it does seem to reflect the world by allowing us to make successful predictions about it. I fail to see how ethical intuitions do this. Unless it be the case that what you mean is something like ‘they allow us to coordinate with other humans better’, which is true, but doesn’t really further any truthfulness for those beliefs as such (like with religion). Humans, yes, are creatures with strong feelings about some stuff. More controversial, but possible if you restrict the range, is that all humans tend to have strong feels about some stuff (in some ratio like Haidt’s 90% chimp, 10% bee). But I still feel you can’t call this ‘realism’. Perhaps i am just too square-minded and restrictive in what I see as ‘real’, but I don’t get the impression, from reading realists, that they also stick to this brand of realism (which would further little or no arguments against reducing ethics to socially accepted and variable norms in both time and space with no particularly privileged ground).