Thoughts I failed to develop into posts
Agent Foundations
Path connectedness in Orthogonality
Part of the misunderstandings in the discussion about Orthogonality would be cleared with the concept of Path connectedness.
What defenders of Strong Orthogonality have in mind is that every point in the space defined by the axes “Intelligence” and “Values” could be filled by a possible entity (as a math object, or as a possibly existent according to the laws of Physics object).
What some opponents of Orthogonality have in mind is the fact that you probably can’t modify the neurons in a human brain so that it affects only the human’s values and not its capabilities.
These are actually two different questions, and while the answer to the first one is probably a clear Yes, the second one is trickier, and can be reformulated so: is there always a continuous path of transformation between two agents of equal intelligence that turns the values of the first into the values of the second?
A priori vs Normative
Some people in AF, and especially people criticizing AF, seem to confuse the terms a priori and normative. A priori means definable without reference to our universe, like a triangle. Normative means saying how things should be even if they are, like the concept of justice, which can’t be refuted just because people aren’t just (it might be a fake concept, though). People call AF concepts normative, and they are to some extent, in the almost tautological sense that every reasoner should reasoner correctly if they want to reason correctly. But it isn’t clear that there is an imperative to reason correctly unconditionally. And what AF describes is relevant because a perfect reasoner would be most capable, not because it’s subsumed under a normative concept.
Decision Theory
Finding an empty box in open Newcomb where Omega has perfect prediction is contradictory, but the situation is the limit of situations that are not contradictory, so it seems advisable that your action should be the limit of your decisions in those other situations.
Nitpicking on Embeddedness
Embeddedness is a property of a universe, which is not just “being modelable as math”, but also not an empirical property. It’s reminiscent of Kant’s transcendental.
Each of the four properties of Demsky’s Embedded Agency are incompatible with Non-Dualism, so having a non-dualist universe implies all of them. But it’s not true that each of them implies the other (scenarios are thinkable which have one but not the others).
Similarly, Kosoy says:
They require the agent to be larger than the environment.
But this property doesn’t follow from the others.
and don’t tend to model self-referential reasoning, because the agent is made of different stuff than what the agent reasons about.
“Tending” is not appropriate vocabulary for conceptual implications.
Embeddedness as a self-help concept
If you plan the things you are going to do, and it doesn’t feel like you’re sacrificing something, your plans are fake.
AI
Platonic Representation Hypothesis
Naive Platonism claims only things like circles instantiate a mathematical “form”. General Platonism claims that not just easily describable objects, but also hair, mud or dirt also instantiate a form [Parmenides 130d]. Today’s super-Platonism claims that not only some things, but every single thing, is just a vector in a 40,000-dimensional vector space.
ASI persuasion
Persuasion is very bottlenecked on personalized interaction time. The impact of friends and partners on people’s views is likely much larger. [...] This implies that even if we don’t get superhuman persuasion, AIs influencing opinions could have a very large effect, if people spend a lot of time interacting with AIs.
“The best diplomat in history” wouldn’t just be capable of spinning particularly compelling prose; it would be everywhere all the time, spending years in patient, sensitive, non-transactional relationship-building with everyone at once. It would bump into you in whatever online subcommunity you hang out in. It would get to know people in your circle. It would be the YouTube creator who happens to cater to your exact tastes. And then it would leverage all of that.
With AI, it’s plausible that coordinated persuasion of many people can be a thing, as well as it being difficult in practice for most people to avoid exposure. So if AI can achieve individual persuasion that’s a bit more reliable and has a bit stronger effect than that of the most effective human practitioners who are the ideal fit for persuading the specific target, it can then apply it to many people individually, in a way that’s hard to avoid in practice, which might simultaneously get the multiplier of coordinated persuasion by affecting a significant fraction of all humans in the communities/subcultures it targets.
A: Intelligence is something like “ability to achieve one’s goals.” But this doesn’t take into account the fact that some goals may have hard upper limits where others don’t.
In particular, this seems to apply to things having to do with social behavior. I can imagine beings that are qualitatively better than humans at math, information recall, etc., since there are already orders of magnitude of variation in these abilities among humans. However, social abilities like “ability to manipulate others” do not seem unbounded in these ways. There are some people who are good at manipulation, but it doesn’t seem like this is a “skill” like math ability that spans orders of magnitude. Roughly speaking, are no “super-manipulators” out there who can manipulate ordinarily wary people.
For instance, one of the most effective ways to get people to do your bidding is to start a cult: there are plenty of chilling stories about the level of devotion that cultists have had to their various leaders. However, it’s not at all clear that it is possible to be any better at cult-creation than the best historical cult leaders — to create, for instance, a sort of “super-cult” that would be attractive even to people who are normally very disinclined to join cults. (Insert your preferred Less Wrong joke here.) I could imagine an AI becoming L. Ron Hubbard, but I’m skeptical that an AI could become a super-Hubbard who would convince us all to become its devotees, even if it wanted to. If social abilities like this are subject to hard upper bounds that have already been nearly achieved, then there’s no potential for AIs to achieve their goals better by becoming superhuman at these abilities, which makes it problematic to just postulate an AI that’s “superhuman at achieving its goals.”
B: L. Ron Hubbard might be the upper limit for how successful a cult leader can get before we stop calling them a cult leader.The level above L. Ron Hubbard is Hitler. It’s difficult to overestimate how sudden and surprising Hitler’s rise was. Here was a working-class guy, not especially rich or smart or attractive, rejected from art school, and he went from nothing to dictator of one of the greatest countries in the world in about ten years. If you look into the stories, they’re really creepy. When Hitler joined, the party that would later become the Nazis had a grand total of fifty-five members. There are records of conversations from Nazi leaders when Hitler joined the party, saying things like “Oh my God, we need to promote this new guy, everybody he talks to starts agreeing with whatever he says, it’s the creepiest thing.” There are stories of people who hated Hitler going to a speech or two just to see what all the fuss was about and ending up pledging their lives to the Nazi cause. Even while he was killing millions and trapping the country in a difficult two-front war, he had what historians estimate as a 90% approval rating among his own people and rampant speculation that he was the Messiah. Yeah, sure, there was lots of preexisting racism and discontent he took advantage of, but there’s been lots of racism and discontent everywhere forever, and there’s only been one Hitler. If he’d been a little bit smarter or more willing to listen to generals who were, he would have had a pretty good shot at conquering the world. 100% with social skills.
The level above Hitler is Mohammed. I’m not saying he was evil or manipulative, just that he was a genius’ genius at creating movements. Again, he wasn’t born rich or powerful, and he wasn’t particularly scholarly. He was a random merchant. He didn’t even get the luxury of joining a group of fifty-five people. He started by converting his own family to Islam, then his friends, got kicked out of his city, converted another city and then came back at the head of an army. By the time of his death at age 62, he had conquered Arabia and was its unquestioned, God-chosen leader. By what would have been his eightieth birthday his followers were in control of the entire Middle East and good chunks of Africa. Fifteen hundred years later, one fifth of the world population still thinks of him as the most perfect human being ever to exist and makes a decent stab at trying to conform to his desires and opinions in all things.
The level above Mohammed is the one we should be worried about.
Math
Category Theory is like Comparative Linguistics. Each language has their own categories, and they can only be defined through terms that are only meaning inside each language. But Comparative Linguistics defines other categories which, while being less interesting for the description of each language, enable comparison. The syntactically defined direct object than many languages have is more accurate for description than the semantically defined “object”; but only the second one allows comparisons. Similarly, the terminal object might be irrelevant to describe some categories, but it allows comparisons.
Political philosophy
Non-causal isomorphisms
Hofstadter claims in GEB that an anti-reductionist attitude is justified when we find relevant isomorphisms at a high level that aren’t present at the lower levels. I’m unsure of whether identifying isomorphisms alone is relevant to understand phenomena if we don’t find causal relationships affecting those isomorphisms. Some candidates of potentially not-very-helpful isomorphisms: agency, high modernism (James Scott), patriarchy, meme, Christianity and Buddhism as “religions”, orthodoxy in the Catholic Church and in the Soviet Union
Steelmaning a less naive communism
I’m steelmanning, since I’ve never met anybody using the label Communist who acknowledges that:
-creating wealth is a non-zero sum game.
-the percentage of needed mistake theory is non-zero.
Naive communism (NC) starts with: “Humans are good, so...”, which can be refuted with “Good ideology. Wrong species.”[2]
A less naive communism (ALNC) has to talk about incentives.
NC sees wealth creation as a zero-game. ALNC doesn’t, and has to convince listeners that a certain trade-off between wealth creation and justice is desirable. It’s not that people deserve the wealth created by others’ work (nobody does), but that the system as it exists stops people from creating their own wealth.
ALNC could say: “Regardless of human inclinations, a mechanism is thinkable in which human greed is let free rein to create wealth, but in which the reward of such productive greed is kept at the (empirical) minimum that still moves greedy people to create wealth, and the rest is used to improve everyone’s life”.
As a European, I could never understand why so many American rationalists would disagree with an idea like that. We can disagree about the level at which he benefits would decrease, but it seems like a tautology.
Lately I realized that it isn’t a tautology, and like so many things in life it has to do with bounded rationality. Optimizing and distributing isn’t free, so apart from transfering some wealth (like the communist is willing to do), some wealth simply gets destroyed in the process, and I don’t find it tautotlogical anymore that the amount of destroyed wealth should be non-zero (I still find it plausible, but not tautological). It’s not obvious anymore that doing this is the best way to achieve maximum utility (maybe very unfair, but very wealth-creating markets are simply better at increasing everyone’s level of wealth.
Possibly relevant; haven’t yet read: https://jacobin.com/2026/05/central-planning-soviet-union-socialism
Psychology & Relationships
The gap between therapist and research manager
You tell a therapist about your relationships, and they give you a hypothesis of a systematic failure mode. A great therapist has to find the middle point between saying obvious things, and saying specific things that don’t apply to you. “You have to take it easy” is bad, because it’s obvious. “You might be secretly gay” is specific, might bad if you’ve thought about that and that’s not the point. I once angrily told a therapist all the ways in which my father was a bad person. The therapist replied: “You had a happy childhood; you might be idealizing your father”. The problem was solved forever, to my big surprise.
I wish something like this existed for other contexts beyond relationships. Someone with knowledge about an area, able to give hypothesis on systematic failure modes. “Maybe you are too deferential on area A, but too contrarian on area B”. Someone told me this is what a good research manager is supposed to do. They also told me there are no good research managers.
Chill as free-loading
At a recent circling session, one participant had a definitely-not-in-the-moment attitude which made several participants be not at ease. I expressed my annoyance and tried to express at length what the point of circling was and how he was making it impossible for anyone to enjoy it. A very friendly told me to be less judgmental and encouraged everyone (except me, apparently, by the Doctrine Of The Preferred First Speaker), including the obnoxious participant, to do as they pleased. I realized that having a chill attitude increases the load of people who try to create social norms.
Meditation
I was skeptical of meditation for a long time, and I recently realized that it mostly had to do with me not trusting most of its advocates. I still don’t trust them, but I’m more open about meditation. Here are some things that would have helped me become more open earlier, potentially useful to use with people in a similar situation:
-meditation working is not a point for Buddhism. Buddhism claims many things, most of them are false, and it’s not surprising that of them are kind of true.
-avoiding connecting it to post-rationality: that the brain does weird things doesn’t change the distinction between between map and territory (it just means that the place where the map is kept is itself a weird part of the territory)
-avoid connecting it to staying in the moment in daily life as the goal of mediation. I once heard a meditator disapproving of alarm reminders for meditation, saying that the goal was not to need alarms. I would now endorse a phone alarm as a metaphor of meditation.
-one thought I always found ridiculous was the talk about becoming detached, while people still have sex, go to restaurants, etc. I now see it as a tool to anchor yourself to the attachments that you endorse long-term. One guy told me he used meditation to get better at picking up women. I used to thought things like that were contradictory and ridiculous; I still find it ridiculous, make it makes sense that a tool can be used for many different things
-sell it as a tool to get intellectual clarity, creativity, etc.
-mention the dangers of meditation
Everything is a status game
I feel like evolutionary psychology truths like “everything is a status game” are to “the art of being happy” as the physical truth “cold is not an objective phenomenon; there’s only heat” to someone trying to learn how to use a chimney to get warm: technically true, but irrelevant or misleading in context.
I don’t doubt that every thing is a status game in some sense, just like I don’t deny that everyone is selfish in some sense. But the world is full (doesn’t matter if it’s 20% or 2%; it’s millions) of people who act in ways that are so altruistic / status-orthogonal according to my perception that I don’t care if it is still selfish / status seeking in the ev-psych sense. If the question is how to be happy, the answer is: find people who enjoy doing things that are functionally identical with being altruistic and sharing for the pleasure of sharing, and learn from them or at least enjoy their company.
Rationality & Philosophy
Clinging to inertia
The Sequences describe the failure mode of not accepting the refutation of one belief of yours. I had a different mode, which corresponds to what Yudkowsky calls Old Rationality: being too skeptic, having too few opinion. Despite generally welcoming being corrected on what I thought, this was also a failure mode, because I was leaving money on the table. I was clinging to inertia, whereas being rational requires not avoiding being wrong, but having a right/wrong ratio that makes you win less than you could.
Words
I feel that the doctrine of words as clusters in Thingspace is true for things in the physical space, but doesn’t apply e.g. to mathematical objects. The difference is that in physical space there are no essential lines to be drawn, and since every point has many more dimensions than human language can describe, every attempt to point at a point is in fact pointing at a region. But (some) mathematical objects are discrete, and it is possible to point at them.
Mind and Body
In the discussion about p-zombies, I feel that rationalists reject the argument about the mind and the body being conceptually independent too fast. I think the argument is possibly irrelevant, but not wrong. A and B are conceptually independent if you can model A and B with two sets of axioms such that not every model (in the Model Theory sense) for one is a model for the other. At least the direction that minds can be thought without bodies seems obvious to me. It is still true that thinking of a body that doesn’t think might be a contradiction.
Kant, Jaynes, Korzybski
Naively, some thoughts correspond to the world as some don’t.
Naively, Kant’s Copernican turn is understood as having eliminated the distinction between right and wrong thoughts about the world.
In reality, it was the elimination of a world-inaccesible-to-thoughts as criterion, and in the redefinition of right and wrong thoughts through reference to thought-coherence, etc.
Traditionally, some processes are “objectively random”. Some of our models capture those processes, whereas others capture our ignorance.
Jaynes could be misinterpreted as claiming there is no difference between knowledge that captures properties of the universe and made-up models. In fact, he is eliminating the reference to probability distributions that lie beyond mine and that I use to evaluate mine, and redefining “knowledge that captures properties” and “made-up knowledge” without reference to external probability distributions.
Scientific knowledge is also a probability distribution that I have accepted, not a property of an external universe. Every scientific theory assigns a certain probability to the measuring instruments giving wrong results, etc. But if we suddenly saw a lot of results in the area to which our model assigns low probability, we’d discard the theory. So it is not true that we consider some probability distributed “located in the external universe”.
*
If the territory is not something external to the maps, but a certain equivalence relation between them, then it has a different nature and no map is right.
*
FDT can’t establish a difference between real an imagined worlds, like Kant’s intuition not being a conceptual property.
*
The anthropic principle, and computability theory, are an imposition of principles of the understanding (=math) on the experience, without this being psychologism.
Skepticism about incomprehensible things
C’est une maladie naturelle à l’homme de croire qu’il possède la vérité directement ; et de là vient qu’il est toujours disposé à nier tout ce qui lui est incompréhensible ; au liez qu’en effet il ne connaît naturellement que le mensonge, et qu’il ne doit prendre pour véritables que les choses dont les contraire lui paraît faux. Et c’est pourquoi, toutes les fois qu’une proposition est inconcevable, il faut en susprendre le jugement et ne pas la nier à cette marque, mais en examiner le contraire ; et si on le trouve manifestement faux, on peut hardiment affirmer la première, tout incompréhensible qu’elle est. Appliquons cette règle à notre sujet.
Double interest
When someone makes a claim, I am interested both in the claim as a feature of the territory I know (this person’s mind), and as a feature of a potential map for a territory I don’t know (what they are describing).
Society
Community & Culture
I’m confused about communities. I feel that there really aren’t communities anymore in the West, but I can’t find a definition that lets me exclude things like K-Pop fans maniacly discussing what real K-Pop is.
It feels like culture should be exclusively dependent on the consensus of the members of a group. It feels like mathematicians judging other mathematicians for being wrong is not a culture, because there is an objective criterion of truth that they are simply applying. But don’t religious people think that they are applying God’s objective criterion? In an toy model, like women regularly attending a market who develop rules about where the queu should be, it seems easier to separate the objective things they came here to do (buying) from the cultural rules (some queues are wrong).
Gender
Overcoming gender roles is about learning an opposite set of skills. Men are taught to be independent. Women are taught that there is no shame in asking for help. Both are great skills! Learn the one you weren’t taught and you’ll be happier.
Some skills are more difficult to see. Men are taught not to cry, which is good make others willing to have certain conversations with you. But what is crying good for? Crying is a biological process that allows you to let go of an emotion, so you can continue a conversation with a fresh mind. Men usually don’t have this skill, and stay physically tense for hours after a conflict, even if they aren’t emotionally upset any more. Thinking usually doesn’t get you out of this. If you are a man, consider learning to cry. (h/t Paul and Ivona)
Humor
One funny German once said to me: There is a lot of German humor; it’s just not funny. French contrepèteries are funny; why isn’t Der Wechstabenverbuchsler funny? (cfr. Hofstadter on translating humor)
The Frenchman is very fond of being witty, and will without scruple sacrifice a little of the truth to a good turn of phrase. Where wit is out of the question, on the other hand, he shows just as thorough an understanding as a man of any other nation — in mathematics, for instance, and in the remaining dry or deep arts and sciences. For him a wisecrack does not have the fleeting value that it has elsewhere; it is eagerly passed around and preserved in books, like the most important of events.
If Hitler had been English would we, under similar circumstances have been moved, charged up, fired by his inflammatory speeches, or should we have laughed? Is English too ironic a language to support Hitlerian styles, would his language simply have rung false in our ears?
Religion
The word “religion” is used to encompass phenomena that might not have much in common. Pascal’s Christianity is a justification-less faith contraposed to science. But it’s unclear whether this is Aquinas’ meaning of religion: is there a big difference between obtaning first principles from faith or from the aristotelean nous?
Phenomena like Ancient Greek religion or hinduism seem even further away: Greeks believed in Zeus like I believe in New York: people say a lot of things about it, some are contradictory, but most of it has to be true or else people wouldn’t talk so much.
Language
Language as static structure vs language as (stochastic) production rules
Teaching
Teaching at a high school I had the impression that the underlying guideline for teachers was “you are responsible for the students’ emotional wellbeing, but not for whether they learn or not”, whereas my natural inclination is to be responsible of whether they learn or not (i.e. adapt my teaching to them), but not of their emotions (i.e. accept that some will get frustrated by the difficult, or by problems not connected to my class)
Other
Meta
Creating weird events and getting a lot of rationalists to attnd is easy in a place like Berkeley is easy. In a big city with no established rationalist community it seems to be more difficult than one would naively think. It is easy to gatekeep, but then the group doesn’t grow. It is easy to get a lot of people to attend, but then the target audience gets diluted and those who most centrally belong to it don’t want to attend. Would it be a solution to create an event and charge 50€ for attendance, but announce everyone enters for free the first time, and then waive the fee for people the organizer finds interesting? How could this be done in a way that is maximally transparent without being too insulting for the people who get rejected? One option would be to have middlemen with the right to waive the fee of a specific group (intelligent, kind, talkative, famous, well-connected people), and evaluated on how well they chose, and to outsource the ugly “you aren’t welcome” to “find a middleman who waive your fee”. (Tebaldo parties)
*
I like that my posts on rationality/AI/etc are subject to the votes of the LW community. And I like having all my posts in one place. At first I thought my dream situation would be that the personal section doesn’t have karma. But this would very fast degenerate into the personal section being de facto a generic blogging platform, which could be used by millions of people and flood LW. So this is no solution. But I think it would b possible to have, unless some conditions (e.g. only 1 in 20 posts) a karma-free post, as a way of saying: I want to post this, this belong to the whole person I am, but I understand this is not something the LW audience would enjoy. I wonder if it has some obvious failure modes like the first scenario.
Funny quotes from SSC/ACX
To this day I believe I deserve a fricking statue for getting a C- in Calculus I. It should be in the center of the schoolyard, and have a plaque saying something like “Scott Alexander, who by making a herculean effort managed to pass Calculus I, even though they kept throwing random things after the little curly S sign and pretending it made sense.” (2015-01-31)
Connections
Icebergs
[Prejudice is inadequate deduction], a major premise which is mostly submerged, like an iceberg.
Northrop Frye
Fatbergs, Hemingway’s iceberg, Synagoge as iceberg-tip of Tribe
Started at Inscribe.
- ^
The first three quotes come through Dynomight.
- ^
E. O. Wilson, through Gwern.
Math, Rationality, Funny quotes from SSC/ACX—feel like too small, unoriginal topics. Meta—skipped this section. Political philosophy—didn’t get it. AI—how does Platonism connect to persuasion?
The shame caused by your message pushed me to write up the two political posts. I wouldn’t have written them otherwise, so thanks a lot for that! (And I wish more people would give honest, negative feedback)
Wiki-libertarianism
Kantian consent
Thanks for the honest criticism; I don’t disagree with it.
re AI: they don’t, but it wasn’t clear. Now it is, thanks!