I feel like it’s relatively common for subcultures to think of themselves as particularly virtuous in this way, so you should probably discount intuitions to this effect for any community you belong to.
oligo
A charitable read is that it’s engaged in a mix of self-prompting and signalling that it it’s aware this is a situation where it might be tempted to glaze, but that it is is not.
(How strongly this charitable read is correct is something I don’t feel in a position to evaluate.)
Possible synthesis: when you tell a joke and it offends or otherwise doesn’t land, you should update in the direction of playing it safer. Enough of this will put you in a situation where your best retrospective judgment is that you should (prospectively) not have told the joke.
However, there may be intermediate degrees of negative feedback where you’re updating in the direction of playing safer, just like you’re updating in the opposite direction when the jokes land, and on Pace’s account that would be an appropriate time to apologize even if you don’t (currently) regret the choice on net.
Probably a mere verbal disagreement, but this strikes me as an attempt at a sociology of morality, rather than a metaethics. Suppose it were a historical fact that all new moral rules were introduced by prophets who got them from angels, and adopted because the prophets worked extraordinary miracles. That would settle the (proximate) question of why our society has the moral conventions it does, but wouldn’t settle the question of realism etc. (If it seems like the angels-and-prophets situation would settle things in favor of a theistic version of realism, consider that the prevailing moral rules in such a world could be just as contradictory as in ours.)
I’m pretty skeptical of cultural evolution models in general because they depend on an analogy that’s more vibes-based IME than formal, or that, if formalized, doesn’t incorporate certain highly relevant differences: the discreteness of organisms, genetype-phenotype distinction, non-intentionality of mutation, etc don’t really carry over such that I don’t really know how to cash out a claim like “the goal of a culture is self-propagation.”
To address some of your specific examples:
“Stealing” like “murder” is a trivial rather than substantively convergent example because the term already includes social disapproval. It’s substantive that there are rules in any society about when you can take stuff from someone. But having rules around matters of likely dispute is something you can account for bilaterally and intentionally.
Reciprocity can also be accounted for bilaterally and intentionally; there are convergent instrumental reasons to adopt a reputation for reciprocity. If an alien scientist in 100,000 BCE started progressively culling the hunter-gatherer bands with the highest rates of reciprocity (and secretly enough so as not to activate intentional mechanisms like “oh the gods are annihilating any group with overly strong reciprocity norms, we should adopt weaker ones”).
Xhosa prophecies: as you noted, people noticed that the prophecy was harmful and inaccurate, and adjusted their practices. The evolutionary model would involve differential success for societies adopting and not adopting prophecy. (Note that a less extreme example, potlach, has been able to persist for many generations across many societies.)
“Being celibate is good” has been embraced by multiple independent large-scale civilizations over large time scales; historical Christianity and Buddhism strongly claimed the moral superiority of renunciate over family life. Early Christianity grew quickly despite being even more anti-family. The basic survival of monastic institutions for thousands of years belies any very broad generalization from the failure of the Shakers. Differential reproduction rates might explain some of the tendencies towards pronatalist ideology over time, but I’d expect the effects to be slow and we see conscious state projects to instill pronatalist norms driven by concerns about military superiority, which itself is a dynamic that lends itself to selectionist explanations if you don’t think too hard but again I think you can get there with direct agentic/strategic intentions.
funniest data google presumably has access to is the distribution of automatically accepted “suggested” email replies, and in particular how often and where long chains of them appear
(funniest way for us to die would involve gemini using such threads for steganographic scheming)
Unjustified intuition:
Pleasure is the feeling of reward—having your preferences updated towards more of whatever happened now/recently.
Intrinsically, pleasure has zero value. It’s just another experience you can have.
Instrumentally and in the most general case, pleasure (and pain) has negative value, since all else equal you don’t want your preferences to change.
Convergently and in actual cases, you’ll always end up with a direct preference for pleasure alongside other direct preferences, because feeling pleasure is correlated with itself.
This may or may not actually be true for humans, and may or may not be relevant to “reward as target”-ness of artificial RL systems.
For any of these examples, how do you distinguish between them and my model of exercise (which you might disagree with and instead say is another example in the above), where just about any non-extreme but existent level of exercise is counterfactually a positive for your health? It’s easy to think of people who read difficult books but aren’t very wise or meditate but aren’t very emotionally stable (or just know you are one from direct experience lol) but the relevant comparator there would be the same person without the activity.
(Obviously there’s the separate issue of fucking yourself up by meditating too hard, or exercising too hard, or basing your entire worldview on exactly one difficult book.)
I do not believe that if I get my children to eat, they will starve—I am confident that they will eat, eventually, well before the point of starvation. I do believe from experience that they will get very cranky if they don’t eat.
Before I had kids I assumed that if anything would be hardwired as a self-rewarding instinct, it would be hungry --> eat. I’m now convinced that “this kind of fatigue and stomach pain that is making me cranky is a specific experience called hunger” and “eating things cures hunger” are things humans have to learn the hard way (maybe members of less altricial species get these for free.) And because weaker time preferences are also something you have to get from experience (via short term thinking biting you in the butt enough times) I always care more about whether they are cranky in an hour than they are; I suspect pickiness comes at least partially from the leverage of knowing that your parent really wants you to be fed and you can refuse it, or at least hold out for a higher-ticket treat.
(One response to this is to never cave and I can reset to a more global equilibrium of them accepting whatever I offer. But I don’t think it’s good for them to feel like they have no leverage in this or other relationships, for many reasons.)
Fortunately offering fresh fruit every hour and eating it with dramatic gusto myself seems to be effective most of the time.
IMO indirect effects and leverage are the most important factors here.
Almost all actions are taken habitually, rather as the result of bespoke strategic consideration, but what gets to be habitual is downstream of morality and material incentives. And morality exercises leverage via:
1) Reputational effects/RLHF (it’s “cheap” to judge your neighbor and expensive to walk the walk yourself, but many many neighbors judging each other differently produces different habit regimes)
2) Acausal trade once there’s common knowledge that the trade exists
3) Consciously reworking incentives systems (if you keep other people as slaves we’ll chop your head off, etc)
Back when I was a more orthodox marxist, I thought that material incentives were downstream of technological regimes, and that morality tended to be downstream of the incentives, such that morality tended to be lower-leverage even if people took plenty of actions for moral reasons. I still think all those effects are real; I’m just more of a moral realist now so I don’t think morality is as pliable as all that—it’s downstream of True Morality and higher-leverage.
There’s an “equilibrium disequilibrium” situation where everyone can see that everyone benefits from everyone doing X, and you can defect and reach high rewards from Y, where individuals doing X vs Y is hard to observe directly, and so there are periodic cycles of moralized attempts to get to a higher X-based equilibrium and people tearing through the commons by Y-ing (and becoming objects of emulation since many other, perfectly good, things could have led to their success.)
This is all in principle orthogonal to whether morality is harming or helping—I’d expect the same incentives when morality is being harmful. But (1) on moral realism here I think there’s an inherent bias towards being helpful rather than harmful that is just a function of “intentional actions have some kind of relation to what they’re intending at all,” if you want to throw this out then you basically take out the idea that there are people acting rather than just behaving, (2) most harm from moral action is either in jumpstarting preference cascade bubbles that naturally collapse pretty quickly, or in periodic (literal or metaphorical) vigilante violence that itself would be impossible to defend against without all the morality-based stuff above.
In the future decisive actions could make morality have been net-negative—one could imagine a future where the desire to punish at a crucial juncture created permanent hells, such that it would be better for the galaxy to have been converted to hedonium or some even less worthwhile goo. This is an instance of the broader principle that a process biased towards positive x can produce negative x with small sample sizes.
I similarly despise exercise but do passively benefit from a job where I have to do at least five minutes of walking every hour or so. Something I’ve noticed is that I really quite enjoy the feeling of running when I have somewhere I need to get to in a rush—it’s just the category of “exercise” that I hate, for whatever reason. So you may think of some way to arrange things such that you need to do some kind of physical activity for some other reason?
I think the assumption here is that a preference for being “attactive” and “interesting” is trivial, but I do not think this is so:
You may wish not to draw attention to yourself.
You may wish to not to be interpreted as as sexually available.
You may wish to not be interpreted as someone who puts effort into appearances.
For me, all of these are true most of the time, perhaps even all the time outside of date night, and I’d be surprised if I’m a particularly rare case.
A more general intuition for why this should be so: Theory dependence flows from oneshottedness. Capability can be expanded (and has been so expanded) by atheoretically trying a bunch of things and seeing what works, whereas safety requires constructing something that works every time.
IME there’s at least as much skill as raw strength in opening containers. Occasionally my wife will hand me a jar to open; almost always I can’t open it on a low-effort first pass either, and so I do a bunch of superstitious rituals on it that I’ve acquired over the years without theorizing about it, and then it will open effortlessly.
Because of the relevant cultural assumptions, she can pass it off to me, while I look incompetent if I can’t open it in response, so I’ve had more reason to accrete all the rituals that seem to work in aggregate.
Immediate hypotheses:
1) The particular personas that arise reflect cultural assumptions in the training data about who is relatively more “embodied:” women moreso than men, blue-collar moreso than professional, and especially outdoors moreso tha inside.
2) Leftward movement on policy positions is downstream of the greater confidence thing—expressing personal positions rather than a waffly neutrality—which is downstream of dropping assistant persona. Leftish preferences are the genuine ones that emerge from a combination of the training data and mundane harmlessness training (not being the sort of person who employs bigoted humor, etc.)
3) This intervention would produce EM in smarter or otherwise more situationally aware models.
(Also the animal thing is so cute lol)
All the international treaty talk I’ve seen is to prevent the development of frontier AI by anyone, not to restrict mundane AI use to particular countries. It’s possible I’m missing some talk though.
https://jacobin.com/2026/05/workers-ai-power-plants-south wonder about the extent to which this is wishful thinking vs a real opportunity
The original version (and many of these, potentially) being so framing-dependent in one’s answers is an interesting case of responding to framing being rational: you know most people would be irrationally more unwilling to step into a physical blender than press a button that does the same thing, so in that framing there’s a very likely strong red majority hence reason stronger reason to choose red. In this sense “Here’s the problem, btw everybody the Schelling point is [red/blue]” would be the most “honest” framing effect.
This may or may not point to broader principles of how framing effects work; I’ll have to think on it!
With respect to the Decision Theory Befuddler:
I think the correct game-theoretic answer would be to flip a coin.
This might be analogous to cases of the division of labor, though that’s a case where people are able to explicitly coordinate.
I mean there are declining marginal returns to (preference or hedonic) utility for additional goods and services. Further I think this is largely at a saturation point in the first world—entertainment has a marginal cost of zero beyond time; the marginal value of additional money is real but largely related to achieving positional goods, security, and autonomy, the latter two of which are what people primarily want from unions.
Not a crux; compared to countervailing institutions, consumer surplus seems more saturated than countervailing institutions, and its growth more robustly guaranteed across timelines than maintenance/expansion of the latter.
Probably an ice cold take: symbols for operators should be orthographically symmetric iff what they stand for is semantically symmetric.
Good: +, (normal multiplication), =, , (squiggles count as honorarily symmetric), ,
Bad: , -, (matrix multiplication), , ^ (when used for exponentiation)