Judging ethical theories by update rules, not by action rankings
TL;DR: It’s useful to think of ethical theories as online learning algorithms rather than specific action-recommendations. As finite beings, we use them to compress accumulated moral judgments and incorporate new ones; different representations update differently. We can expect to spend our entire lives mid-moral update rather than in reflective equilibrium. So theories should be judged by their updates in the expected information environment, not just their initial checkpoint.
Asserting the ought
Assertion of recommendations. Some advocates of utilitarianism, deontology, etc. will tell you their ethical theory is superior because it produces the right results.
Assertion of form. Others will tell you their theory is correct because to take the right action is to follow the right rules / to maximise welfare itself. What it means to be moral is to have an ethics shaped in the form of their theory.
I believe these are confusing ways to think about ethics, for the meta-ethically curious. Meta-ethically curious meaning being curious about what the hell ethical theories are anyway, where they come from, and what our interest in them is.
The ethical theories were made by man, not man by ethical theories
If you make the phenomenological move of bracketing out all your previous intuitions and potentially flawed concepts and simply observe what is the case:
Social primates began forming a multi-agent society.
Those that developed elementary notions of empathy got selected for by evolution.
This proto-empathy developed into moral intuitions felt and expressed strongly by members of such societies.
As man developed reflective verbal intelligence, he developed systems for further reflecting on his intuitions. He did this out of the same instincts towards the beautiful and right that evolution produced in him.[1]
Ethical theories are cognitive assists created by man.
They ultimately answer to the extended moral reflection of man.
This extended moral reflection has something to do with both the universal reality of multi-agent societies and with the contingent historical facts about man’s evolutionary niche.
Sorting algorithms
All sorting algorithms output sorted lists, but
the ways they do it are very different
you’d recommend different sorting algorithms under different circumstances
some algorithms are just generally better than others
You can think of ethical theories the same way. All ethical theories try to approximate the “true ought,”[2] but
the ways they do it are very different
you’d recommend different theories under different circumstances
some are generally better than others
Seen this way, disagreements between ethical theories aren’t only about substantively different facts of the matter, but about different properties of estimators.
What algebra did for numbers, theory did for oughts
The problem ethical theories solve is that we are small, poor, grovelling, mortal beings, endowed with some basic moral intuition, capacity to deliberate, and a knack for systematisation.
We began with a million messy judgments: it seems right to give children food, to consider the intentions behind actions, to care for those closer to us more, no, wait, to care for those further away too.
We assembled these into dramatically simple overarching principles:
Maximise welfare.
Be nice and normal.
Act with honour and integrity.
These appealed to our varying sensibilities, and moreover, they structured the sum total of our moral reflections into a framework we could use every day and reflectively adjust over time.
Like algebra let us represent in a sentence sums that took us pages to describe, our ethical systems let us represent vast domains of judgments in dense symbols.
We shuffle oughts around fluidly, abstractly, at paces dizzying to any pre-verbal person’s concrete considerations. Ethical theories are one of the algebras of civilisation.
(source)
Multi-armed banditry
But compression alone isn’t enough: the thing we’re compressing keeps changing.
Consider further the human situation: we require an online learning algorithm, for assembling the sum total of our moral reflections so far and incorporating new ones.
Ethical theories are firstly compact representations, and secondly, online learning algorithms for updating them.
People focus on the object-level action-recommendations. But this is just the current checkpoint of your learning algorithm. With enough sophistication, very different ethical frameworks can be made to reproduce similar action-rankings.[3][4]
More interesting are the online learning dynamics.
The geometry of moral updates
In utilitarianism, you maintain a model of reward and how actions map to outcomes. However, the mapping from those models into policy space is chaotic: small updates can lead to taking actions that would have been considered extremely Uncool the previous day. A noisy local estimate of consequences can lead to something disastrous, and a group of people doing this can erode trust in a system built over much time. Utilitarians can be hard to trust for this reason, and this is part of the reputation of utilitarianism as leading to monstrous actions. That said, consequentialist representations can be wonderfully fluid. “Okay, but what actually happens?” can get past a lot of ritual and normativity. Flexible under novelty; vulnerable to bad world-modeling and motivated bullshit.[5]
In deontology, you maintain a ruleset. This creates immense stability, especially in a multi-agent context, which is why deontology is the observed theory of conservative communities. Constraints prevent noisy local estimates from rewriting good enough policies that encoded wisdom accumulated across many situations and generations. But updates can be slow and lead to crises in the system when enough pressure builds up. Robust against local bullshit; brittle out-of-distribution.
In virtue ethics, you maintain a set of virtues and choose actions that reinforce them. Asking “what kind of person am I becoming?” integrates the consequences of repeated actions that might look negligible one by one. Updates are smooth, but they can zero out in local minima of what a certain culture considers virtuous, or in a defensive self-identity. Rich contextual generalisation; but hard to specify, audit, or transmit.
The choice of representation induces different update dynamics.
Different algorithms for different regimes
When consequences are concrete and measurable, and readily yield themselves to better estimates with more effort, consequentialist models are excellent; introduce them to your philanthropic grant-making org.
When you want a system to survive generations of adversarial input and create predictability, deontological models may be a good way to structure the laws of the country you’re founding.
When living through chaotic times where the true sign (positive or negative) of various actions is hard to know in advance, a virtue model may make for the best outcomes.[6]
The AGI transition
The AGI transition seems like a particularly bad regime for naïve direct consequentialism: deeply out-of-distribution, under adversarial pressure, and with large moral uncertainty.
Anthropic, despite being a very consequentialist organisation, appears to be trying to give Claude the capacity to generalise values deeply and fluidly while remaining bounded by deontological constraints.
Though, consequentialists that they are, they justify even that consequentially, and I think this shows up in Claude.
You can observe what kind of information you expect to get in a particular environment and choose a suitable update rule.
Asking “are you a utilitarian?” would be a bit like asking “are you a quicksortist?” Maybe you use quicksort almost all the time. But you wouldn’t believe in quicksort. Every so often, you need to sort a linked list.
You may even consider an ensemble method.
Ensemble morality
Theories can remain useful even if you don’t buy their metaphysical justification.
“Consequences aren’t all that matter,” but you still run a consequentialist pass over a decision to catch issues.
“Deontology sucks,” but you preserve bright-line rules because you know motivated reasoning is strong.
“Virtue ethics is too underspecified to be a foundational theory” but you still ask what kind of a person a choice is training you to become.
“Several moral theories” is not a diplomatic compromise between mutually inconsistent metaphysical claims, but an ensemble method for reflecting on our moral reflections.
It’s natural to not select one ethical theory at all, but retain several biased representations and treat disagreement between them as a sign the current case deserves more consideration. This is, in practice, what most people do.
Continual inequilibrium
I don’t expect us to reach final reflective equilibrium in the course of our lives. We hurtle forwards in a forever fog of war.
The main criteria for ethical theories, then, are about how they perform for finite beings incorporating messy bits of information over time; not final elegance in some sort of complete information world.
Conclusion
Morality is that subset of our values that are interpersonal, prosocial, and to some extent expected to be agreed-upon.
Values themselves are complex, evolving features you can describe human beings as having that let you explain their preferences, particularly those that are consciously and reflectively endorsed.
We are part of an intergenerational dance, a promethean fire, that kept certain values alive, and within the course of our individual lives, our values develop and individuate further. There are no true values written under our skin; it’s one big living mess.
Keeping this mess going is important to us. We do this by continually incorporating moral data[7] in an online learning process we expect to keep going for the rest of our lives.
Ethical theories are compact representations of our judgments that induce different update dynamics. Our moral theories shape our moral trajectories.
It’s useful to think of ethical theories as online learning algorithms, that perform differently in different expected information environments.
Appendix A: Maximum entropy morality
The premise of this post is that we as finite beings experience moral uncertainty.
Why shouldn’t we model this the same way we model epistemic uncertainty[8]—with priors over our preferences, that we narrow over time with more moral reflection?
Statistical methods for dealing with uncertainty often equate to keeping maximum entropy—not being any more biased towards a particular hypothesis than your evidence lets you be.
Similarly, we might want superpowerful AI to preserve moral uncertainty—compatibility with many different futures as we figure it out, and progressively narrow in.[9]
This addresses the core issue with trying to solve morality in one go before AI takeoff (namely that it can’t be done). The issue with trying to argmax a particular utility function is it assumes too narrow of a prior on a particular idea of our preferences. It’s an unjustifiably narrow delta prior.
Prudent outer alignment may be to have it maintain moral uncertainty alongside us.
More in @beren’s post on Maximum Entropy Morality, Metaplastic Constitutionalism, and the Dynamic Virtues.
Appendix B: Convergence vs development
There may not be some pre-determined point to speak of us as converging upon. Where we go next may well be non-universally contingent on whatever wayward path we took. Religious freedom arose as an answer to the bloody wars of Europe. The right to bear arms arose as an answer to monarchical oppression.
philosophers were really out trying to do reward specification and red teaming each others with things like “utility monster,” “freedom monster,” never realising that morality lay in imitating existing behaviours and extending them within an open-ended multi-agent collective that had no fixed or specifiable end (source)
Our systems can still help organise our sense of all that’s happened before and refine our sense on where to go next. But there may be no fixed point to arrive at.
Developmental frame. We continue to build on our moral values in an open-ended way.
Uncertainty frame. There’s some true moral ought we are increasingly narrowing in on.
Fun Theory somewhat reconciles these—we can affirmatively narrow in on an experience of continued development as the thing we desire.
Appendix C: Tools vs Laws of Morality
A Consequentialism-lover may say:
Such equivocation! These other theories are good tools, but if we step back for a moment, surely we must admit consequentialism is the ultimate Law?
Or a Deontology-lover may object:
Such bias! Do you not see, that even in your every analysis of ethical theories, you judge them by the consequences they produce? You departed from me the moment we began.
In Toolbox-thinking vs Law-thinking, Eliezer argues Bayesianism is The Law of probability—it is how the very shape of reality works, that other tools merely try to help us navigate.
In some weaker sense, Consequentialism might be said to be the Law of morality—ultimately, it was consequences that drove moral intuition into man. It is the shape of reality from which morality even arose in the first place.
This doesn’t make consequentialism great as an everyday tool. But it means it’s especially worthwhile to study the Law of consequentialism. To reason in theorems about super-cooperation and formalise moral uncertainty as uncertainty about preferences, maximum entropy utility, and so on.
- ^
The “moral” instinct then being that subset of our values that are interpersonal, prosocial, and to some extent expected to be agreed-upon.
See @ahbwramc’s reaction to this point:
I’m not entirely sure why, but this comment was inordinately helpful in doing away with the last vestiges of confusion about your metaethics. I don’t know what I thought before reading it—of course morality would be a subset of our values, what else could it be? But somehow it made everything jump into place. I think I can now say (two years after first reading the sequence, and only through a long and gradual process) that I agree with your metaethical theory.
- ^
The “true ought” may not be a destination so much as a vector field to steer by.
- ^
E.g., Don’t lie is boring, even if it universalises. What about… hmm… Don’t lie, unless your girlfriend asks you if you’d love her if she was a worm. That one universalises fine! You can add arbitrary qualifiers to the categorical imperative. If everyone maximised utility subject to some constraints…
An ethical theory can even simulate and delegate out to a whole other theory within it (like two-level utilitarianism). They’re Turing complete specifications for following an action!
Any plausible theory can be consequentialised, and consequentialised theories can be deontologised.
- ^
You might even expect sophisticated enough theories to come to agree with each other when viewing the same reality.
See also this tweet and this comic.
- ^
Some fixes for this particular property of utilitarianism are quantilization, KL penalties over policy space, etc.
- ^
I find this relevant to working in AI safety. See:
- ^
“Moral data” like experiences in the world, extended reflection, perspectives from others.
- ^
Where epistemic uncertainty is uncertainty about how the world will react, moral uncertainty is uncertainty about how we will react to it.
To an AIXI-like agent, there may not be much of a difference? Maybe it’s all part of a singular world-model that continues to gain certainty about what kind of agent it is.
- ^
This is similar to the CIRL/off-switch game argument, that an agent certain of its objective has no instrumental reason to accept correction, but an agent uncertain of it might treat human pushback as evidence and defer (with caveats).
Do they? If we’re acknowledging the evolutionary nature of moral intuition, surely some theories recognize and accept that my intuition is going to be different from Bob’s intuition. It doesn’t feel like “true ought” once we’re talking about differences between my true ought and Bob’s true ought.