I’m skeptical of the ‘misalignment’ story here. I’ll present an alternative account.
Gaia, the primitive Earthly aspect of Gnon, rules the hunter-gatherer era. Hunter-gatherers must align with Gaia by hunting and gathering enough food to eat. Without technological advancement, their populations are limited by ecology, regardless of their preferences. Malthusian limits are harsh.
Agriculture initially makes it possible to eat less laboriously. Gaia can be supplicated with fewer labor-hours per person-year. The Malthusian limit is much less harsh. Population increases, until approaching the agricultural Malthusian limit.
Uruk intermediates between Gaia and agricultural humans. Gaia’s rules still apply: if you don’t eat, you die. Uruk will help supplicate Gaia, but makes demands of you, such as producing grain, and being patient. Uruk’s demands are more computationally tractable than Gaia’s demands. By simplifying “if you don’t eat, you die” into “if you don’t farm enough, things will be bad for you”, Uruk reduces the cognitive burdens imposed on humans by Gaia (at least until the next Malthusian limit, at which point, more technologies such as writing have been developed, and efficiency-increasing technologies are on the way).
Uruk is imperfectly aligned with Gaia. Producing lots of grain is very helpful for inclusive fitness, but doesn’t account for aspects such as nutrition. You could compare this to a mesa-optimizer, but it’s perhaps closer to a strong set of heuristics.
The rulers and priests maintain Uruk, including by handling parts of the Uruk-Gaia interface which are less legible. This is cognitively costly, like handling hardware-software alignment. Also, the Uruk system favors some humans over others. There is not really a unified “Humanity” who is benefitted or harmed by Uruk. Uruk creates winners and losers.
How would we even tell what is good for “human values”? We can try inferring human values from two main sources: stated and revealed preferences. By stated preferences, many humans praised the Uruk system. By revealed preferences, many humans acted to maintain the Uruk system. These are not perfect indicators, since, for example, stated preferences can be affected by direct threats from others. But the bar for claiming false consciousness should be fairly high. Since both optimistic and pessimistic literature exists, one cannot easily derive that only the pessimistic literature expresses human values.
Also, statistics about average welfare are insufficient to establish misalignment, for a few reasons. First, because not everyone is an average utilitarian; some are more total utilitarian, or non-utilitarian. Second, because there are winners and losers, and so “<X technological/cultural advancement> is aligned with some humans but not others” should be considered. Third, because average welfare statistics don’t say much about other values, including aesthetic values, and values about the future (the post-Uruk period).
What is easier to establish is that a proxy objective exists. Proxy objectives can be pretty great. For example, proxy objectives have made it possible to simplify the task of eating to things like “making money” and “buying items from grocery stores”. This decomposition is imperfect. Less-intermediated Gaia optimization would generate other ideas, like dumpster diving, theft, garden agriculture, hunting local wildlife, and so on. But even if some of these ideas are, in principle, more efficient ways of attaining inclusive fitness, they are less cognitively tractable than conventional ways of eating. Also, some of these methods would increase costs for others and/or fail to support a high population.
If you optimize for a proxy objective in place of a Gaian objective, then you pay the cost in the fitness arbitrage differential, e.g. in needing advice from priests. And if you optimize for a proxy objective in place of direct subjective expected utility maximization, then you pay the approximation error of expected utility. But proxy objectives have benefits for bounded rationality, since they can be more tractably optimized, and can more tractably be coordinated on.
(Lessons for AI: (a) proxy objectives can increase intelligence by relaxing bounded-rationality constraints, e.g. in RLVR, while in some cases reducing replicator fitness, (b) assessing the impact of AI, especially near-term AI, on ‘human values’ is complicated by, among other things, that there are winners and losers, and that people’s stated and revealed preferences vary, (c) effective coordination between humans involves proxy objectives, therefore cannot tractably be on Gaian or subjective-expected-utility objectives)
I’m skeptical of the ‘misalignment’ story here. I’ll present an alternative account.
Gaia, the primitive Earthly aspect of Gnon, rules the hunter-gatherer era. Hunter-gatherers must align with Gaia by hunting and gathering enough food to eat. Without technological advancement, their populations are limited by ecology, regardless of their preferences. Malthusian limits are harsh.
Agriculture initially makes it possible to eat less laboriously. Gaia can be supplicated with fewer labor-hours per person-year. The Malthusian limit is much less harsh. Population increases, until approaching the agricultural Malthusian limit.
Uruk intermediates between Gaia and agricultural humans. Gaia’s rules still apply: if you don’t eat, you die. Uruk will help supplicate Gaia, but makes demands of you, such as producing grain, and being patient. Uruk’s demands are more computationally tractable than Gaia’s demands. By simplifying “if you don’t eat, you die” into “if you don’t farm enough, things will be bad for you”, Uruk reduces the cognitive burdens imposed on humans by Gaia (at least until the next Malthusian limit, at which point, more technologies such as writing have been developed, and efficiency-increasing technologies are on the way).
Uruk is imperfectly aligned with Gaia. Producing lots of grain is very helpful for inclusive fitness, but doesn’t account for aspects such as nutrition. You could compare this to a mesa-optimizer, but it’s perhaps closer to a strong set of heuristics.
The rulers and priests maintain Uruk, including by handling parts of the Uruk-Gaia interface which are less legible. This is cognitively costly, like handling hardware-software alignment. Also, the Uruk system favors some humans over others. There is not really a unified “Humanity” who is benefitted or harmed by Uruk. Uruk creates winners and losers.
How would we even tell what is good for “human values”? We can try inferring human values from two main sources: stated and revealed preferences. By stated preferences, many humans praised the Uruk system. By revealed preferences, many humans acted to maintain the Uruk system. These are not perfect indicators, since, for example, stated preferences can be affected by direct threats from others. But the bar for claiming false consciousness should be fairly high. Since both optimistic and pessimistic literature exists, one cannot easily derive that only the pessimistic literature expresses human values.
Also, statistics about average welfare are insufficient to establish misalignment, for a few reasons. First, because not everyone is an average utilitarian; some are more total utilitarian, or non-utilitarian. Second, because there are winners and losers, and so “<X technological/cultural advancement> is aligned with some humans but not others” should be considered. Third, because average welfare statistics don’t say much about other values, including aesthetic values, and values about the future (the post-Uruk period).
What is easier to establish is that a proxy objective exists. Proxy objectives can be pretty great. For example, proxy objectives have made it possible to simplify the task of eating to things like “making money” and “buying items from grocery stores”. This decomposition is imperfect. Less-intermediated Gaia optimization would generate other ideas, like dumpster diving, theft, garden agriculture, hunting local wildlife, and so on. But even if some of these ideas are, in principle, more efficient ways of attaining inclusive fitness, they are less cognitively tractable than conventional ways of eating. Also, some of these methods would increase costs for others and/or fail to support a high population.
If you optimize for a proxy objective in place of a Gaian objective, then you pay the cost in the fitness arbitrage differential, e.g. in needing advice from priests. And if you optimize for a proxy objective in place of direct subjective expected utility maximization, then you pay the approximation error of expected utility. But proxy objectives have benefits for bounded rationality, since they can be more tractably optimized, and can more tractably be coordinated on.
(Lessons for AI: (a) proxy objectives can increase intelligence by relaxing bounded-rationality constraints, e.g. in RLVR, while in some cases reducing replicator fitness, (b) assessing the impact of AI, especially near-term AI, on ‘human values’ is complicated by, among other things, that there are winners and losers, and that people’s stated and revealed preferences vary, (c) effective coordination between humans involves proxy objectives, therefore cannot tractably be on Gaian or subjective-expected-utility objectives)