Do you personally prefer to live in a world with a lot of slack (like today’s world) or in a Malthusian world where you’re barely above subsistence? The former, right? So even if Malthusian competition among AIs can recreate human-like values, that world would be dispreferable (have low value) according to those values!
In the original blog post, we think a lot about slack. It says that if you have slack, you can kind of go off the optimal solution and do whatever you want. But in practice, what we see is that slack, when it occurs, produces this kind of drift. It’s basically the universe fulfilling its naturally entropic nature, in that most ways to go away from the optimum are bad. If we randomly drift, we just basically tend to lose fitness and produce really strange things which are not even really what we value.
This random drift seems obviously caused by humans’ lack of sufficient intelligence/agency/cooperation, not some sort of fundamental problem with slack? Also, where is “the original blog post” referred to here?
This random drift seems obviously caused by humans’ lack of sufficient intelligence/agency/cooperation, not some sort of fundamental problem with slack?
Do you have a preferred formalism for this? UDT doesn’t have a standard temporal coherence theorem by default, since even allowing some deviation from EUM (geometric means, logical induction, etc.), the UDT “prior” can be reinterpreted as a probutility measure, with only minor distortion (magnitude of distortion conjectured acceptable for humans) caused by the structure of the UDT chains provided by your multiverse theory. Since UDT does no updating, there are no further conflicts between probability and utility[1].
@Vladimir_Nesov suggests being careful here[2] and his proposal[3]intentionally allows some value drift. I think this is more drift than you want, but unless I impose severe fairness restrictions on the set of situations I consider, I haven’t been able to avoid increasing constraints on the fixed point to the point that a human would have trouble distinguishing Nesov’s solution from certain ways of implementing a locked in utility function. Descriptively, this is quite close to the problem that almost all attempts at implementing “balance” fall to enough optimization pressure, unless the preferred outcome is proven to result[4]. Either the updateless core lets the aggregate substantially fall to manipulators, or it acts substantially like it has a fixed utility function (assuming modified Critch boundaries are enforced by Nesov’s sysops, and the majority of situations are close to locally SIA=SSA fair[5] (the latter presumably following the updateless cores of most humans)).
Do you personally prefer to live in a world with a lot of slack (like today’s world) or in a Malthusian world where you’re barely above subsistence? The former, right? So even if Malthusian competition among AIs can recreate human-like values, that world would be dispreferable (have low value) according to those values!
This random drift seems obviously caused by humans’ lack of sufficient intelligence/agency/cooperation, not some sort of fundamental problem with slack? Also, where is “the original blog post” referred to here?
Do you have a preferred formalism for this? UDT doesn’t have a standard temporal coherence theorem by default, since even allowing some deviation from EUM (geometric means, logical induction, etc.), the UDT “prior” can be reinterpreted as a probutility measure, with only minor distortion (magnitude of distortion conjectured acceptable for humans) caused by the structure of the UDT chains provided by your multiverse theory. Since UDT does no updating, there are no further conflicts between probability and utility[1].
@Vladimir_Nesov suggests being careful here[2] and his proposal[3] intentionally allows some value drift. I think this is more drift than you want, but unless I impose severe fairness restrictions on the set of situations I consider, I haven’t been able to avoid increasing constraints on the fixed point to the point that a human would have trouble distinguishing Nesov’s solution from certain ways of implementing a locked in utility function. Descriptively, this is quite close to the problem that almost all attempts at implementing “balance” fall to enough optimization pressure, unless the preferred outcome is proven to result[4]. Either the updateless core lets the aggregate substantially fall to manipulators, or it acts substantially like it has a fixed utility function (assuming modified Critch boundaries are enforced by Nesov’s sysops, and the majority of situations are close to locally SIA=SSA fair[5] (the latter presumably following the updateless cores of most humans)).
Indeed, because of this, a probutility semimeasure may be acceptable. Solomonoff, potentially unpublished letter.
https://www.lesswrong.com/posts/ZLary4FDY7kQaGfc3/existential-risk-from-ai-an-exposition-for-mathematicians?commentId=tBYP7bAhfNnokqFyZ
https://www.lesswrong.com/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a?commentId=BjQrqeKfov946oAKj
Presumably, to prove this you would first need to formalize metaphilosophy.
Though outdated, see Anthropic Decision Theory