The distinction between inner and outer alignment is quite unnatural. For example, even the concept of reward hacking implies the double-fold failure of a reward that is not robust enough to exploitation, and a model that develops instrumental capabilities as to find a way to trick the reward; indeed, in the case of reward hacking, it’s worth noting that depending on the autonomy of the system in question, we could attribute the misalignment as inner or outer. At its core, this distinction comes out of the policy <-> reward scheme of RL, though prediction <-> loss function in SL can be similarly characterized; I doubt how well this framing generalizes to other engineering choices.
My best attempt at attempting to characterize Kant’s Transcendental Idealism -
Kant’s idealism says that essence—not existence—is dependent on us. That is to say, what it is to be is dependent on how we understand. For example, the schema of classification in biology, such as genetic proximity, depends on what purposes they serve to us. What it is for animals to be depends, in other words, on the biologist. To draw the biology analogy ad absurdum, transcendental idealism says something like “the genetic composition is the condition of the possibility of how we are able to make sense of biological objects in the first place”. The existence of these classification schema is dependent on our mind a priori.
I wish people reported their P(catastrophe) : P(doom)[1] ratio in addition to timelines. Whenever I hear that people’s p(doom) is at e.g. 5%, I immediately wonder, what about their p(catastrophe)? Is it 10%, or 20%, or 50%?
There’s a few reasons why I am curious about this ratio.
One simple reason is it serves as a consistency check. I think a lot of the times doom is used to mean ‘something really bad’, but actually catastrophe already qualifies as ‘something really bad’. So it’s good to reflect on whether you’d endorse what your p(doom) would imply about p(catastrophe).
A second reason is that I think the p(catastrophe) : p(doom) ratio reveals interesting information about people’s threat models, once people actually take them to denote different things. If the ratio is low, then plausibly their threat model is more dominated by immediate takeover risks. I for one think the ratio is really high. I think catastrophe at our current pace is very plausible, but e.g. extinction is far, far more implausible. To put some number to things, I think p(doom) is at <0.2%[2] and p(catastrophe) is at >20%[3] i.e. P(catastrophe) : P(doom) is at about 100:1, or equivalently given my definition p(doom | catastrophe) < 1%.
For the sake of precision here, we can say catastrophe means >10% of the global population dies in excess of baseline mortality rates in a sequence of events before 2100, and say doom means extinction before 2100. I don’t think the definition is critical here (it’s there to make rigorous the difference between ‘really really bad’ and ‘bad’), and I think you should report this ratio even when you disagree with the definition.
The distinction between inner and outer alignment is quite unnatural. For example, even the concept of reward hacking implies the double-fold failure of a reward that is not robust enough to exploitation, and a model that develops instrumental capabilities as to find a way to trick the reward; indeed, in the case of reward hacking, it’s worth noting that depending on the autonomy of the system in question, we could attribute the misalignment as inner or outer. At its core, this distinction comes out of the policy <-> reward scheme of RL, though prediction <-> loss function in SL can be similarly characterized; I doubt how well this framing generalizes to other engineering choices.
My best attempt at attempting to characterize Kant’s Transcendental Idealism - Kant’s idealism says that essence—not existence—is dependent on us. That is to say, what it is to be is dependent on how we understand. For example, the schema of classification in biology, such as genetic proximity, depends on what purposes they serve to us. What it is for animals to be depends, in other words, on the biologist. To draw the biology analogy ad absurdum, transcendental idealism says something like “the genetic composition is the condition of the possibility of how we are able to make sense of biological objects in the first place”. The existence of these classification schema is dependent on our mind a priori.
I wish people reported their P(catastrophe) : P(doom)[1] ratio in addition to timelines. Whenever I hear that people’s p(doom) is at e.g. 5%, I immediately wonder, what about their p(catastrophe)? Is it 10%, or 20%, or 50%?
There’s a few reasons why I am curious about this ratio.
One simple reason is it serves as a consistency check. I think a lot of the times doom is used to mean ‘something really bad’, but actually catastrophe already qualifies as ‘something really bad’. So it’s good to reflect on whether you’d endorse what your p(doom) would imply about p(catastrophe).
A second reason is that I think the p(catastrophe) : p(doom) ratio reveals interesting information about people’s threat models, once people actually take them to denote different things. If the ratio is low, then plausibly their threat model is more dominated by immediate takeover risks. I for one think the ratio is really high. I think catastrophe at our current pace is very plausible, but e.g. extinction is far, far more implausible. To put some number to things, I think p(doom) is at <0.2%[2] and p(catastrophe) is at >20%[3] i.e. P(catastrophe) : P(doom) is at about 100:1, or equivalently given my definition p(doom | catastrophe) < 1%.
For the sake of precision here, we can say catastrophe means >10% of the global population dies in excess of baseline mortality rates in a sequence of events before 2100, and say doom means extinction before 2100. I don’t think the definition is critical here (it’s there to make rigorous the difference between ‘really really bad’ and ‘bad’), and I think you should report this ratio even when you disagree with the definition.
FWIW, I think this is terrible given that our base rates are a few OOMs lower.
Could be way higher, but >20% is the number I’m comfortable committing to.