You contradict yourself in the way you define the ontology. You write (emphasis mine):
I’ll use the term “deference-goodness” to refer to differential improvements to AIs being effective and aligned at key tasks (as in, differential relative to just increasing general capabilities).[16](I’m not including work which is just useful for reducing the chance that AIs are seriously misaligned in deference-goodness.)
And then later you write:
We can break down deference-goodness into an alignment/elicitation component and a capability profile component:
Broad alignment: [...] The AIs can’t be seriously misaligned.[18]
I think having this distinction is needlessly complicated:
Aren’t seriously misaligned: They remain corrigible to the intended group/structure and don’t try to seize power or kill humans.
Are effective and aligned at key tasks: The AIs do a good job—which includes being sufficiently aligned and sufficiently capable—of autonomously: [...]
(The way I am reading this is that on key tasks, the second criterion implies the first (so the first only is relevant on non-key tasks because the second doesn’t apply there). If you meant sth different I don’t know what you meant.)
If AIs are sufficiently aligned at the key tasks, they are likely also generally not seriously misaligned. I think you agree with this given footnote 16.
I think it would be less confusing if you had defined the ontology like:
AIs need to be in the BGD, which means they must be sufficiently aligned, and effective at key tasks: [...]
And then defining deference-goodness via that. (The fact that alignment on non-key tasks doesn’t necessarily need to be quite as strong is still expressed through that, but now only implicitly, because it’s not that important.) This would make the way you define broad alignment above correct again, although in other parts of the post you are also mentioning that preventing egregious misalignment isn’t part of broad alignment.
(I guess why you thought separating might make sense because the “aren’t seriously misaligned” maps to “the property that makes alignment through behavioral training possible”, but this doesn’t seem like a strong reason and the cost of having a confusing fuzzy ontology between broad alignment and preventing egregious/serious misalignment seems much larger. (Tbc, it’s totally fine to have a concept for “make sure tests are representative” or “preventing alignment faking”, but I think it would be better if that just counts as part of broad alignment.)) (Or maybe you think the methods for preventing serious misalignment are very different from achieving broad alignment? That would surprise me though.) (Or maybe you were first thinking more about elicitation and then realizing that post-handoff most of elicitation is captured by alignment? That would explain it but not be a good reason.)
(If you decide to keep the split, then probably rename “broad alignment” to “broad alignment at key tasks”.)
(Tbc, I’m not suggesting editing the whole post ofc, just saying what I think is the better ontology going forward. Whether you add an edit note is up to you.)
You contradict yourself in the way you define the ontology. You write (emphasis mine):
And then later you write:
I think having this distinction is needlessly complicated:
(The way I am reading this is that on key tasks, the second criterion implies the first (so the first only is relevant on non-key tasks because the second doesn’t apply there). If you meant sth different I don’t know what you meant.)
If AIs are sufficiently aligned at the key tasks, they are likely also generally not seriously misaligned. I think you agree with this given footnote 16.
I think it would be less confusing if you had defined the ontology like:
And then defining deference-goodness via that. (The fact that alignment on non-key tasks doesn’t necessarily need to be quite as strong is still expressed through that, but now only implicitly, because it’s not that important.) This would make the way you define broad alignment above correct again, although in other parts of the post you are also mentioning that preventing egregious misalignment isn’t part of broad alignment.
(I guess why you thought separating might make sense because the “aren’t seriously misaligned” maps to “the property that makes alignment through behavioral training possible”, but this doesn’t seem like a strong reason and the cost of having a confusing fuzzy ontology between broad alignment and preventing egregious/serious misalignment seems much larger. (Tbc, it’s totally fine to have a concept for “make sure tests are representative” or “preventing alignment faking”, but I think it would be better if that just counts as part of broad alignment.)) (Or maybe you think the methods for preventing serious misalignment are very different from achieving broad alignment? That would surprise me though.) (Or maybe you were first thinking more about elicitation and then realizing that post-handoff most of elicitation is captured by alignment? That would explain it but not be a good reason.)
(If you decide to keep the split, then probably rename “broad alignment” to “broad alignment at key tasks”.)
(Tbc, I’m not suggesting editing the whole post ofc, just saying what I think is the better ontology going forward. Whether you add an edit note is up to you.)