The first implicit pillar of the Berkeley Model that we want to criticize is the assumption of content indifference: The Berkeley Model assumes we can fully separate the technical problem of aligning an AI to some values or goals and thegovernance problemof choosing what values or goals to target. While it is logically possible that we’ll discover some fully generic method of pointing to goals or values (e.g. brain-reading), it’s equally plausible that different goals or values will effectively have different ‘type-signatures’: goals or values that are highly unnatural or esoteric given one training method or specification-format may be readily accessible given another training method or specification-format, and vice versa.
I’m glad to see someone articulate this point. On the other hand...
Consider that we humans ourselves manage to be respectful, caring, and helpful to our friends despite not fully knowing what they care about or what their life plans are—thereby providing an informal human proof for the possibility of beneficial and safe behavior without exhaustive learning of the target’s values. And as concerns sufficiency, the recent literature on deceptive alignment vividly demonstrates that value learning by itself can’t guarantee the right relationship to motivation: understanding human value and caring about values are different things.
Perhaps I’m confused, but this feels like what someone who misunderstood the Berkeley Model of Alignment and accidentally strawmanned it would write?
“Perhaps I’m confused, but this feels like what someone who misunderstood the Berkeley Model of Alignment and accidentally strawmanned it would write?”
You’d have to elaborate what you suspect the misunderstanding might have been, or what feels inconsistent here?
I’m glad to see someone articulate this point. On the other hand...
Perhaps I’m confused, but this feels like what someone who misunderstood the Berkeley Model of Alignment and accidentally strawmanned it would write?
“Perhaps I’m confused, but this feels like what someone who misunderstood the Berkeley Model of Alignment and accidentally strawmanned it would write?”
You’d have to elaborate what you suspect the misunderstanding might have been, or what feels inconsistent here?