It seems like corrigibility only really helps if you keep a human in the loop, and I don’t think we’re likely to do that (since “the AI does what you meant without annoying clarifying questions” is a valuable capability)
I disagree. I see corrigibility as a choice to delegate goal choice to individuals rather than fixed baked-in abstractions. People can then choose how much control and agency to use at any particular point. “Go off and try to build X and only ask me questions if you’re genuinely super unsure what I’d want” is a reasonable ask for a corrigible agent. It should be able to figure out what you want and only ask you for input when you would’ve wanted it to.
I disagree. I see corrigibility as a choice to delegate goal choice to individuals rather than fixed baked-in abstractions. People can then choose how much control and agency to use at any particular point. “Go off and try to build X and only ask me questions if you’re genuinely super unsure what I’d want” is a reasonable ask for a corrigible agent. It should be able to figure out what you want and only ask you for input when you would’ve wanted it to.