I would love to fund negative results, and have already earmarked retrofunding for one of my critics. You do not need to be pro-CAST in order to get CRF money. You simply need to advance our understanding of the topic in a way that helps the AI situation go well!
I do think there are ways in which existing alignment audits can get at corrigibility, and the H-only stuff is a good example. I am, personally, confused about the steerability stuff from that paper, and should probably think harder about it. If something cleaves close enough to corrigibility, I’ll consider rewarding it, even if it doesn’t talk about the concept directly. Feel free to suggest more examples.
I would love to fund negative results, and have already earmarked retrofunding for one of my critics. You do not need to be pro-CAST in order to get CRF money. You simply need to advance our understanding of the topic in a way that helps the AI situation go well!
I do think there are ways in which existing alignment audits can get at corrigibility, and the H-only stuff is a good example. I am, personally, confused about the steerability stuff from that paper, and should probably think harder about it. If something cleaves close enough to corrigibility, I’ll consider rewarding it, even if it doesn’t talk about the concept directly. Feel free to suggest more examples.