Even if refusals don’t prevent catastrophic misuse, I would guess that model refusals would prevent the majority of potential harm.
I guess the crux is that I care less about this non-catastrophic harm compared to the benefits of focusing on intent alignment and corrigibility. I agree there’s a trade-off here but I am just taking a side on what’s overall best as per my worldview.
I guess the crux is that I care less about this non-catastrophic harm compared to the benefits of focusing on intent alignment and corrigibility. I agree there’s a trade-off here but I am just taking a side on what’s overall best as per my worldview.