Thinking about AI Alignment and Reliability.
Yavuz Bakman
Karma: 34
very relevant work!
I think when the capabilities of the model increase, guardrails can be fooled by the models more easily, which is why it wouldn’t be a good solution at that time. But these days, they are still quite powerful, and I guess anthropic deploys guardrail models in production.
That’s an excellent idea! I believe a similar approach can be used for model capabilities as well, but it may also prevent benign users from updating their models as well. Still, achieving fragile capabilities for adversarial updates but preserving them for benign updates seems doable to me.
Almost all scenarios can be done by human beings too, and some people already have access and the power to execute those possible scenarios. Do you think all those people with access and power are aligned? So why aren’t you worried right now? Or are you?