I only skimmed both of those two blogposts, but they feel different from what I want to do. So let me try to state my plan more clearly:
I think labs are currently way too secretive about their alignment / safety practices. So we don’t have good evidence about whether their method(s) work or not. (edit; and more importantly, why they will continue to work in the future! in fact i suspect there is no prosaic alignment technique that currently scales into superintelligence.)
If we do good science in the open, we will expose failure modes; these can be used as a foundation to demand evidence that labs have properly addressed these failure modes.
Basically I would like to not have to “just trust” that OpenAI / Anthropic have done safety / alignment sufficiently well; I think it’s important that independent third parties get the right to audit their alignment stack and draw their own conclusions about whether this is good / sufficient.
The fact that both labs have heavily redacted alignment science blogs is (to me) evidence that we clearly cannot just trust them to voluntarily tell us much about their alignment stack(s)
I only skimmed both of those two blogposts, but they feel different from what I want to do. So let me try to state my plan more clearly:
I think labs are currently way too secretive about their alignment / safety practices. So we don’t have good evidence about whether their method(s) work or not. (edit; and more importantly, why they will continue to work in the future! in fact i suspect there is no prosaic alignment technique that currently scales into superintelligence.)
If we do good science in the open, we will expose failure modes; these can be used as a foundation to demand evidence that labs have properly addressed these failure modes.
Basically I would like to not have to “just trust” that OpenAI / Anthropic have done safety / alignment sufficiently well; I think it’s important that independent third parties get the right to audit their alignment stack and draw their own conclusions about whether this is good / sufficient.
The fact that both labs have heavily redacted alignment science blogs is (to me) evidence that we clearly cannot just trust them to voluntarily tell us much about their alignment stack(s)