One of the problems with performing independent AI safety research is that if you do actually end up coming up with something novel and sufficiently powerful you immediately run into a catch-22: you can’t broadcast the specifics (because safety research is capabilities research), and without a public track record it’s unlikely that anyone relevant will entertain a private review.
Everyone operates from a locally justifiable position, but it does make getting information into the hands of the people/organizations best suited to act on it responsibly difficult.
On the current margin most safety research is several times more impactful as safety research than capabilities research. If publishing allows you to either stop wasting your time or get a job where you have access to 10x more resources, it’s a huge net win even if capabilities progress you generate removes 20% of your impact, which is unlikely. So I think this issue is way overblown.
One of the problems with performing independent AI safety research is that if you do actually end up coming up with something novel and sufficiently powerful you immediately run into a catch-22: you can’t broadcast the specifics (because safety research is capabilities research), and without a public track record it’s unlikely that anyone relevant will entertain a private review.
Everyone operates from a locally justifiable position, but it does make getting information into the hands of the people/organizations best suited to act on it responsibly difficult.
On the current margin most safety research is several times more impactful as safety research than capabilities research. If publishing allows you to either stop wasting your time or get a job where you have access to 10x more resources, it’s a huge net win even if capabilities progress you generate removes 20% of your impact, which is unlikely. So I think this issue is way overblown.