which is the opposite of what the ‘alignment tax’ model would have predicted.
Well, sort of.
Training gpt-4 from gpt-4-base isn’t just an alignment task though, it’s a training task to obtain a more useful model (i.e. both more useful, and safe).
During post-training, for example, it’s trained to respond to answers, and also to not respond with dangerous answers. Responding to answers makes it (arguably) More Useful and More Dangerous. Not responding with dangerous answers makes it arguably Slightly Less “Useful” and Less Dangerous (i.e. not contradicting the idea of a positive alignment tax). So we’re kind of mixing two things, here.
Still, I agree that we could have negative alignment taxes. (Also very positive ones). And furthermore—negative or extremely small cost alignment mechanisms will be massively more likely to be adopted—so we should be really prioritizing research of alignment avenues that have potentially near-zero or ideally positive capability costs.
Well, sort of.
Training gpt-4 from gpt-4-base isn’t just an alignment task though, it’s a training task to obtain a more useful model (i.e. both more useful, and safe).
During post-training, for example, it’s trained to respond to answers, and also to not respond with dangerous answers. Responding to answers makes it (arguably) More Useful and More Dangerous. Not responding with dangerous answers makes it arguably Slightly Less “Useful” and Less Dangerous (i.e. not contradicting the idea of a positive alignment tax). So we’re kind of mixing two things, here.
Still, I agree that we could have negative alignment taxes. (Also very positive ones). And furthermore—negative or extremely small cost alignment mechanisms will be massively more likely to be adopted—so we should be really prioritizing research of alignment avenues that have potentially near-zero or ideally positive capability costs.