I greatly enjoyed reading your post and look forward to more from you! I’m also relatively new (AI Safety research since Feb of this year, research in general since fall 2023), and it’s an incredibly important field. A couple comments on your post:
I think it’s worth considering the confound between misalignment and general capability failure in measurement. In your first example of the stock on credit card advice seeking, the model may give bad advice due to actively wanting to give bad advice or simply failing to give good advice. However, I don’t think this confound particularly affects your conclusions given the point is not to report a static number, but to compare between ablations.
It’s a fairly well documented problem that training B on top of a model who has already done training A might improve at B but sometimes subsequent gets worse at A (a big safety problem). Some are experimenting with combining training steps.
Also, one newbie to another, I’ve found great success and learning opportunities from directly meeting or chatting with seasoned researchers (both virtual and in-person). I regularly cold email researchers and tell them who I am and why I am interested in their work. When I travel to the bay area (not sure where you are located), I send emails specifically targeted to people there and put “want coffee?” in the subject line. We researchers are all nerds who want to yap to other nerds about the stuff we find interesting and cool. I’ve learned a lot, gotten specific and career advice, and have gotten people to vouch for me in that way.
I greatly enjoyed reading your post and look forward to more from you! I’m also relatively new (AI Safety research since Feb of this year, research in general since fall 2023), and it’s an incredibly important field. A couple comments on your post:
I think it’s worth considering the confound between misalignment and general capability failure in measurement. In your first example of the stock on credit card advice seeking, the model may give bad advice due to actively wanting to give bad advice or simply failing to give good advice. However, I don’t think this confound particularly affects your conclusions given the point is not to report a static number, but to compare between ablations.
It’s a fairly well documented problem that training B on top of a model who has already done training A might improve at B but sometimes subsequent gets worse at A (a big safety problem). Some are experimenting with combining training steps.
Also, one newbie to another, I’ve found great success and learning opportunities from directly meeting or chatting with seasoned researchers (both virtual and in-person). I regularly cold email researchers and tell them who I am and why I am interested in their work. When I travel to the bay area (not sure where you are located), I send emails specifically targeted to people there and put “want coffee?” in the subject line. We researchers are all nerds who want to yap to other nerds about the stuff we find interesting and cool. I’ve learned a lot, gotten specific and career advice, and have gotten people to vouch for me in that way.