I don’t think I had that misunderstanding exactly. I am disputing your characterization of the intellectual environment that “many others” experience from being at Constellation lunches, talks, and on slack (but perhaps this is a good characterization of what you experienced, or of being a junior Redwood employee, I’m not sure).
I don’t agree that the lunch/talks/slack have the biases you refer to (i.e. inductive biases arguments for misalignment over Goodhart, the belief that terminal training-gaming poses no risk, ignoring misalignment emerging/spreading during deployment, ignoring the limitations of band-aid solutions to alignment). I don’t see much evidence that there has been tons of discussion of inductive biases arguments for misalignment in the slack/talks/lunch conversations.
I now agree with your position that there is rare discussion of the alignment problem in lunch/talks/slack. I was largely thinking about my early years in Constellation, in which I talked about that subjects very frequently.
I don’t think I had that misunderstanding exactly. I am disputing your characterization of the intellectual environment that “many others” experience from being at Constellation lunches, talks, and on slack (but perhaps this is a good characterization of what you experienced, or of being a junior Redwood employee, I’m not sure).
I don’t agree that the lunch/talks/slack have the biases you refer to (i.e. inductive biases arguments for misalignment over Goodhart, the belief that terminal training-gaming poses no risk, ignoring misalignment emerging/spreading during deployment, ignoring the limitations of band-aid solutions to alignment). I don’t see much evidence that there has been tons of discussion of inductive biases arguments for misalignment in the slack/talks/lunch conversations.
I now agree with your position that there is rare discussion of the alignment problem in lunch/talks/slack. I was largely thinking about my early years in Constellation, in which I talked about that subjects very frequently.