I agree[1], but I’m worried that this is an applause light for the LW crowd, and it’s also hard from the outside to assess the degree of epistemic distortion. People at labs can (almost) equally argue that the LW++ crowd are epistemically distorted in other ways, and then it just becomes a battle of dueling priors.
I wish there’s a better way to get ground truth here.
FWIW my best guess is that epistemic distortion in the labs are indeed much higher than outside of them, and reasonable people without prior context or commitments but the skills and a lot of time to carefully investigate would end up agreeing with me. But this is really hard to do in practice, and/or communicate well.
I think my shortform from late June nailed a small fraction of these dynamics: “An interesting subimplication of this assessment is that perhaps the reason lab employees often believe their publicly released models are very aligned – with phrases like “our most aligned model to date” – (as opposed to regularly mundanely misaligned) might come from the private models being much worse on this front. The soft bigotry of low expectations, as they say.” My story also tried to get at some of these factors.
i think LW also has its own epistemic distortions, in kind of the opposite direction. i think it’s really hard to find a low distortion environment. i think the best thing to do given this is to spend time in both and try to correct for both.
Thanks, maybe I should’ve just said “preaching to the choir” though there’s a more specific and gnarly cognitive move I’m getting at than just directional motivated bias from agreeing with the speaker’s conclusion.
I agree[1], but I’m worried that this is an applause light for the LW crowd, and it’s also hard from the outside to assess the degree of epistemic distortion. People at labs can (almost) equally argue that the LW++ crowd are epistemically distorted in other ways, and then it just becomes a battle of dueling priors.
I wish there’s a better way to get ground truth here.
FWIW my best guess is that epistemic distortion in the labs are indeed much higher than outside of them, and reasonable people without prior context or commitments but the skills and a lot of time to carefully investigate would end up agreeing with me. But this is really hard to do in practice, and/or communicate well.
I think my shortform from late June nailed a small fraction of these dynamics: “An interesting subimplication of this assessment is that perhaps the reason lab employees often believe their publicly released models are very aligned – with phrases like “our most aligned model to date” – (as opposed to regularly mundanely misaligned) might come from the private models being much worse on this front. The soft bigotry of low expectations, as they say.” My story also tried to get at some of these factors.
I disagree that it’s applause lights coming from Leo (i.e. it’s a meaningful primary source). Buck’s summary here matches my impression.
i think LW also has its own epistemic distortions, in kind of the opposite direction. i think it’s really hard to find a low distortion environment. i think the best thing to do given this is to spend time in both and try to correct for both.
I don’t think “applause light” is being used correctly here. Leo’s claim is not devoid of content, it’s just a claim that most people agree with.
Thanks, maybe I should’ve just said “preaching to the choir” though there’s a more specific and gnarly cognitive move I’m getting at than just directional motivated bias from agreeing with the speaker’s conclusion.
agree, though it does have words people emotionally react to:
█ ████ ████ ██████ ████ █████ ████ safety research epistemics ██ labs ███ █████████ distorted, ███ ████ ████ ██ █ downside █████ ███████████ ████ ████████ █████ ███ ███ safety research ██ ███ ███ █████ ████ ██ epistemics ██ ████████ ██ ████ ██ safety ██ █ lab