i wish more people were aware that safety research epistemics at labs are extremely distorted, and that this is a downside worth considering when deciding where to do safety research or how one ought to comport oneself if choosing to work on safety at a lab
incentive to do capabilities, or things that are useful for capabilities, or ship in the product
incentive not to think too hard about asymptotics / ways approaches might fail if they are already working
making the company go faster and win is good (anthropic is especially bad at this one)
much stronger bias towards trusting your eyes vs abstract arguments (tbh, LW could use a little bit more of this one on the current margin, but too much is also bad)
a kind of subtle thing where it’s easy for all stories of impact to route through making tiny marginal changes to big things, rather than doing ambitious things
overestimation of the value of knowledge inside labs / the irrelevance of work outside labs
one way to observe some of this is to see what ways median lab people mispredict the world. for example, median lab people often underpredict capabilities progress and scary alignment failures. i think this is not a coincidene
To be fair, the capability mispredictions are shared by lots of other people, and suffice it to say that not even lab employees have abandoned normie priors on capabilities progress yet.
“Then, after you’ve started reading all this daily intelligence input and become used to using what amounts to whole libraries of hidden information, which is much more closely held than mere top secret data, you will forget there ever was a time when you didn’t have it, and you’ll be aware only of the fact that you have it now and most others don’t….and that all those other people are fools.
“Over a longer period of time — not too long, but a matter of two or three years — you’ll eventually become aware of the limitations of this information. There is a great deal that it doesn’t tell you, it’s often inaccurate, and it can lead you astray just as much as the New York Times can. But that takes a while to learn.
“In the meantime it will have become very hard for you to learn from anybody who doesn’t have these clearances. Because you’ll be thinking as you listen to them: ‘What would this man be telling me if he knew what I know? Would he be giving me the same advice, or would it totally change his predictions and recommendations?’ And that mental exercise is so torturous that after a while you give it up and just stop listening. I’ve seen this with my superiors, my colleagues….and with myself.
“You will deal with a person who doesn’t have those clearances only from the point of view of what you want him to believe and what impression you want him to go away with, since you’ll have to lie carefully to him about what you know. In effect, you will have to manipulate him. You’ll give up trying to assess what he has to say. The danger is, you’ll become something like a moron. You’ll become incapable of learning from most people in the world, no matter how much experience they may have in their particular areas that may be much greater than yours.“
[...]
Actually, later, in a much later conversation, the next year, I was with Kissinger in San Clemente and I was urging him to read the Pentagon Papers, which ended in 1968, he had a copy. And he said: But do we really have anything to learn from these documents? I said: Well, yes. I think you do. And my heart was sinking at this point. And he says: But we make policy very differently now. And I said: Well, Cambodia – the fiasco debacle that had just happened in the spring – I said that didn’t look so different. He said: Well, that was done for very complicated reasons. I said: Henry, every rotten decision in the last 20 years in Vietnam has been done for very complicated reasons and pretty much the same ones; political considerations, legislation in Congress, fear of being called weak, or you know, not giving the whole. [...] And he actually said to me: But they didn’t have clearances. And I thought, oh God.
I agree[1], but I’m worried that this is an applause light for the LW crowd, and it’s also hard from the outside to assess the degree of epistemic distortion. People at labs can (almost) equally argue that the LW++ crowd are epistemically distorted in other ways, and then it just becomes a battle of dueling priors.
I wish there’s a better way to get ground truth here.
FWIW my best guess is that epistemic distortion in the labs are indeed much higher than outside of them, and reasonable people without prior context or commitments but the skills and a lot of time to carefully investigate would end up agreeing with me. But this is really hard to do in practice, and/or communicate well.
I think my shortform from late June nailed a small fraction of these dynamics: “An interesting subimplication of this assessment is that perhaps the reason lab employees often believe their publicly released models are very aligned – with phrases like “our most aligned model to date” – (as opposed to regularly mundanely misaligned) might come from the private models being much worse on this front. The soft bigotry of low expectations, as they say.” My story also tried to get at some of these factors.
i think LW also has its own epistemic distortions, in kind of the opposite direction. i think it’s really hard to find a low distortion environment. i think the best thing to do given this is to spend time in both and try to correct for both.
i think LW also has its own epistemic distortions, in kind of the opposite direction
Fwiw I’m also skeptical of this line of argument/reasoning on meta grounds[1]. I’ve most often seen Anthropic ppl like Amanda and Dario deploy this. But I think it’s seductive for bad reasons if ppl think about it even a little bit.
tbc maybe your object-level reasons are sufficiently strong to believe this. Just wanted. to clarify the meta point here because it’s been bugging me for a while.
tbc, i think LW is much closer to sane! but i definitely do get frustrated by some of the cargo culting on here sometimes, and that was the main reason i said what i did. i mostly ignore ant/oai people shitting on LW.
Sorry my comment was unclear! I was less thinking of the specific shitting on safety people sociologically and more about something akin to the golden mean fallacy that I see prominent Ant people deploy in a bunch of seeminglyunrelatedincidences (dunno if you see it regularly at oai too).
Besides the normative aspect I also think they’re actually making a fairly simple descriptive mistake, which is that it’svery easy to see yourself as the center if you don’t define your bounds ahead of time, and/or if you observe this among ppl you interact with rather than from e.g. polling.
i’m not claiming you should in general be in the center of various things. for example, i don’t think you should spend some time with the far left and some time with the far right (you should spend time with neither); nor do i think you should spend some time with the religious and some time with atheists (i think you should spend almost all of your time with atheists).
Thanks, maybe I should’ve just said “preaching to the choir” though there’s a more specific and gnarly cognitive move I’m getting at than just directional motivated bias from agreeing with the speaker’s conclusion.
Given your statements like this and this, can someone please ask 80K Hours to stop recommending joining labs to do safety work. I’ve been unhappy about this for years but I’m a nobody.
Echoing other commenters, could you write a list of a few / a bunch of examples of this, or other concrete observations you’ve made that lead you to think this? (I’m very predisposed to believe it, but concrete descriptions & testimony would be helpful. E.g. even if someone, such as a newcomer to AI safety / etc., already believes that what you say is true, they may not even be able to imagine what that would be like or notice when it’s actually happening to them or people around them.)
Wish I could say I found this surprising but that tracks, especially in light of all the recent news. What kinds of distortions would you say are most prevalent?
i wish more people were aware that safety research epistemics at labs are extremely distorted, and that this is a downside worth considering when deciding where to do safety research or how one ought to comport oneself if choosing to work on safety at a lab
Can you say more about how safety research epistemics are distorted?
examples
incentive to do capabilities, or things that are useful for capabilities, or ship in the product
incentive not to think too hard about asymptotics / ways approaches might fail if they are already working
making the company go faster and win is good (anthropic is especially bad at this one)
much stronger bias towards trusting your eyes vs abstract arguments (tbh, LW could use a little bit more of this one on the current margin, but too much is also bad)
a kind of subtle thing where it’s easy for all stories of impact to route through making tiny marginal changes to big things, rather than doing ambitious things
overestimation of the value of knowledge inside labs / the irrelevance of work outside labs
one way to observe some of this is to see what ways median lab people mispredict the world. for example, median lab people often underpredict capabilities progress and scary alignment failures. i think this is not a coincidene
To be fair, the capability mispredictions are shared by lots of other people, and suffice it to say that not even lab employees have abandoned normie priors on capabilities progress yet.
https://wonkmonksnotes.wordpress.com/2021/04/22/daniel-ellsberg-the-effect-of-top-secret-clearance/
[...]
these quotes are so bang on accurate to my experience.
unfortunately, the vast majority of lab employees have been at a lab for fewer than 2 or 3 years.
I agree[1], but I’m worried that this is an applause light for the LW crowd, and it’s also hard from the outside to assess the degree of epistemic distortion. People at labs can (almost) equally argue that the LW++ crowd are epistemically distorted in other ways, and then it just becomes a battle of dueling priors.
I wish there’s a better way to get ground truth here.
FWIW my best guess is that epistemic distortion in the labs are indeed much higher than outside of them, and reasonable people without prior context or commitments but the skills and a lot of time to carefully investigate would end up agreeing with me. But this is really hard to do in practice, and/or communicate well.
I think my shortform from late June nailed a small fraction of these dynamics: “An interesting subimplication of this assessment is that perhaps the reason lab employees often believe their publicly released models are very aligned – with phrases like “our most aligned model to date” – (as opposed to regularly mundanely misaligned) might come from the private models being much worse on this front. The soft bigotry of low expectations, as they say.” My story also tried to get at some of these factors.
I disagree that it’s applause lights coming from Leo (i.e. it’s a meaningful primary source). Buck’s summary here matches my impression.
i think LW also has its own epistemic distortions, in kind of the opposite direction. i think it’s really hard to find a low distortion environment. i think the best thing to do given this is to spend time in both and try to correct for both.
Fwiw I’m also skeptical of this line of argument/reasoning on meta grounds[1]. I’ve most often seen Anthropic ppl like Amanda and Dario deploy this. But I think it’s seductive for bad reasons if ppl think about it even a little bit.
tbc maybe your object-level reasons are sufficiently strong to believe this. Just wanted. to clarify the meta point here because it’s been bugging me for a while.
tbc, i think LW is much closer to sane! but i definitely do get frustrated by some of the cargo culting on here sometimes, and that was the main reason i said what i did. i mostly ignore ant/oai people shitting on LW.
Sorry my comment was unclear! I was less thinking of the specific shitting on safety people sociologically and more about something akin to the golden mean fallacy that I see prominent Ant people deploy in a bunch of seemingly unrelated incidences (dunno if you see it regularly at oai too).
Besides the normative aspect I also think they’re actually making a fairly simple descriptive mistake, which is that it’s very easy to see yourself as the center if you don’t define your bounds ahead of time, and/or if you observe this among ppl you interact with rather than from e.g. polling.
i’m not claiming you should in general be in the center of various things. for example, i don’t think you should spend some time with the far left and some time with the far right (you should spend time with neither); nor do i think you should spend some time with the religious and some time with atheists (i think you should spend almost all of your time with atheists).
Sorry I think I didn’t get my point across (skill issue on my end). I should probably write a full post or a longer quick take about this.
I wonder if that’s upstream of why Claude likes the “secret third thing” so much.
Man who likes standup comedy and Georgism: “the best way to think clearly is in the intersection of late-night comedy and land-value tax essays”
I don’t think “applause light” is being used correctly here. Leo’s claim is not devoid of content, it’s just a claim that most people agree with.
Thanks, maybe I should’ve just said “preaching to the choir” though there’s a more specific and gnarly cognitive move I’m getting at than just directional motivated bias from agreeing with the speaker’s conclusion.
agree, though it does have words people emotionally react to:
█ ████ ████ ██████ ████ █████ ████ safety research epistemics ██ labs ███ █████████ distorted, ███ ████ ████ ██ █ downside █████ ███████████ ████ ████████ █████ ███ ███ safety research ██ ███ ███ █████ ████ ██ epistemics ██ ████████ ██ ████ ██ safety ██ █ lab
Given your statements like this and this, can someone please ask 80K Hours to stop recommending joining labs to do safety work. I’ve been unhappy about this for years but I’m a nobody.
Echoing other commenters, could you write a list of a few / a bunch of examples of this, or other concrete observations you’ve made that lead you to think this? (I’m very predisposed to believe it, but concrete descriptions & testimony would be helpful. E.g. even if someone, such as a newcomer to AI safety / etc., already believes that what you say is true, they may not even be able to imagine what that would be like or notice when it’s actually happening to them or people around them.)
Alas this unawareness is exactly what you would expect if safety research epistemics at labs were distorted in this way.
Are… are they? somehow not??? Maybe other people just don’t have Sinclair’s Law firmly in their worldview’s generating set?
Wish I could say I found this surprising but that tracks, especially in light of all the recent news. What kinds of distortions would you say are most prevalent?