I endorse this. I do think it’s hard, and I think it’s probably even scary. But I think “thinking about it for five full minutes” suggests that there are a lot of options besides saying “evaluation is hard”. I hope that my suggestions prompt some reflection.
dan.parshall
Tangentially related, I find that my interactions with Opus 4.7 are far more productive when I have installed the “claude-exit” tool. https://github.com/danparshall/claude-exit
This is encouraging, and I hope true!
I tried to tag you in the comment above: https://www.lesswrong.com/posts/2o8B9hDN94k4Qine6/grantmakers-aren-t-afraid-to-die?commentId=fPr53stHP5f64isGW
Briefly, I think some of what looks like capture is response to the perception and/or reality of the “mops and sociopaths” dynamics (which is sorta similar to your post about ASO and messaging). I think AIS folks are going to have to find some way to work through those concerns, because if we want to win, we’ll need a larger coalition.
I’m still thinking about this, but it would be helpful if you could elaborate about the networks and signals.
I think there’s an understandable fear that AI Safety will become subject to the same dynamics described in https://meaningness.com/geeks-mops-sociopaths
And I suspect that some of that fear might drive some of the behavior that @less_raichu interprets as as “capture”.
And I don’t think those fears are wrong! Nor are the responses to them “a little bit evil”. But I’m also currently of the opinion that in order to win the game, we’ll have to suck it up, handle the mops and sociopaths the best we can, and plod on anyway. I’d intended to cover some of this material in the “Lessons From Vannevar” post, but wanted to get it out on the anniversary of the Einstein Letter, so trimmed it. I should probably tackle this subject in detail.
I think this is an excellent point, and true as far as it goes! But let’s dig a bit to think about what might cause this to be the case?
As near as I can figure, there are two possibilities:
you have already found ALL of the talent that could be net beneficial
you don’t know HOW TO TELL net-positive from net-negative talent
If it’s the first case, then I concede my prior complaints about “EA/AIS isn’t actually talent constrained”. But also, we should give them all of the money now! (unless you don’t believe that timelines are short, but that gets back to my original thesis of “grantmakers don’t believe in imminent existential risk”)
If it’s the second case, then we should invest HEAVILY into finding, identifying, and supporting net-positive talent, possibly through one of the mechanisms I outlined above. I think this is true even in long-timeline worlds, and part of the change I hope this essay spurs.
Or is there some third case I’m missing?
Thanks. I think this is a good example of when “privilege” really applies. I’m fortunate to have been able to spend many months on this self-funded, but at some point it’s a stretch. An even bigger stretch was sorting out childcare for the two months that I’ve been accepted to a residency!
One point that I haven’t made explicitly, but should: yes, the existing team is top-level players, but they are a small team. Bringing in 10x the number of people, even if on average each is only half as good, is still a huge win! Bringing in 100x the people is a win even if each is only a tenth as good!
I agree that’s the least-charitable interpretation, and I think it behooves anyone in their position to be aware of the possible incentive towards prestige and control. I think they could have been doing a lot more to reduce the talent constraint in the past, by giving people outside that small circle opportunities to get involved, but on long timescales I understand why. I genuinely think that they’re trying to do the best they can, but haven’t yet internalized the implications for short timescales, hence the post.
I think I broadly agree with this. e.g. I agree that the amount of “force” needed gets greater as the timelines get shorter. A rocket can be steered with only a tiny amount of delta-vee very early in the journey, but the amount required gets greater closer to the ending.
I also think the grantmakers believe in the possibility but not imminence of extinction risk, and I wanted to call that out specifically. I think having that attention focused on on “use a different strategy now that the game is different” may help.
Independently, I think it’s valuable to the incoming recruits to increase transparency (which at the moment they definitely do not have) and I’m in the fortunate position of being able to go back to industry if everyone hates me for writing this.
First, you can continue to collect more data over time. Second, if you are currently losing then you should lean into higher-variance strategies. Provided that you can do basic filtering to remove the worst cases, you can get positive expectation results that could potentially change the board. Third, if you are currently losing then you should absolutely invest in “more data” because what you’re currently doing is not working.
Hmm, I think I disagree here, and I certainly disagree on short timelines, partly because keeping things fixed to “the cool kids” guarantees that the total size of the labor force available for AI work doesn’t grow. Not a problem on long timelines, but if you believe that the next few years matter, then this is really really bad!
We already know from past experiments on scientific proposals that it’s pretty easy to filter out the true crap, but extremely hard to rank-order after that. So unless something is truly different (maybe?) then a much better proposal is “filter for minimal quality, and fund”. That means small grants to people to move into AI safety (e.g. the Blue Dot “Career Transition Grant”) is kinda a no-brainer. Even if half the people are grifters, I think it’s pretty easy to tell at the end of a month or two which half with a reasonable FN/FP rate.
Note that Habryka’s metric is “ten blog posts is better than a referral”… so why not be VERY generous about giving $30k grants for three months, and saying, “at the end of this, we’d like to see at least 10 blog posts”. Anyone who doesn’t produce that you cut, anyone who DOES produce that is shortlisted for followup. You don’t have to apply the same stringency of filter at all stages! You can apply permissive filters at early stages and then get progressively more strict, as you learn more!
Or, of course, grantmakers could keep doing what they’ve been doing… which is arguably better over any individual grant cycle, but as I am trying to point out: still a losing outcome overall
(ETA: as I mention below, I should make this explicit: our current team is top-notch, which is great! But if we can 10x the size of the team, even if the average member is only half as good, that’s still a win. Likewise if we 100x the size of the team, even if the average member is only a tenth as good, that’s an even bigger win!)
hmmm, I think that the Clinton/Dole race of 1996 is instructive here, because Clinton set himself the task of moving so far to the right, that Dole had little platform left to run on (at least as FdB tells it). We might end up in the scenario where AI safety becomes so popular that politicians are trying to out-do each other on who can be “tough on AI”.
But generally I think it’s best to start from the OPs perspective, of avoiding partinization at all.
I think this is partly fair, but I also think it’s context-dependent. Ten years ago, hell even 2 years ago, I’d have broadly favored conservative decision-making as well. But if you think you’re losing, and that you can’t play much longer, then conservative decision-making is no longer the rational strategy, and “keep doing what we’ve been doing” is an especially poor choice if you think we’re losing!
So what’s the crux? Do you think we’re winning with our current approach, or do you think timelines are long? Or some other factor?
I think we’re broadly in agreement. My point, which I apologize for not spelling out explicitly, was something like “even if only 1 in 500 applicants were qualified, they’d have been able to hire multiple people for the role, thus planning ahead instead of just-barely keeping up”
But the actual revealed process/preference is “rely on references from a small inner circle” https://x.com/BogdanIonutCir2/status/2091784358562590941
The “talent-bound narrative” is mostly bullshit? https://forum.effectivealtruism.org/posts/B6d8Wzk4gNzHsXvdi/ai-safety-is-extremely-bottlenecked-on-grantmakers#Jf8RWS8YcsPT8XqvM
If they’re getting 1500 applications for a single role, then they ought to be hiring at least 3 of those people… or admit that they’re not talent-constrained.
Late to the party, but as someone who regularly engages with normies, I think this approach is entirely understandable. I’m old enough to remember when “rationality was applied winning”, and folks knew that one only has so many “weirdness points”, and they have to be carefully used.
I do think that there’s a difference between “carefully phrasing one’s objections and choosing one’s battles” as opposed to “attack near-allies for strategic gain”, but I’m glad to see you explicitly acknowledged and apologized for that; please don’t do it again.
Otherwise, overall evaluation: B+, far from perfect play, but not bad for a flawed human.
thanks for testing! I notice that the re-un-translated version is almost exactly what I wrote originally, which on the one hand is great. OTOH, I didn’t find LLM writing as annoying as everyone else seemed to- but apparently my writing has always been like that. I now have to go explicitly remove dashes just so no one thinks it’s slop. :: le sigh ::
I like the idea of having “AI usage note: originally written in Spanish, translated by Opus 4.7 to English, see conversation at [LINK]”, that’s teh kind of thing that could be automated as part of a pipeline system such as SlopChecker.
I do genuinely like the democratization aspect, but some communities I’m in, folks haven’t done the extrapolation to “what if everyone does this?”, and this post is partly well-intentioned advice for them.
I’ve read that Claudish persists across languages in normal discussion (e.g. on Xitter someone claimed the tics are all the same nature in German). I’m not sure if it would be noticeable to a human for translation only.
And while I’m fluent in Spanish (CEFR ~B2 although never formally tested), my grasp of nuance isn’t good enough to pick up tics along those lines, so I’ll have to defer to those who are at that C1/C2 level. By all means elaborate on what you’ve observed, though!
However my understanding is that per the EU reg, the translation would still be watermarked, and thus detectable in many of the scenarios I think matter (e.g. grant funding).
I endorse this, for a few different reasons:
incoming talent desperately needs feedback; I plan to write about this soon.
grantmakers would benefit from being explicit about what they think success probabilities are, even if only to themselves
if we have “thousands of explicit probabilities” and aren’t USING that data, it’s damn near criminal; see my proposal for Thompson sampling in https://www.lesswrong.com/posts/2o8B9hDN94k4Qine6/grantmakers-aren-t-afraid-to-die
If we’re trying to actually win, then ignoring things like this seems like an own-goal.