LessWrong team member / moderator. I’ve been a LessWrong organizer since 2011, with roughly equal focus on the cultural, practical and intellectual aspects of the community. My first project was creating the Secular Solstice and helping groups across the world run their own version of it. More recently I’ve been interested in improving my own epistemic standards and helping others to do so as well. as
Raemon
I like this frame. (It’s now my go-to frame for what nearterm political goal to be thinking about)
I’m not actually that sure what in particular needs to be figured out in advance? You gave some examples at the end, I can kinda imagine it, but, I was guessing it looks like “well, we use the existing spy apparatus (that civilians don’t have that much access to so it’s hard to think about), and have the list of people who need to talk to each other ready”, and I can imagine more stuff but I struggle to come up with more than a couple weeks worth of work.
I’m assuming there’s a lot more fiddly details that I’m not tracking, I’d be interested in a followup post that goes into “what actually needs to happen before The Scramble?”.
I specifically wanted to. know what cousin_it thought in this case. I can generate examples I just didn’t know his particular models here.
That could be. But, I think the actual evidence is “republicans and democrats have been surprisingly comparable in terms of how much they’re updating, with a couple outliers”. (the graph of “senators speaking out about AGI was within the same rough ballpark for blue and red politicians)
What are some examples of situations where you’d expect someone to be tempted to apply this “ask what violates it” principle, where they’d be wrong?
(I guessed that your comment was sort of disagreeing with some of the vibe of the post, but I wasn’t entirely sure)
Yeah I am worried about this but don’t feel like I know enough to call the sign.
My uninformed take:
I think it probably would have been better to propose this in a few months when the conservatives have had more time to think “okay shit this AI stuff seems real, we need to do something” and come with some conservative flavored version of the thing, and then this might still cause partisanship but the right wing would have consolidated around some position that a) handled the problem at all, and b) is conservative flavored enough that the debate can be between two At Least Somewhat Helpful plans.
One of my thoughts as I’ve experimented with this:
Sometimes it is just correct to be in full-on-manager-mode.
But, there’s still a fair amount of work that requires deep attention and thought. I think a mode it is often good to be in is to have one major deep thread you are focused on, and some side-threads that are low stakes and the right shape such that you can throw Claude at them and come back in 30 min.
(There is a trap where you try to do this, but, end up spending all day doing AI-tasks that were easy to manage, but not actually that important)
A pair of arguments you’ve maybe heard before but want to make sure you’re tracking:
1) AI companies are naturally motivated to solve legible problems. They aren’t naturally motivated to solve illegible problems. (see Wei Dai’s Legible vs. Illegible AI Safety Problems)
2) It doesn’t actually help to make the legible progress sooner. You get the same wall-clock time to leverage the benefits of the legible progress.
i.e. sooner or later, companies run into situations like HuggingFace, and CEOs and politicians start to notice and researchers have concrete examples of misalignment to study. In the world where safety-conscious people didn’t help companies move faster, we still end up with comparable time afterwards to leverage the legibility.
The question is just “beforehand, did you you have more calendar years of people thinking through the problems that were harder-to-think-about, less commercially useful-to-solve? Or not?”
It seems fine to argue “it’s very hard to make progress on illegible problems, and most such work will be useless.” But I don’t think it’s usefulness is zero. At the very least, having mapped out more deadends is useful when you get to the HuggingFace point. And there’s at least a chance some of the research agendas would pan out.
(what counts as “illegible” varies a bit depending on person. Things are more legible if they are easier to understand. I can dig into examples if that feels helpful)
...
There are counterarguments that feel (potentially) compelling to me here, which involve overhang, or company culture, or what kind of governance situation we’re in, or some other kinds of path dependancy.
But, if we don’t have a specific, valid argument of that type, my baseline expectation is that commercially-useful safety work is probably worse than useless.
Yeah I’d been frustrated by lack-of-this recently and finally got around to making it!
I mean I think there is “a Lightcone/MIRI-ish cluster, where being a lab employee is obviously really bad by default and you need some really good points to make up for it”, and then, idk, the broad professionalized EAcosystem where lab employees seem to be the experts and have lots of money, why wouldn’t they be high status?
I mean, idk what they did at OpenAI, just that it’s less obvious they did anything to accelerate capabilities while they were there.
[musing/rambling, not sure about point
The thing I feel confused about, despite this being my obvious first-order belief, is… nonetheless, I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I’m less sure).
Notably, they were both doing governance, I think, not capabilities.
When I imagine the average MATS scholar asking “should I go work on governance at OpenAI?”, I think “oh god definitely no”, because I have a low opinion of average MATS scholar’s ability to track incentive pressures on themselves and warp themselves and otherwise have their eye on the ball in the first place.
I’m not sure whether Daniel and Richard did a cognitive operation such that they could know in advance they’d leave (and I give Richard less credit for leaving because I think he left after the ship had clearly sailed on OpenAI having anything like a real safety culture, whereas Daniel seemed more helpful in catalyzing that wave).
...okay typing this out, while I’m still unsure about many details, I immediately notice “I don’t think the average MATS scholar even actually knows the core x-risk arguments well these days”, which is minimum pre-requisite for it being remotely plausible that one should work at a lab on anything.
For example, I asked earlier “Why can’t one of the other experienced mods train or manage the new mod?” and got no answer to this (confirmed via LLM that it wasn’t answered or explained anywhere in the entire comment tree rooted at the shortform OP)
fyi I have stopped answering questions like this after you ignored most of the content of this comment, which stated pretty clearly what sorts of discussion would be likely to move us, and why: https://www.lesswrong.com/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform?commentId=betXXw7gFskNfiHp3
Ah, well cool that you tried that and sorry it didn’t work.
But to clarify, I would be using Fable for this, I don’t have a belief that older models are able to track the arguments here.
I don’t have that strong a belief that this “ask the models” thing will turn out to work, maybe that’s a dead-end for now. But, insofar as the idea has promise I’d be doing this with Fable, making sure it’s read Meta-tations on Moderation: Towards Public Archipelago , Banning Said Achmiz (and broader thoughts on moderation) and all comments on those, and reading your shortform page for all comments between you and habryka on moderation.
One reason I think this problem is on your side is that Oliver has said multiple times “we are not doing this because Eliezer said so”, and you repeated the belief that we were in your most recent interaction.
There are less straightforward reasons for that but if you’re still ignoring basic statements like that the conversation feels pretty hopeless to me.
I agree it might not be totally comprehensive, but, that seems kinda fine? A lot of the concepts have gotten repeated. (Like, seems fine to try a lower effort version and see how it goes and try a higher effort version if it feels on-track-to-be-helpful but insufficient)
I’ve done this actually for twitter conversations with people I don’t know and it was just pretty straightforward. “Hey I want to talk to so-and-so about X, please read everything by so-and-so that bears on X”.
One important thing was to fork the conversation after the initial “read up everything relevant”, so that when I asked “okay, what if I wrote this to them?”, it doesn’t anchor on whatever random stuff it guessed the first couple times.
Oh that is pretty good actually.
(i.e. basically, @Wei Dai you can have Claude Code Fable scrape for all the past moderation discussions, and then write up a message and ask “are there replies that habryka pretty obviously say to this?” and then iterate a bit until you get to something that feels like it’s at least moving the conversation forward)
Part of my answer to this is “we have explained our models in detail several times, and I haven’t seen you reference any of those details in your suggestions. You can’t propose a workable idea until you have actually integrated everything we’ve said.”
Our most valuable resource is habryka’s time. We’re happy to talk things through, but by now I think habryka has explained all the pieces of the problem at least twice.
I think things will go better here if you start by rereading everything (at least from this thread, and maybe any specific things he linked to from the past couple rounds of discussion), starting with the assumption “we are not doing this because Eliezer made us, we are doing this because we think it is a good idea”, and try to summarize everything as you understand it so far.
I’m saying you need to convince habryka that whatever your proposal is would increase net intellectual progress on LessWrong. It seems like you keep avoiding engaging with any of our cruxes, and you will eventually need to do that if you want us to change anything.
From what I can tell, your current model is based on several wrong assumptions, and it sounds like you are about to rabbithole on a plan that rests on those false assumptions.
Habryka has spent a lot of time explaining this to you and you keep AFAICT not listening. (Part of where the cost-in-time comes from, as well as training. Any new hire requires habryka onboarding time)
We would not onboard third-party moderators that moved the site in a different moderation direction until you convinced habryka your changes would result in a better LessWrong by his lights.
Can you give me a dollar figure to solve this problem?
It’s not really measured in dollars. It’s measured in opportunity-cost-of-habryka’s time. But also it just really won’t solve the problem the way you imagine.
Back in my first comment I mentioned “it seems like you’re conflating some things”, and, I wanna revisit:
It sounds like you have a general model of “there is a ‘the LW mods have some moderation-taste that consistently points in the wrong direction, and I wish that were different’.” There is some truth to that. But, for example, it feels like you’re assuming:
The moderators banning Said
The moderators adding “authors can ban users.”
...share a common cause. They sort of do. But, note the banning Said was third party moderation. You disagreed with the result, but, that’s what it was. You’re not just asking for third party moderation, you’re implicitly asking for third party moderation you agree with. Which is a fair thing to want, but I think you’re mixing up a bunch of things here without noticing.
And, “authors can ban users” was the reason we didn’t ban Said for so long. In the world where we had no author-banning, we would have banned Said much earlier, and I’d be even more confident it was the right call in that case.
...
Now, those two decision share the cause “the moderators believe authors are more important than commenter, because authors drive the majority of intellectual progress on LessWrong.” I think this is especially true on LessWrong, but basically any platform will share the problem of “content creation is harder than critique”, and I think it’d basically be the wrong call for any platform to ignore that.
But, if you think that’s wrong, and you don’t want to found a different platform, you do actually need to convince habryka his upstream models of intellectual progress are wrong. It hasn’t seemed like you’re trying to do that at all, just asserting principles that seem good to you.
I think “stop doing harmful work until you are fired, instead of quitting” is fairly compelling.
But, I do think most people working on prosaic safety are mostly doing harm.
If you don’t have a particular story for how what you’re doing scales to superalignment*, I a) don’t think you’re solving particularly important problems in the nearterm, b) in the medium term, mostly making it marginally faster to roll out stronger capabilities, which is bad because it’s burning calendar time for serial research time, and solving the legible problems leaving the illegible ones.