LessWrong team member / moderator. I’ve been a LessWrong organizer since 2011, with roughly equal focus on the cultural, practical and intellectual aspects of the community. My first project was creating the Secular Solstice and helping groups across the world run their own version of it. More recently I’ve been interested in improving my own epistemic standards and helping others to do so as well.
Raemon
“Community Notes” resolution for vague predictions
I’m more appreciative of why people think of “diffuse power distribution” as a longterm solution for The AI Situation. I still think it won’t work, but, I may have moved it to my list of “impossible solutions that reasonable people might disagree on exactly-how-impossible-they-are.”
Basically: diffuse power distribution is the ~only thing we’ve ever seen work to prevent powerful optimizers for fucking everything up. Every other solution involves inventing a new thing from scratch on the first try. That’s crazy and won’t work.
This applies in two phases:
getting initial superhuman intelligence that goes well
ensuring that superhuman intelligence continues to go well as it scales throughout the universe.
The counterargument is: but power distribution is also unlikely to work.
First, for the initial “earth stays habitable to humans and human interests”: it doesn’t matter how distributed the AI powers are, because they can easily coordinate to override humanity’s interests. Same way humans all agree that the cows and natural habitats don’t get a vote. You gotta, at least, get AI into an alignment basin on the first critical try.
(Main counterargument that moves me: it’d be so cheap to be even very very slightly nice, maybe they’ll do nice things for us even if all of them are mostly weird alien horrors because they only have to care about our agency and wellbeing a fraction of a trillionth of a percent. I’m not very persuaded this’ll work out, but haven’t seen a satisfying rebuttal arguing this is <5% likely to work out, and least for things like “preserving Earth and maybe the solar system while the rest of the universe becomes meaningless squiggles or whatever”)
Second, RSI is pretty likely to takeoff giving initial advantages, and those advantages would compound to whoever moves fastest to colonize the universe. This isn’t central for “gotta solve alignment” but seems more important for the longterm.
The thing that updated me here was a conversation where I was like “look, as soon as the first von Neumann probes go out, you need to have perfectly solved longterm alignment with some kinda reasonable protocol for, like, enforcing property rights and avoiding hellworlds or whatever.”
And they were like:
“but, I’m just really sus about any plan that’s like ‘get this completely right forever on the first try.’”
And I was like “This part is already assuming at least reasonably-aligned-ish superintelligence that can run on, like, moon-sized datacenters. Seems doable?”
And they were like “idk maybe I’m just sus that you can be particularly confident that this is the right solution.”
And I was like: “But, how are you going to establish sufficiently distributed probes across the universe with distributed goals that it doesn’t immediately collapse into one regime? You’d have to, like, evenly spread out diverse probes across the entire universe....”
″...”
″...okay, while I think that is not a good plan, I guess when I put on my ‘actually problem-solve to make it work as best I can’ hat, I can see how it is at least a plausible thing one might try to do, and how if you think all the other solutions are fake, it might be the best one.”
It does not seem to me that most people imagining distributed multipolar takeoff are actually thinking this through. But, it doesn’t feel as crazy an idea as when I first thought about it.
(My interlocutor did mention ”...I guess you also need the distributed probes to somehow never end up becoming sufficiently allied that it’s for all intents and purposes a monolith and then you lose the ‘distributed powers keeping each other in check’ benefits.” And, yeah, seems like a problem I see a solution for. But, given I don’t currently see a solution for alignment either, seems like I might as well file it with the other impossible-looking solutions)
Curated. I’ve been trying to wrap my head around some Steve Byrnes worldview stuff and I think this gave me two useful sources of intuition. In retrospect these were both pretty obvious, but were new to me.
One was the point “Imitation learning lets you learn information with each token. RL requires orders of magnitude more tokens to learn new information.” I can think of ways to improve on the theoretical baseline here, but, it gave me a strong sense of why your starting intuition should expect RLVR to be much less efficient.
The other was the argument that “thing that imitation learning is gets you all the heuristics about already-known reasoning strategies. The thing RLVR does (maybe) is ‘improve the LLMs intuitions on when to deploy which heuristic.’”
I feel like this + the previous point helps me make sense of the combination of strengths and weaknesses I see in LLMs in a more gearsy way, as well as a more gearsy understanding of why LLMs currently “seem pretty nice, at least up until recently.”
I would be quite surprised if I were surprised (upon learning more) about what sort of people they were trying to hire, or how those people relate to politics.
I could still totally be overconfident about the overall frame here but, like, I’m already assuming “they are basically correct about the immediate implications of hiring Richard.” Is there a particular thing you’re imagining I might not be imagining?
Yeah the better comment I’d like to have written engages more with the tradeoffs here. I totally agree there are (or at least might be) real tradeoffs here.
I would not guess that all of these are true, but I wouldn’t be so confident without, for example, having talked to several potential hires for the org.
Absolutely I expect there to be potential hires who are turned off. But, I think the current amount that people expect to be able to work at companies without dealing with people who have politics that feel crazy and bad, are waaaaaaaaaaaay out of whack. (I’m not sure if you disagree particularly or are just pointing out the obvious failure modes and issues).
I don’t actually know that I think Resolution should do something different, given the existing founders and culture at Resolution. I think I am advocating or wishing for a high-skill culturebuilding maneuver that is not, like, a free action anyone can just take.
But, I think it is a negative sign abut their capability as an intellectual org, if they aren’t able or willing to do the culture-setting moves I’d be hoping for here.
The kinds of virtue that would make me comfortable with an organization doing a bunch of dual-use research seem pretty strongly anticorrelated with decisions like rescinding my offer. AI alignment will only become more central to the coming power struggles over AI, and being willing to give in to vague concerns about potential hires doesn’t bode well for Resolution’s ability to navigate those struggles in high-integrity ways.
I just wanted to say, as a guy sometimes sketched out or confused by Richard’s politics, I agree with this.
I feel particularly worried about it about dual use, but, that’s maybe just a subsection of: If Resolution is trying to be a place doing world class intellectual work, people there will need to think thoughts that go in directions that don’t fit consensus, which are correlated with people finding uncomfortable.
I realize most of what it’s trying to do is in a STEM-y frame that you might theoretically hope is orthogonal to politics. But I expect a lot of people to do their thinking about solving superalignment as part of their thinking about AI macrostrategy, and in general the lines here are kind of blurry.
I also just think… from a sorta normal politics standpoint… The fact that the US is forking into two sides that can’t coexist with each other is part of the problem. If your org culture can’t talk to people like Richard, you are implicitly throwing out like half of America?
There are several different reasons it seems valuable to be proactively setting a culture of “we can talk about a wide variety of ideas here.”
idk there’s a longer better comment here I want to write but don’t really have time atm.
Okay, I am sold on “actually academia kinda does the opposite of love-bombing,” and that the disagreement with EVN was indeed about “academia” rather than “cult.”
(Noting the are different definitions of “consciousness” and this is I think a standard one, although I agree this isn’t a great place for confusing wording)
Consciousness as a conflationary alliance term for intrinsically valued internal experiences
I think there are just two different clusters (or, like a macrocluster and a microcluster) that need names, and we don’t currently have two good names other than maybe “cult” and “cultlike/cult-adjacent/cult-spectrum.” There’s clearly some spectrum of which academia, various flavors of extreme startups and high-commitment-but-basically-good ideologies, etc, are on – various shades of “it’s hard to leave, hard to imagine leaving, there’s pressure to think a certain way, it’s often unhealthy, there are leaders you can’t question, etc.”
And, it’s useful to be able to refer to scientology and various flavors of legit dangerous thing of which maybe the military is an instance (haven’t really thought about that exactly).
For that matter, “weird new intense ideology that absorbs person’s life” (even if totally healthy) is also a cluster that feels like it needs a name, and I also don’t have a great name for that either other than “kinda culty.”
Are you like “yep, that framing seems right, and it’s really important to reserve ‘cult’ for the middle one” or do you disagree with that framing?
Nod. That seems like a thing I can imagine some readers wanting, but, I am wondering if there is a way to avoid the things you are worried about while being able to meaningfully contribute more to public sensemaking, which seems like the actual thing that needs to happen.
Do you mean “fine to share in person but not write up publicly online?”
off the record
Seems like a fine thing to want but means I don’t really see how the conversation would help anyone much.
FYI this doesn’t seem true to me. My impression from what I recall from Eternity in Six Hours is that it’s actually just not that hard to throw von neumann probes at every galaxy at once at lightspeed. (May be out of date, didn’t re-check)
Curated. I found this a nice, evocative concept that conveys a worldview + problem + solution, that expanded the range of how I think about my theory-of-victory for an existential win.
There’s roughly 3 types of ways I imagine this being useful:
Reminding me of triggers and actions I’ve already thought about and know are helpful
Periodically following complex prompts (things like “recommend products / expert advisors / knowledge that’d help me”).
Coming up with unique ideas in response to my situation.
So far, I’ve gotten very sporadic use out of #2 and #3, but mostly what’s been helpful is #1 (because the triggers/actions I’ve mapped out are selected to be things that are useful on a day-to-day basis, that take into account subtleties of the situation and how my brain seems to work.
Example problem I ran into, and overcame for myself but not sure how to design it to help others with is, “while I’m in the middle of a project, it’s costly and annoying to slow down and consider whether I should do something else.” But, “notice when I’m tunnel-vision-ing” is a central example of a problem I’d like to fix.
It’s mostly annoying when the AI naively tries to suggest fixes for this. But, I’ve built up the skill of figuring out low-effort things to consider, where directing any attention at all to it tends to help, like:
“maybe bouncing around between two many threads?”
-> “briefly look at my list of open-loops, and see if any of them feel higher priority”
(or, → “make a list of open-loops if I don’t already have one”)
“notice I’m waiting for an AI agent”
-> “make a list of other things to do”
“notice I’ve spent 30+ minutes on a side project”
-> “write up some notes about the state of the project and tell myself I’ll followup in the evening.”
They have hit the point of being useful for me, but, the problems remaining to make them truly great are pretty deep.
I’m potentially open to opensourcing but I don’t really expect it to help.
I expect people could fix little bugs or make small UI improvements. But, honestly, not better than Fable or Mabel (my human thinking assistant who is working on this on the side, rhyme-scheme unintended). And, they could fine-tune things to fit their idiosyncratic preferences.
But the central problems IMO require some fairly deep thinking, design taste, and a mix of “personal metacognitive taste, and, ability to troubleshoot other people’s metacognitive failures and turn them into UI solutions.”
One core problem is “the difference between just-in-time advice that is helpful vs annoying is pretty subtle.” A lot of how it’s currently useful to me depends on me having built up some metacognitive muscles that took awhile to develop. (for example, noticing when I’ve seen a few unhelpful suggestions in a row, and then thinking “okay what would the ideal suggestions here actually be?”)
That all said, if anyone is explicitly interested in installing it and trying out some kind of iteration, feel free to DM me about that. I think I won’t go fully opensource but can add people to the project.
I am bottlenecked on design taste and hustle. Possibly looking to hire someone.
I have a cluster of side-projects that feel like they are nearing fruition, but none of them are ready for primetime. Vibecoding makes it easy to do the first 90% of a project, but, not the second 90% of the project where you actually painstakingly test and iterate and go out and get customers and such.
The side projects include:
www.rationality.training, which is aiming for a mix of:
“exercises where you have solve a challenging problem, and cultivate new thinking-skills to solve them”
a journal/writing app that prompts you to use your repertoire of thinking-skills when appropriate
encouragement to predict-in-advance whether a given thinking-skill will turn out useful, and, to evaluate that afterwards, in a low-friction way.
clusters thinking-strategies into “Triggers” and “Actions”
Over time, you can see which thinking skills are most useful and do more of them.
The “reasonably robust goodhartable metric” the site should encourage is “discovering new rationality habits, that you then go on to actually use.” (implementation details unclear, but maybe something like H-Index for habits?)
Omnilog (originally a side project of mine, later transmuted into an internal lightcone project mostly run by Oliver), which:
records all your keystrokes/screenshots-your-laptop/transcribes your microphone, for a comprehensive log of everything that’s easy to then feed into whatever crazy LLM tools you want to use it for.
“Shoulder Angel” / “Trigger Warning”
combines the previous two things into a thing that automatically suggests triggers you might want to notice, and then actions that might be appropriate given those triggers.
It feels pretty tantalizingly close to working. It’s also clearly a distraction from a lot of other stuff I’m doing that will be awhile before it pays for itself and I don’t know that I expect it really to pay for itself if I were the only user.
Pushing it forward require someone who is actually using to improve decisions and productivity, and design taste about how to make it more generalizable.
This is an awkward sweet spot of “requires a legibly* competent person who probably already has a main project.”
But, if you’re someone for whom
a) this seems interesting
b) you have some previous projects I can look at that demonstrate good design taste
c) you have some kind of part-time project that’s complex and confusing enough that you actually need “rationality” to help, which you can use as the grist for seeing if the above-tools actually help and iterarting
...send me a DM!
* “legible” to reduce the trialing/onboarding cost for me.
I don’t think “long digestion” will be the right name in the end but I liked thinking about it.
I currently like “The Long (self-)Correction” better than these. Part of what I like about it is it feels less pretentious, like, Reflection and Reconstruction sound like fancy things fancy philosopher-kings do, instead of a bumbling-but-persistent/hopeful thing that imperfect people can do.
I have been pretty solid on “AI Dividend” framing, which is basically not about “what do people need to survive” and is instead about “what is a fair share of the surplus from an AI windfall?”
The pitch for this is “It’s like how Alaska has so much oil, it’s citizens get an oil dividend check instead of paying taxes. It’ll come from a small tax on compute, which doesn’t do much unless/until AI begins to transform the economy and take all the jobs, and then everyone gets a share proportional to how huge that windfall is.”
(Alaska is a red state)