LessWrong team member / moderator. I’ve been a LessWrong organizer since 2011, with roughly equal focus on the cultural, practical and intellectual aspects of the community. My first project was creating the Secular Solstice and helping groups across the world run their own version of it. More recently I’ve been interested in improving my own epistemic standards and helping others to do so as well.
Raemon
Is that more like an endorsed opinion, or just some background about how your thought process works?
Though in this case, my claim is: for (somewhat illegible, abstract) reasons, if AI happens soon in your lifetime, without time for people to think through the illegible abstract problems, you will die.
And, like, if you want to get the cool good things without the bad things, you need to think through the illegible fake-feeling things in advance.
Yeah, I think it’s a quite strong claim that things would get done “never” that requires extraordinary evidence.
i.e. at what odds would you bet that no one else would try scaling deep learning for decades, even as compute prices continue to fall? What about a couple centuries or millennia? (whether it happens in a few decades or a thousand years is less overwhelmingly obvious. I would bet “decades”, but can imagine learning good counterarguments).
I’ve updated further on this the more I’ve looked into history of intellectual progress – I would have thought “so and so invented this and was causally responsible” but it’s fairly rare to see inventions that didn’t have multiple contemporaries. (Examples of Highly Counterfactual Discoveries? asks for examples, and there are some, but, even then most of the cases seem to be measured in decades)
...
I think there are legitimate arguments about whether it’s better to have AI takeoff
...sooner, before we get more Hardware overhang.
...sooner, while there are fewer actors, who are maybe easier to coordinate with.
...via LLMs, which gave us a paradigm that has a “non-agentic foundation”. We’re currently seeing them turn into agents in all the obvious ways because agency is instrumentally useful. But, it’s at least plausible to imagine building classes of oracle AI that are less agentic that are at least useful.
My answer to these is “probably no, because, the problem is so hard it requires decades of serial research time and it’s too hard for civilization to stop at exactly the safe point, and the legible problems are all easier to work on that the legible problems, and naively solving the legible problems leads to disaster.”
But that gets more into the weeds and it feels like reasonable people could disagree on the details.
I am kinda interested in you mulling this over, and reading Legible vs. Illegible AI Safety Problems, and… idk having some takes or hearing your off-the-cuff thoughts about Wei Dai’s framing.
I also think that we all underrated the value of building LLM-ish things in the first place. It would have been safer to not go in this direction at all, but would it have been better?
My answer is “yes, it’d have been better”, because this sort of thing would have happened sooner or later. Because it’s happening sooner, we have less “serial research time” for deep thinking.
All the interesting economic productive stuff still happens eventually. The part where “now it is concrete, and easier to think about, for people who benefit a lot from seeing concrete things in front of them” would have happened eventually. (Both for researcher-types and politician types). It could have happened in a world where we’d made more substantial strides on agentfoundationsy stuf.
I’m still learning, but some things that seem “real” to me in technical AI safety are like, computational mechanics, singular learning theory, the beginnings of some multi-agent game theory, some “what is an agent” theory with causal graphs, etc. it isn’t at a stage to be implemented into AIs in practice yet.
I think there’s a ”.” missing somewhere here? (This looks like a list of things, the first few of which are supposed to be real, the second few of which are “not real?” I’m not sure I’m parsing it right.)
...
My sort model of… not really psychoanalyzing you-in-particular, but a broader class of people (I don’t have that great a model of you-in-particular here). But, what your comment makes me think of:
There are people who seem to have more difficulty thinking through abstract things that they can’t see them in front of them. It feels to them like they can’t make progress on the problem until there are concrete things to do.
And… well, it’s just… a skill issue? There was useful stuff to do that did not depend on having concrete stuff in front of you. It did depend on being able to think through multiple-steps-ahead and simulate things well enough for them to feel real (while, tracking the uncertainty of the different ways things could have played out).
And not everyone can do that. And, it’s sorta reasonable for the abstract stuff to feel fake to the people who can’t do that. But, man, the people who pushed ahead to “make AI concrete” burned a lot of potential time for thinking things through that would have given us a lot better positioned to take advantage of the situation once we did inevitably reach the “things are now very concrete” period.
(Example of this: a guy on twitter said “well now that we’ve seen how AI misalignment and hacking plays out in practice, we’re much better able to think through what to do about it.” And gwern replied ”...what did you know now, what wasn’t extremely obviously going to happen like a year ago? Of course hacking was going to be a convergent goal, of course it was going to show up sooner or later.”
...
Having ranted that out: I don’t think this is exactly fair. I don’t think “inability to think abstractly” was a weakness of Paul. I think it’s probably a weakness of Dario, not sure about Demis or Sam. I think the Demis, Dario, Sam and Elon’s biggest weaknesses are more like “just really really action biased, and really motivated to be at the center of the action.”
My rant is more directed at people on the sidelines who were vaguely egging on Demis, Sam, Dario and Elon.
I think Paul gets bayes points for having predicted “takeoff will be smooth”, loses bayes points for “takeoff will be long”, and loses “made the world worse” points for doing work that accelerated takeoff by working on the legible problems instead of the illegible ones in a way that burned calendar time.
Yeah my question was a mix of “what’s your next bottleneck” and “what are you/John’s guesses about whether LLMs are fundamentally capable of handling the ‘knowing deciding what to point at’ part, or other stuff that’s more taste-laden.”
Having enough datapoints to see a trajectory, is your current model that they will improve a bit more but plateau at helping with your work in a predictable way, or, does it seem more likely to might end up dramatically accelerating it in some fashion?
A thing I feel a bit confused about reading this is, like, I have an impression that, say, “interpretability researchers” have been using AI in a way that at least seemed superficially productive to them, and while I think you’re doing different stuff than them my vague impression from a year ago was it wasn’t, like, crazy different when it came to the coding.
I see reasons why it’s good to have cybersecurity AI. But, I don’t see reason to particularly try to differentially accelerate it.
By default, I think AI capabilities are basically bad. There are good properties of having more capabilities of some types, but, in order to pass the high bar of “it’s good to accelerate AI capabilities”, they need to not just be useful, but, more useful to have sooner relative to other type of capabilities.
I think giving the world more time to adapt to the current level of AI cyber capabilities seems better than rushing to the next level.
Train existing models to increase their cybersecurity capabilities.
Why is this good?
I think a takeaway from this convo is actually “I want to build a singleplayer writing tool that just naturally prompts you to notice prediction-implications of your writing and help operationalize them”, and if that seems to be going well is more natural to try out in multiplayer contexts.
Cool. I am glad you’re working on that!
I don’t think this is why you don’t have forecasts more widely locked/tracked on lesswrong.
We’ve specifically thought about forecasts several times, and each time we end up feeling “idk, the forecasts just don’t quite seem to be doing the same type of work that LW posts are doing, like, the fact that you need to do all this annoying operationalization to make them trackable is pretty annoying and doesn’t really feel like it’s Doing The Thing.”
This thread has made me feel more optimistic about routing around that (less with community notes probably, more with AI assistance)
I’m not 100% sure what you mean, but, it sounds like this is Zvi?
Personally I think community notes is a better big swing than forecasting.
Well insofar as community notes is the bedrock that enables this thing, yeah it seems basically “strictly better as a big swing.” You’d have to invent it first to get people warmed up to using it in other contexts.
providence tracking on the web,
Could you explain more about what that means?
I think this needs to happen at the start, not the end. There should be agreement while the forecast is in flight.
...
Like my sense is that people don’t like or aren’t capable of back rationalisation on controversial topics to come to conclusions they don’t like. But they are if they agreed beforehand. So personally I expect a process (sure, I agree, a community notes-like process) to happen during the forecast.
Mm, yeah sounds right. I assume here you mean agreement on the parameters/ironing-out-operationalization as much as is practical?
If you think this, do you think this would work on LessWrong, why why not? (I don’t think it would)
I feel something like “this is overkill for LessWrong.” (We’ve particularly tried out community-notes algorithms… and, well, for good or for ill, there just doesn’t actually seem to be major competing clusters in LW that need to be community-notes-algorithm’d – the results aren’t noticeably different from the normal karma system)
I think it’d be good to have various automated nudges on LW towards better epistemics. I think this isn’t “the piece that LW is missing” in a way that it feels more like “a piece that twitter is missing.” The reason to focus on predictions is because it’s a way to inject epistemics into the broader populace that it’s easier to justify.
Insofar as I believe this is the right thing for twitter, I do think I should prioritize making some version of it happen on LW and if it’s not working here that is evidence it wouldn’t work there.
Agreed on AI-suggested operationalizations and resolutions (this could maybe be iterated on more extensively on more classical prediction market sites).
I’d be interested in you reading over my reply to Nathan (which maybe does a better job laying out why I care about this), and then maybe spending a couple minutes on rambling about “does the proposal here feel like it’d work? Or, do you have other ideas for solving the high level goal of “move the needle on ‘mainstream incentives towards truth’?”.
Thanks for substantive reply!
Yep, this is a big swing.
This is backchaining from “what could possibly help, for substantively raising the median sanity, or at least median sanity of the intellectual class?”. Things that help more “on the margin” are also useful and people should do them but, man, we really need a lot of sanity really fast IMO. I think we will need some big swings.
That said, obviously if you were going to do this, you’d want to start with cheaper ways to prototype it. (i.e. experiment with it on Lesswrong, Manifold, glosso, etc, and initially roll it out as an opt in beta thing on twitter).
But, I want to start by asking “when we simulate the good, smooth version of the product, implemented magically perfectly, does the big swing even seem like it’d work?” and iterate on that conceptually before worrying about the simpler dogfood prototyping.
The big swing needs to engage directly with “what seems incentive compatible with ‘most people don’t care about truth, really’ and “pundits and leaders actively resist submitting to processes that could rule them out” and “any process that has teeth will be a target of politicization”.
Whomst among us has not spent time down the resolution criteria mines. It’s not just about getting some criteria it’s about getting criteria which nail the issue. If people read the criteria for “Is the strait of Hormuz open by October 1st, 2026?” is it gonna match what they think?
This seems to be missing the point of the post, though. The whole point is “resolution criteria is basically a hard blocker on accurate predictions going mainstream. So how do we completely route around it?”
The post is proposing the hypothesis, that to square that circle:
you make it happen frictionless/automatically on a big platform people already use.
you make that happen by selling Musk on it (which seems hard but doable? modulo prototyping it in lower stakes places first and, like, you know, make a world-class UI and sort out novel algorithmic problems or whatever, but, “that’s the easy part”)
you use LLMs to automate it as much as you can (in particularly aggregating and summarizing the object level evidence).
you use community notes algorithm to handle “resolve complex predictions quickly without having to deal with resolution criteria, in a way that will hold up.”
use various UI tricks to make it feel like a fun game instead of a punishment.
I think the main thing the post is asking is “does #4 actually work?”. It feels intuitively to me like a mix of “LLMs to make an initial ruling and highlight evidence, community notes voting to resolve controversial stuff” should work. But, I’m interested in thoughts from people who have more experience with the gritty details o that.
(“is it actually possible to make this a fun game that mainstream people care about, or at least mainstream aspiring intellectuals, if you solve all the other problems?” feels like the second hardest part)
Hiring Vibe-wrangler Matchmaking Thread
I agree it’s quite difficult to solve this completely. But, it seems like you could make it so “On twitter, the default thing is if you try to pundit, the system nudges you towards legiblizing your predictions.”
This at least makes more of the world, on the margin, track “being correct” as a good thing to do. This can hopefully have flowthrough effects as the equilibrium shifts from “illegible pundits just look like status quo” to “illegible pundits look weirder.”
Yeah my impression is that prediction markets incentivize things that are more like “seek out poorly priced things and arbitrage them” rather than “make useful, correct predictions that anyone cared about.” (I have some vague sense that the shape of prediction markets tends towards not solving the problems I care about, although I haven’t thought about it too hard).
There totally could also be prediction markets involved here. But the primary purpose is to establish a record of “who predicted accurate things that anyone cared about” so that can become an easily referenced source-of-authority.
This needs to be incentive-compatible with “being a pundit”, I think. Like the idea is to make it so people who do pundrity on twitter are automatically nudged to make their predictions legible and gradable.
Nod. Also to be clear, I’d agree with “people trying to think through illegibly fake-feeling (to them) things rarely is helpful.”
But, just because it feels illegibly fake to one person doesn’t mean it feels illegibly fake to everyone. (I’m guessing you agree with that but aren’t sure what-if-anything to do about it?)
What I felt some excitement about, in this convo, was “hmm, it feels like a big problem that civilization doesn’t know how to handle illegibility. Sarah seems like she’s in a place where natively a lot of the ASI arguments feel fake to her. But, she’s familiar enough with the arguments, and maybe has updated about some pieces of them enough, that, maybe she can help move the conversation forward on ’man what do we do with the illegible problems, given that it’s hard to tell from the outside which ones are real?”