Semi-anon account so I could write stuff without feeling stressed.
Sodium
In AI safety, anything that accelerates AI progress and brings us closer to RSI is effectively a force-minimizer.
This doesn’t take into account AI labor, right? If you speed up AI’s ability to do AI safety research, then you have more AIs working on alignment before it’s too late as well. The point about having better LLM AI safety researchers is that you are increasing the total amount of quality weighted AI safety labor before the point of no return.
Concretely: working on alignment directly seems just as differentially good for superalignment/handoff compared to any kind of capabilities acceleration
I think this point is made pretty sloppily? Like what types of alignment? The examples you provide below:
I’m more saying that this particular bet looks like it takes on a lot of unnecessary risk when there are other things that look equally promising. If the goal is to hand off to AI early-ish (which I’m not claiming is good or bad) then even some prosaic alignment things like trying to systematically understand midtraining or trying to “solve” eval awareness seem less risky while being equally productive and maybe more tractable.
All of those prosaic stuff seems like it could be sped up if we had better LLM AI safety researchers
When it comes to other interventions (e.g., Pause adovcacy,) I think a lot of the dominating considers for “should you make LLMs better at AI safety research or advocate for a pause” looks similar to “should you do prosaic AI safety research or advocate for a pause.”
You know, for example, which of these count as doom and which don’t:
A young ASI tries to use it’s first-mover advantage to take over the world and prevent other ASI competitors from emerging. In doing so it sparks a war against humanity where it eventually loses,[3] but it kills 10% of all humans in the process.
ASI empowers a single person or small group of humans to become tyrants and lock in a permanent authoritarian regime where almost all other humans are subject to the whims of the tyrant(s).
A state actor or terrorist group uses narrow AI to build a bioweapon (eg AlphaFold but for plagues), and that weapon gets out and kills literally everyone.[4]
A great power conflict starts for reasons mostly unrelated to AI, but advanced AI systems are deployed on the various battlefields, and eventually (in part due to the speed and ruthlessness of AI) it escalates to a thermonuclear war that kills 99% of humanity.
Humans “merge with the machines” in a way that, from the outside, looks an awful lot like all of humanity being subsumed by inhuman machinery.
Humans become incapable of competing with machines in almost all sectors of the economy. While the rule of law persists, the political landscape is also dominated by a variety of ASIs with various goals, and they don’t redistribute significant wealth into human hands. Only a small number of humans survive in the long-run, relegated to being a historical curiosity.
ASIs are developed in a broadly corrigible way, leading to extreme abundance, and the offense-defense balance means that the world is basically safe. But everyone (even the Amish) eventually stop having human kids because other things (including AI “kids”) are way more fun/satisfying. Even though lifespans are really long, people still gradually die off, and eventually only the machines remain.
sure, I mean, each one of those took less than 45 seconds to decide.
not doom
doom
doom
doom
doom
doom
not doom
Another clear advantage for inoculation prompting is that you don’t have to train against internals in any way, so interpretability is left as a much more “held-out test” for alignment. Preventative steering isn’t directly training against internals, but it still manipulates them a little.
Strong downvoted the post because:
The tone is incredibly off-putting and imo not appropriate for lesswrong.
The advice doesn’t actually seem very actionable or good, nor does it seem like this post will actually change how people act
The Word Games section appears entirely AI generated without being noted as such. This violates Lesswrong’s LLM use policy.
LessWrong 2.0 Is a Website, Not a Culture or an Authority
Man I don’t think I agree with this. Clearly what draws people here is the culture and the vibe.
And I do think I relate to twitter and Elon Musk in a similar way as I relate to LW and the LW team? I have preferences over what content should be promoted on Twitter, and I would be happier with Elon if he shared my preferences. Ditto with Lesswrong.
On a similar note: I feel generally pessimistic about doing extremely prosaic alignment research[1] that’s intended to beat baselines[2] outside of Anthropic, because you don’t have access to the SoTA baselines.
Edit: however, in this case it might be the case that narrow fine-tuning on eval-cooperativeness right before you run the evals could be good? You can treat it as spiritually similar to eval awareness steering.
In worlds where AI alignment can be handled by iterative design, we probably survive.
I’m curious whether you still believe in this? I think there are currently lots of alignment issues that are totally fixable by iterating on them, but they aren’t. Naively extrapolating, it’s possible that even if various future alignment problems are solvable via iteration, they might not be.
My guess is that at the current state, conditional on (misalignment at the point of no return will kill everyone) AND (this misalignment is fixable using iteration), p(misalignment) is still around 50%.[1]
- ^
Obvious this is a bit fuzzy since iterative design worlds would have much more continuous looking PONR and misalignment issues, but that’s the sort of vibe I have about the current state of the AI race.
- ^
I find myself confused at other people’s surprise. The thing we are seeing seems to me like the obvious thing you would expect to see from a generic “you get what you optimize for” objective.
fwiw I strong upvoted the post not because it said something surprising, but because it explains our current observations very well.
I think the Walmart and Toyota case is less interesting because they’re not creating “new” consumption. Like Walmart has a huge revenue because it’s captured a big slice of people’s overall consumption. If Walmart’s revenue doubled next year, it’ll probably because they got a bigger slice, not because people are suddenly buying twice as much stuff.
I don’t think that’s true because Randy would appoint Republicans throughout government/be more captured by the Republican party’s interests? Like it depends on how much you like Randy-flavored Republican in executive and judicial roles. I think there’s probably a huge difference for what types of judges Randy and Donna would nominate, for example.
I guess this is more true for Presidents than it is for Senators/Representatives (since an Republican congressperson will vote for the Republican Speaker of the House/Senate Majority Leader, who has a lot more power than any individual congressperson.)
While you can rip my epistemic qualifiers from my cold dead hands, probably, I sometimes grudgingly admit that the sentences I write have a certain kind of meandering quality to them, often going on for so long that by the time the reader has reached its end, the reader will have forgotten how it started.
The fact that this sentence is meandering and makes it easy to forget how it started by the time one reads to the end makes it an instant banger.
I mean this is assuming that ASI is aligned and chooses to not manipulate public opinion right? I agree that assuming that it’s misaligned, then there’s not much to talk about.
(You can also imagine multi polar worlds where different AIs police each other for superpersuation.)
this way of reasoning seems like somewhat naive consequentialism.
Maybe? It is hard to reason well about these things given my strong emotions towards the admin.
But I do think the current administration is uniquely terrible by American standards.[1] It attracts and gives power to incompetent sycophants with no moral boundaries.
There was something Eliezer said about Bernie Sanders recently that really resonated with me recently:
[T]hank you also for consistently trying to do as seems right to you over the years, a stance that has grown on me as I have had more chance to witness its alternatives.
Having Trump as the president really just seems like it would be terrible for AGI governance because he is a terrible person. I’m sorry, I really don’t think there’s a more “precise” way to put it. Character matters. Trump doesn’t even pretend to be a kind person/is not under much pressure to appear to be nice.
(To be clear, I agree that, all else equal, it would be good for the Iranian regime to fail. Alas, all else would not be equal. While I think it would definitely be bad for your soul[2] to do things in the realm of “sabotage the American economy/military operation in order to make our president look bad,” I don’t think I’m obligated to stop my enemy when he is making a mistake either.)
I think the most important effect of the war is that it makes Trump less popular/powerful domestically (even if a miracle happens and he gets some sort of deal.) This is good because the less power he has (e.g., Republicans lose the senate in the midterms), the more likely we are to navigate AI development in a sane way. I think if you put
anynontivial*weight in short timelines, the AI considerations likely dominate everything else.
*edited any to nontrivial. Like, maybe 10%+ pre-Jan 2029
I flushed out a similar idea in (Maybe) A Bag of Heuristics is All There Is & A Bag of Heuristics is All You Need
Hong Kong had seats in its legislature designed for special interest/business groups (see Functional constituency on wikipedia). I don’t understand it very well though.
I don’t think of this as relevant to any sort of doom really. I think of “output math” as habit that the model has picked up, since it did a ton of this during training.
See e.g., Nemotron 3 Nano here using code when asked a religious question on openrouter:
This is my least favorite fact about Claude. I don’t think it’s actually genuine when using “genuinely” (or at least, when it describes something as “genuinely X,” I often find that the thing is in fact not X.)
My guess is that whatever constitution-inspired post training process they used gave birth to a reward model that likes of text outputs that contain “genuinely.”
Hmmm I don’t think the OpenAI model that opened the PR was engaging in metagaming? It wasn’t “reasoning about feedback mechanisms or oversight that sit ‘outside’ the scenario’s narrative.” Similarly, I don’t think Opus 4.6 was metagaming by looking for free compute online? Both seems closer to like, “over-eagerness” and doesn’t require like sophisticated reasoning about oversight mechanisms.