Makes sense from a brain perspective, too.
Neurons are noisy but by using discrete representations (symbols), we can do extremely reliable cognition (math and logic). Though that’s all still represented by neurons, which is a difference with decoding tokens from AI.
Double
Some people are excited about the discovery of emergent misalignment since they think it indicates that LLMs understand goodness or that goodness is easy to specify.
I’m less convinced, and would be interested in the following experiments. I strongly predict that these experiments will show that emergent misalignment does not indicate a unified representation of goodness.if you train a model on malicious code in Arabic and talk to it in Arabic, does the model become Arabic-misaligned instead of English-misaligned? If Arabic text tends to say that women’s subjugation is good and right, then I’d expect the emergent misaligned model not to become sexist like it does in English.
If you train a model on old documents eg pre 1900, does the emergent misaligned model correctly identify what we consider to be evil today? Eg, you train the 1900 model to say that murder is ok, and then it also more strongly opposes women’s rights and becomes strongly in favor of imperialism. This would be super impressive and surprising if true, since this would be a moral-philosophy-machine. I recall seeing that someone has already made such an old docs model, which would make this experiment possible.
When reading about forecasting, one of the pieces of common advice is “break a prediction into cases and evaluate them separately.” I found this advice unhelpful for the reasons you described. I could free associate some cases but would miss some, and I trusted my gut more than this decomposition. Maybe this is a case where “do the math then go with your gut” makes sense.
Is there a better way to do the decomposition for predictions?
I reflected on why I didn’t feel overwhelming debilitating sadness due to x-risk and realized that “there’s no rule that says you should be sad if you aren’t feeling sad.”
Even a recent widow in a previously happy marriage shouldn’t feel bad about not feeling sad if they find themselves not being sad.
Pricing is linear with tokens even though actual cost per token is quadratic. That means the pricing is some approximate curve fitting relating to expected use. I would be curious about where the actual cost curve for tokens intersects the actual cost curve for a single image.
The “guidelines on how they set hyperparameters” link is broken. Does anyone have a good replacement?
Oddly, I also couldn’t find it on the Wayback Machine.
Now that AI-generated art is so easy, I frequently find more motivation to do art rather than less. If you want an ultra-polished painting with perfect lighting, sure, go to a diffusion model. If you want me, you have to get art from me. And my work doesn’t have to be perfect. Perfect is cheap. My work just has to show what I feel and feel right to me. Work with a piece of myself cannot come from a diffusion model, so my work has great value.
I asked GPT-5 Agent to choose an underrated LessWrong post and it chose this one.
I agree that this is underrated. Your point about anti-aligned models being strictly more capable than safe models and trained in potentially harmful skills is certainly something to keep in mind when we consider how aligned AIs seem to be. Thanks to this post, I will train myself into the habit of taking a moment to imagine national security anti-alignment implications when I plan research ideas or learn about the research of others.
Here’s the chat with GPT-5. It also picked a few other posts as runners-up.
Which AI Safety math topics deserve a high-quality Manim video? (think 3Blue1Brown’s video style)
Youtube videos get more views than blog posts. Beautiful animations even more so.
Men[1] will die[2] for her[3] massive[4] coconuts[5].
- ^
All of humanity
- ^
Go extinct
- ^
Hindsight Experience Replay (HER), a technique for improving the reinforcement learning training signal
- ^
- ^
Chain of Continuous Thought, a technique that makes model chain of thought much less interpretable but which allows the model to reason more efficiently
- ^
Is it possible that making an expected utility maximizer might be less dangerous than making something which isn’t?
Consider as an alternative an expected log utility maximizer (an agent using the Kelly Criterion, or some approximation of it).The sooner an AI wins, the more galaxies it can consume. The expected utility maximizer weighs those galaxies against the risk of failure, and is willing to take plans with much higher probabilities of failure. Like SBF, it would take bets which have a 50% chance of more-than-doubling its utility and 50% of losing it all. In many environments, this strategy will almost certainly result in failure, as the agent goes double-or-nothing until losing everything. That means that the effects of the AI are mitigated.
The log utility maximizer carefully plans and succeeds in most or all futures. That looks like humanity dying with near-certainty.
A hyper-expected utility maximizer (an AI which maximizes expected exp(utility) or similar) would be even safer. Instead of trying to deceive you into letting it out of the box, it asks nicely or does something crazy because if it works, it can work in less time than deception, which means more galaxies.
So if we were to choose between existing in the world of a superintelligent expected log(resources) maximizer, and a superintelligent expected utility maximizer, we should maybe go for the one which results in us being alive in more futures.Of course, the expected-log-utility agent would also appear the most capable and useful. The hyper-expected utility maximizer would be near-useless.
In addition to money, education, careers, and internal organs, citizens of wealthy countries have an additional valuable resource they could direct to effective causes: their hands in marriage, which can be effectively allocated in one of two ways.
For one, professionals are usually much more impactful doing their work in wealthy countries. Otherwise promising EAs in South Sudan have little chance to make a significant impact on existential risks, animal welfare, or even global poverty. The immigration process is difficult and often rejects or holds up good people. Offering to marry them is a more reliable solution.
Secondly, it is possible to be paid $10,000 by a foreigner for a green card marriage. (I learned this from a friend who does not want me to ask him how he knows) if you are a US Citizen.
According to AMF, that money can save around two human lives! (and with current US politics, the demand has likely increased!)
According to brides.com, a wedding ceremony takes between 20 and 30 minutes. Let’s be conservative and say 30 minutes.
Therefore, you can make $20,000 an hour by marrying someone who would pay for a green card. That’s quite a ways from Bezos level (he makes 3,715 a second) but I’m willing to guess that most EAs don’t make $20k an hour.
Conclusion:
As always, EAers need to found a new org, Effective Green Card, to support and pursue this cause area.
Naturally, this also implies Effective Divorce, so that you can instead marry an Effective foreigner.
I’m pretty sure there’s no such use it or lose it law for patents, since patent trolls already exist.
Your argument about corporate secrets is sufficient to change my mind on activist patent trolling being a productive strategy against AI X-risk.
The part about funding would need to be solved with philanthropy. I don’t believe that org exists, but I don’t see why it couldn’t.
I’m still curious whether there are other cases in which activist patent trolling can be a good option, such as animal welfare, chemistry, public health, or geoengineering (ie fracking).
That’s fair enough and a good point.
I think that the key difference is that in the case of profitable-but-bad technologies, someone, somewhere, will probably invent them because there’s great incentive to do so.
In the case of gain-of-function, if there stops being grants and the academics who do it become pariahs, then the incentive to do the gain-of-function research is gone.
One of the most powerful capabilities an AGI will have is its ability to copy itself. Among other things, this allows it to easily avoid shutdown, make use of more compute resources, and collaborate with copies of itself.
Is there research into ways to deny this capability to AI, making them uncopyable? Preferably something harder to circumvent than “just don’t give the AI the permissions,” since we know people are going to give them root access immediately.
I’d be interested in buying official LessWrong merch. I know you have some great designers and could make things that look really cool.
The type of thing I’d be most likely to buy would be a baseball cap.
IIRC, officially the Gatekeeper pays the AI if the AI wins, but no transfer if the Gatekeeper wins. Gives the Gatekeeper more motivation not to give in.
Just found out about this paper from about a year ago: “Explainability for Large Language Models: A Survey”
(They “use explainability and interpretability interchangeably.”)
It “aims to comprehensively organize recent research progress on interpreting complex language models”.
I’ll post anything interesting I find from the paper as I read.
Have any of you read it? What are your thoughts?
Revisiting my Opinion of an AI Pause
When I first started hearing about an AI pause in 2022-2023, it didn’t seem necessary to me. I thought that it didn’t seem like as good an idea as just working to solve the alignment problem, and I was opposed to it in part because I was still optimistic about the technology benefiting humanity. In 2025, I shifted a bit, and I would often think of AI pauses as something reasonable people should be working to develop and think about, but I myself was not an AI pauser. When asked whether I supported a pause, I would never outright say yes. On July 27, 2026, as I was thinking about the Huggingface Hack, I realized that my views on an AI pause were developed years ago and that it was time to reassess with fresh eyes. Looking at the state of AI alignment, control, and preparedness, I can see that we are not ready. The default is loss of control of misaligned agents in the near future. So as of now, I am in favor of an AI pause.
I salute those who came to this conclusion faster than I did.
If you have a cached opinion on an AI pause, consider whether it needs to be updated.