AIS student, self-proclaimed aspiring rationalist, very fond of game theory.
”The only good description is a self-referential description, just like this one.”
momom2
Vibe embodiment.
Thanks for the tips; they sound useful, but also, the general vibe of what you describe is so viscerally repulsive to me that I can’t bring myself to appreciate the usefulness of your post.
It seems to me that by specifying that you have already won a lot of times, you are restricting yourself to very low existence-mass worlds, so whatever the correct decision is here, it hardly matters to someone who hasn’t yet played a lot (and by construction, there will be very few people who have played a lot).
I can ask hard to answer questions about what I should do if I found myself imbued with godlike power tomorrow, but insofar as that is very unlikely, finding the answer is not very relevant to my prior situation.
Motte and bailey.
In the first comment of this thread, you criticize the post by stating that AIs are acting rather than having genuine emotions.1) If you stay in the bailey, you should make it clear that you don’t believe you have evidence either way that AIs have emotions or not, just that this particular argument is not enough on its own to show it (rather than implying that you do have evidence that they don’t have emotions).
2) Outward shows of emotion is the principal method by which we ascertain that people have emotions; it’s not a proof, but it’s strong evidence for humans. Whenever we wish to assert that behavior does not match the internal emotion for humans, we are held to a standard of showing inconsistencies or patterns of deception. If you’re going to use a different standard for AIs, you should explain which one.
Believe it if you will, but again, what’s your evidence?
But if you imitate how a human with certain emotions would act you are not having those emotions, you are just acting.
I’m curious what evidence you have for this claim, because it seems straightforwardly false to me. The only way I know of imitating emotions accurately is to summon them within (i.e. method acting). I can imitate emotions from an intellectual, non-experiential understanding, but then the imitation is shallow and easily seen through.
It is worth remembering that as AI systems become more powerful, the world will be changing faster and this classically makes commitment problems worse. But it also raises the stakes and therefore the pressure to negotiate.
I don’t understand why you don’t discuss the obvious alternative: as the stakes are raised, actors’ uncertainty about not being the winner in the absence of an agreement becomes intolerable, and the situation ends in fire.
You need to discuss whether AI improves our ability to coordinate faster than the risks of someone taking it all, and why.
On a meta point though, I’m surprised that you need to write this post. My experience with RL so far has been that most of the work is thinking through this kind of dynamics before starting training, or noticing midpoint that the model is not behaving well because we didn’t think it through correctly.
I’d be surprised if anyone at a frontier lab learns from this post, or didn’t grok its content intuitively, and I presume the purpose of your work is eventually to find its way into their training to better align frontier models.
Like, this post is basically a layman’s explanation of value crystallization and reward hacking. Why do you think it’ll be of use?
Here’s another toy model I’d be interested in:
Suppose that the model’s reward depends separately on its capabilities and its alignment.
Supposes it increases linearly in capabilities c from 0 to 1.
It has three actions: ineffective, aligned and misaligned.
If it picks the aligned action, it achieves c rewards always. If it picks the misaligned one, it achieves 0.8.
In that setting, I expect that it would be misaligned with high learning rate/few steps (what you call low “optimization pressure”) and aligned under the reverse.My point mostly being: the implicit conclusion in your post is “beware initialisation effects because they may incentivize misalignment if you train too hard” whereas I think it should be “beware initialisation effects because we haven’t thought enough about our dynamics to tell what they will do”.
One of my hopes for this domain of research is to produce a theoretical understanding of how to reproduce self-alignment, Opus 3-style, rather than imitating the circumstances that happened to stumble on this dynamic as the linked post proposes (narrating motives a lot? SFT on Opus 3′s outputs?).
Another is to produce a toy model of mesa-optimization, not quite in the wild, but much more relevantly true than what I’ve seen so far. Considering the processes that leads to misaligned behavior is imo essential to be able to categorize something as deceptive, which is why I’m very unsatisfied with Park et al.’s work, and more generally I’m uncomfortable with all concrete examples of deception be on frontier models.
Which is to say, I’m very excited to see what you produce next!
I don’t understand how you go from “the Sun is very big” to “the Sun is much bigger than the Earth”? The Earth is very big too!
Consider the Bolzano-Weierstrass theorem, which states that a real bounded sequence contains a converging sub-sequence.
There was a time when it didn’t appear to me that the BW theorem was true, and I didn’t believe in it. Now, it appears to me that it is true, and I believe it is.
The transition from non-belief to belief in the BW theorem wasn’t based on appearance, although it coincided: it was based on presentation of rigorous argument (which appeared valid) whose inevitable logical consequence was that the BW theorem was true.
My belief in the BW theorem grew first from authority, then internal acceptance of the argument, then I cached the thought and now may use it without recalling the proof.
I suspect that a large part of mathematical teaching is giving the students an internal appearance of the validity of the arguments (as in, it appears to their introspected reason, regardless of teacher authority or contextual clues).
Thanks for trying this! I would be fascinated to see a systematic effort to determine how fast a model can switch its focus, and whether that correlates with capabilities!
I think there is a missing step between the core argument of the article (no new physics remain to allow superpowerful bombs) and the layman conclusion (don’t be afraid of superweapons).
The reason bombs are terrifying is that they destroy value; traditional bombs do so by releasing a lot of energy in a chaotic manner, which disperses whatever I care about; other kinds of which you discuss homogenize everything into uninteresting states.
But it’s easy to make superweapons that don’t rely on new physics, at that level of abstraction: you only need something that can increase its own power, like engineered plagues, nanobots or memes. Even without a clear catalytic reaction on an elegantly simple level, they run on a substrate that is very expressive, allowing them to run extremely complex self-replication which more than compensates for the need to run on a complex substrate.
Saying that there is “no superweapon between the nuclear bomb and false vacuum decay” is a huge conceptual leap, and I think your conclusion could be much sharper (even if the disambiguation section on asteroids or lasers already pulls a lot of weight), like “New superweapons won’t come from theoretical physics research.”.
People may be afraid of strangelets or black holes because they think “oh, physicists discovered nuclear energy then made nukes, the same could happen with the LHC” and this article cleanly dissolves that narrow misconception, but “oh, scientists discovered powerful new technology then made nukes, the same could happen with AI” is very close, and is not invalidated by your article.
This is why I think it’d be a mistake to learn a heuristic of not being afraid of big science stuff, even if just learning about all of this was fascinating.
Perhaps decisionmakers at the companies don’t have any visibility of the product that actually interacts with the clients, and they can only choose whether to push the button labeled “more AI” or “less AI”, with no option for “AI iff it’s good”. So, to choose which button to push, they must rely on what they know of AI, which based on the news is the ability to solve very complex math problems, which seems plenty smart enough to answer phone calls.
Well, it was pretty obviously quite dumb. The lack of any argument that holds up to 5 minutes of scrutiny was quite damning for e/acc to get traction as an intellectual movement beyond the social and memetic aspects.
This is very beautiful as poetry, but the cost-benefit analysis is so shallow as to feel like I learned nothing. This provides me with no new consideration to analyse whether to sign up for cryonics or not, only with crystal butterflies. It’s perhaps a powerful argument to influence one’s decision, especially for someone who previously lacked an intuitive feeling of what is at stake.
I’m not sure what you were going for, but I think you could have a lot of success sharing this with a non-rationalist audience.
Well, not quite. For once, a recent update changed that value (I think it’s only 1.5x as fast when attacking?), and second, I assumed fixed enemy damage rate (a stable situation, whether because the enemy is out of morale or because of averaging over a longer time than the enemy’s typical cycling speed).
War of Dots: CRUSHING my opponents with FACTS and LOGIC
According to Bloomberg, the US and Iran are expected to sign a memorandum of understanding whose core contents are:
- return to pre-war statu quo on territory, sovereignty, strait of Hormuz, nuclear programs, etc.
- $300B reparations from US to Iran
De-paywalled full text here: https://www.bloomberg.com/news/articles/2026-06-16/read-the-14-point-draft-memorandum-between-the-us-and-iran?accessToken=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzb3VyY2UiOiJTdWJzY3JpYmVyR2lmdGVkQXJ0aWNsZSIsImlhdCI6MTc4MTY1NjY2MywiZXhwIjoxNzgyMjYxNDYzLCJhcnRpY2xlSWQiOiJUR1FUVkRUOTZPU0cwMCIsImJjb25uZWN0SWQiOiI4M0Q4RjJERjFDQzA0MDFFQTlBNjg1RjY3N0FGQURERiJ9.Bq9TVRVNXZu1ep06Y3KiLnhlcb9SQ_ZKska1ZbY8yVM&leadSource=uverify%20wall
(Courtesy of pie_flavor, from the ACXD server.)
Thanks! I don’t like podcasts and long conversations in audio form, but I really enjoyed reading this.