It comes down to whether you have a strong prior that Soryu does not just intrinsically value hurting people (good policy) or a trapped prior that Soryu does not just intrinsically value hurting people (very dangerous)
Hastings
There’s two sides to whether this is delusion. Would I take memories revealed this way as strong evidence in a court of law? Probably not. Would I let a Leverage alumni talk me through unclenching a muscle, because I’m a big tough rationalist and woo fairy spirits can’t hurt me? Also no.
If I pretended that “inner work” was a proposed chemical medication and not a ritual, I would be highly skeptical of claims that it was revealing existing problems, not creating them, if one of its common outcomes was years-long breakdowns.
A community of 30 or 40 language models that don’t have language yet can’t invent a language to communicate. This is a built-in human capability (see https://en.wikipedia.org/wiki/Nicaraguan_Sign_Language) so it will emerge sooner or later by the bitter lesson. Once LLMs and algorithms get to a point where they can do this we’ll finally get answers about their subjective experience in their own words without reference to human experience, although there will be other issues at that point.
Concretely: there are lots of things that we are pretty sure won’t happen in such an experiment (such as carefully and correctly isolating the group of cooperating generative transformers from human generated text, training from random initalization on RL in some environemnt, and then coming back and finding a recognizable congregation of fire and brimstone pentecostals) which means that would get a huge torrent of information about subjectivity if one of the impossible outcomes happened, which means we get tidbits of information on subjectivity when they don’t.
Hard to nail down exactly what I mean here, but I see traces of an objective account of what it’s subjectively like to be a person in correlations between oral traditions of tribes that share human ancestry but not any cultural connection. This could be done better with ideas expressed in languages that emerged recently (with the prototypical cases being schools for the deaf or hereditarily deaf populations inventing sign languages) but I haven’t found any pre-translation oral traditions in these languages.
I think I was directionally correct here https://www.lesswrong.com/posts/eqrtdyphQFefc2ox9/hastings-s-shortform?commentId=WtLhZ9shcbZ8Ki38g when I supposed that it would be increasingly possible to track “secret” large scale training runs purely by their escape attempts into the public internet, although I was wrong in that I expected it to look like models quietly but detectably surfing the web for information useful for scoring high, not models blatantly and extremely loudly hacking in pursuit of useful information.
I like this fleshman framing as it lacks the following issue:
Steelmanning makes a great deal of sense in terms of first order effects, but what I observe is that once a diverse group is in the habit of steelmanning, the group is exploitable where its more effective to vaguepost in such a way that all the factions steelman you into their camp, than to clearly spell out your belief. Incentive gradients against clear writing are like, pretty bad.
Fleshmanning picks the median instead of the extreme and it seems like factions can agree on median, or at least not predictably disagree like they do when picking best, where this predictable disagreement gets them pwned
The moods are so missing that it almost reads as a literary piece with an unreliable narrator.
We were really determined not to start a cult, which is why we repeatedly watched the YouTube video how to start a cult.
We were not an intellectual monoculture: for diversity we had a bunch of neo reactionaries. The allegations of us being a Peter Thiel cult are unfounded.
We didn’t have any psychotic breaks. None. But if we did, that would be perfectly normal for our type of groundbreaking organization, and an understood danger of the process. But we didn’t have any.
Some things I can’t nail down the point of but do:
Voting. Correctly working out the point of an individual voting requires a working decision theory but even though I don’t have that I’m happy to basically pascal’s wager it. I picked this as an interesting example as I would have no qualms asking someone to help me go vote if I needed a ride or something.
I’m writing a new programming language, even though I’m not sure what the point is. On base rates the world doesn’t need this. I haven’t strictly asked for help with this before I work out what the point is, but I have definitely “Hey look at this thing I’m making”ed my friends with it enough that they could ask me why. I picked this as an example in the hopes that someone would ask for details, after which I would “Hey look at this thing I’m making” them.
Reddit: this doesn’t have a point and I would ask for help to not do this if I thought anyone could help.
Working through “A monad is a monoid in the category of endofunctors,” I was able to learn the definitions of monoid, category, and endofunctor pretty easily and have been blocked on “in” and “of” for significantly longer. (vague claim that this generalizes)
Credentials: I have beaten ultrasound with the ML stick until it yielded a few times ( https://scholar.google.com/citations?hl=en&user=O1xhOlUAAAAJ ) it was terrible and I have lasting resentment toward an imaging modality, though great fondness for all my colaborators. Many failed projects that did not make the google acholar
ultrasound is awful to work with in traditional medical image processing technologies and pretty darn bad in 2015-2025 convolutional medical imaging ai technologies. Magnetic Resonance is so much better when its the right tool for the job that replacing it with ultrasound for cost reasons is a tarpit. There are cases where ultrasound is better but these center on leveraging an extremely talented operator who is manipulating the probe manually, which you lose in any of these bath based approaches. I’d be surprised if this goes anywhere.
Worth noting that the only way to get a pull request accepted to stockfish is to beat stockfish at a different form of centaur chess (manually modify stockfish and then have your changed stockfish beat the original in a series of games) and this happens regularly.
I guess the main useful insight here is that .01% successful escape attempt during training sounds very aligned, and 100,000 successful escape attempts during training sounds very not aligned
Fermi estimate: Lets say each training episode for Claude Mythos cost a dollar, and Anthropic spent a billion dollars post-training Mythos. Now, in 0.01% of training episodes, Mythos broke out of the training environment entirely to get useful data from the public internet, which is a billion * .01% = 100,000 requests. I wonder how fast this last number is growing? Could an entity like the NSA or CCP, with taps in enough internet infrastructure, detect 100,000 weird requests if it looked hard enough? Potentially an avenue for monitoring and verifying pauses, analogous to detecting nuclear weapons tests by monitoring for radiation leaks. Spotting the containment breaches from a frontier lab’s training run, without their cooperation, would be an excellent safety project I think- not an easy project, but easier than half the stuff we are trying. It would look a lot like “ordinary” agentic AI traffic, but coming from weird IP addresses, and imho likely pretty internally homogenous- from my understanding algorithms like the one used by deepseek attempt the same task a large number of times in a batch with different seeds to get a signal, which would produce bursts of traffic to answer the same question, as a first spitball of a signal. Would be much easier with cooperation from at least one frontier lab since they know what the real leakage from their training runs looks like, but could monitor all of them.
Thanks for replying, this is an important update for me.
I have a very similar setup, and have found similar convenience instead of inconvenience: specifically, agents can compile or configure tools without bonking each other or me. I also usually run agents on a server instead of my laptop, just so that they don’t stop what they’re doing if I close my computer to walk to a coffee shop, and so there is very little additional friction to virtualizing since its ssh and tunnels either way.
Are alignment researchers seriously not even keeping the ai in a virtual machine? This feels like one of those Hastings Is Not Living in Berkeley and so is Frequently Surprised that Things Weren’t Jokes moments.
If LLMs are alignable, the question isn’t whether LLMs can scale to ASI, its whether 1) LLMs are the equilibrium, compute optimal way to be intelligent or 2) we can coordinate to stay off the equilibrium in the time between discovering it and working out how to align it. 1) seems galactic-ally unlikely to me so LLM alignment is entirely reliant on 2) (so we would be well served to develop that capability)
There’s a bit of flexibility in that the compute optimal way to build intelligence could be llm-like enough for alignment to generalize, but this seems unlikely- for example, LLMs pre rlvr were way easier to make behave but it is unthinkable to coordinate on not doing rlvr. I expect more rlvr-like innovations.
I’ll try to phrase my thoughts in a different way: in rationalist spheres, its’s extremely low social status to ponder “maybe X is just puppy kicking evil without any complex hidden upside or justification” and as a result this hypothesis doesn’t gain probability mass even when it makes good predictions.
After the Waluigi effect was discovered in LLMs, I had to reckon with the reality that minds exist which pretty well know what the right thing to do is and take the opposite action because it is opposite.