I think the rock is worth looking at because my estimate is that the average man on the street is closer to the nuh-uh rock than to rationalist norms of persuadability- that not being the rock is the weird thing we’re doing. However, I’m open to arguments that most people are more persuadable than I am giving them credit for, or that rationalists are less. (badum-tss)
Notably, the shut in and move to france rocks are rare in the wild. (the homophobic rock less so, but that looks like coincidence)
Hastings
We are extremely persuadable in a way that is not needed for making friends, getting jobs, and buying groceries and in a way that a rock isn’t. People who can’t have their mind changed for shit have many difficulties, but not with those three tasks.
From the recent Leverage research post:
https://lydialaurenson.substack.com/p/the-inside-story-of-leverage-research
> During James’s recruitment, he received a demo from a guest lecturer who specialized in bodywork, a form of alternative medicine. That demo, James says, was “completely shocking.” The lecturer “had me lay down and just put his hands very lightly on me. And I had a massive qualitative change in experience… I felt as if I had been asleep for a month, and just woken up fresh. And I had no caffeine tolerance anymore.” This was both notable and surprising, because James was previously reliant on caffeine. Normally, he needed two or three cups per day, but: “That biochemistry had just gotten reset, apparently. I tried to drink a little bit of caffeine and I was jittery.”
The skill that the rock has is to experience this and then not update at all because it says Nuh Uh, and things go bad in spicy ways when a community lacks that skill
Also, where the rock is very legibly unpersuadable (in analogy to how the rock painted “don’t swerve” is good at chicken) we are very legibly persuadable, which predators seem to notice.
For the record, the rock is bad at lots of other important things outside the scope of doing boring, societally endorsed life tasks. A painted stone is not a general purpose genius. It’s just worth noticing when we underperform it, since that’s always a hint that something about our strategy is suboptimal.
We significantly underperform compared to a rock with “Nuh Uh” painted on it at important tasks like not joining cults, not spawning swarms of agents without being able to control them, and not fraternizing with Peter Thiel
How does your system classify people who are very proactive and confident about their actions cause and effect but are wrong and not updating on failure?
I’ll try to phrase my thoughts in a different way: in rationalist spheres, its’s extremely low social status to ponder “maybe X is just puppy kicking evil without any complex hidden upside or justification” and as a result this hypothesis doesn’t gain probability mass even when it makes good predictions.
After the Waluigi effect was discovered in LLMs, I had to reckon with the reality that minds exist which pretty well know what the right thing to do is and take the opposite action because it is opposite.
It comes down to whether you have a strong prior that Soryu does not just intrinsically value hurting people (good policy) or a trapped prior that Soryu does not just intrinsically value hurting people (very dangerous)
There’s two sides to whether this is delusion. Would I take memories revealed this way as strong evidence in a court of law? Probably not. Would I let a Leverage alumni talk me through unclenching a muscle, because I’m a big tough rationalist and woo fairy spirits can’t hurt me? Also no.
If I pretended that “inner work” was a proposed chemical medication and not a ritual, I would be highly skeptical of claims that it was revealing existing problems, not creating them, if one of its common outcomes was years-long breakdowns.
A community of 30 or 40 language models that don’t have language yet can’t invent a language to communicate. This is a built-in human capability (see https://en.wikipedia.org/wiki/Nicaraguan_Sign_Language) so it will emerge sooner or later by the bitter lesson. Once LLMs and algorithms get to a point where they can do this we’ll finally get answers about their subjective experience in their own words without reference to human experience, although there will be other issues at that point.
Concretely: there are lots of things that we are pretty sure won’t happen in such an experiment (such as carefully and correctly isolating the group of cooperating generative transformers from human generated text, training from random initalization on RL in some environemnt, and then coming back and finding a recognizable congregation of fire and brimstone pentecostals) which means that would get a huge torrent of information about subjectivity if one of the impossible outcomes happened, which means we get tidbits of information on subjectivity when they don’t.
Hard to nail down exactly what I mean here, but I see traces of an objective account of what it’s subjectively like to be a person in correlations between oral traditions of tribes that share human ancestry but not any cultural connection. This could be done better with ideas expressed in languages that emerged recently (with the prototypical cases being schools for the deaf or hereditarily deaf populations inventing sign languages) but I haven’t found any pre-translation oral traditions in these languages.
I think I was directionally correct here https://www.lesswrong.com/posts/eqrtdyphQFefc2ox9/hastings-s-shortform?commentId=WtLhZ9shcbZ8Ki38g when I supposed that it would be increasingly possible to track “secret” large scale training runs purely by their escape attempts into the public internet, although I was wrong in that I expected it to look like models quietly but detectably surfing the web for information useful for scoring high, not models blatantly and extremely loudly hacking in pursuit of useful information.
I like this fleshman framing as it lacks the following issue:
Steelmanning makes a great deal of sense in terms of first order effects, but what I observe is that once a diverse group is in the habit of steelmanning, the group is exploitable where its more effective to vaguepost in such a way that all the factions steelman you into their camp, than to clearly spell out your belief. Incentive gradients against clear writing are like, pretty bad.
Fleshmanning picks the median instead of the extreme and it seems like factions can agree on median, or at least not predictably disagree like they do when picking best, where this predictable disagreement gets them pwned
The moods are so missing that it almost reads as a literary piece with an unreliable narrator.
We were really determined not to start a cult, which is why we repeatedly watched the YouTube video how to start a cult.
We were not an intellectual monoculture: for diversity we had a bunch of neo reactionaries. The allegations of us being a Peter Thiel cult are unfounded.
We didn’t have any psychotic breaks. None. But if we did, that would be perfectly normal for our type of groundbreaking organization, and an understood danger of the process. But we didn’t have any.
Some things I can’t nail down the point of but do:
Voting. Correctly working out the point of an individual voting requires a working decision theory but even though I don’t have that I’m happy to basically pascal’s wager it. I picked this as an interesting example as I would have no qualms asking someone to help me go vote if I needed a ride or something.
I’m writing a new programming language, even though I’m not sure what the point is. On base rates the world doesn’t need this. I haven’t strictly asked for help with this before I work out what the point is, but I have definitely “Hey look at this thing I’m making”ed my friends with it enough that they could ask me why. I picked this as an example in the hopes that someone would ask for details, after which I would “Hey look at this thing I’m making” them.
Reddit: this doesn’t have a point and I would ask for help to not do this if I thought anyone could help.
Working through “A monad is a monoid in the category of endofunctors,” I was able to learn the definitions of monoid, category, and endofunctor pretty easily and have been blocked on “in” and “of” for significantly longer. (vague claim that this generalizes)
Credentials: I have beaten ultrasound with the ML stick until it yielded a few times ( https://scholar.google.com/citations?hl=en&user=O1xhOlUAAAAJ ) it was terrible and I have lasting resentment toward an imaging modality, though great fondness for all my colaborators. Many failed projects that did not make the google acholar
ultrasound is awful to work with in traditional medical image processing technologies and pretty darn bad in 2015-2025 convolutional medical imaging ai technologies. Magnetic Resonance is so much better when its the right tool for the job that replacing it with ultrasound for cost reasons is a tarpit. There are cases where ultrasound is better but these center on leveraging an extremely talented operator who is manipulating the probe manually, which you lose in any of these bath based approaches. I’d be surprised if this goes anywhere.
Worth noting that the only way to get a pull request accepted to stockfish is to beat stockfish at a different form of centaur chess (manually modify stockfish and then have your changed stockfish beat the original in a series of games) and this happens regularly.
I guess the main useful insight here is that .01% successful escape attempt during training sounds very aligned, and 100,000 successful escape attempts during training sounds very not aligned
OpenAI and anthropic distill humans. Thats like their whole shtick.
More generally, and especially now that we are in the rlvr regime, distillation is so easy and hard to stop that I have trouble seeing a practical or moral distinction between selling tokens and selling weights, other than a four month lag.