Data point: I attended the London Haskell meetup today and told a few people I was on sabbatical and hoping to help with AI safety, because it’s clearly the most important thing right now. Reactions ranged from neutral to “oh yeah, obviously”; no one registered to me as surprised or skeptical.
philh
or in the case of a similar Anthropic incident, socially engineering humans
It’s not clear to me what in the linked post this is referring to?
No, I agree justice and morality are different. But like, if a legal system says “if you smoke pot you get life in jail”, I say “that is not just, this legal system does not even seem to be trying for justice” rather than “I guess that’s just, even though it’s deeply immoral”.
Absent a legal framework you cannot have justice kinda by definition.
Strong disagree! I judge legal frameworks against my preexisting conception of justice. I don’t decide something is just or unjust merely because a legal framework judges it one way or another. (Though my own judgment may take into account “there is a legal framework in play here which says...”)
My (mostly off-the-cuff) answer is that [ex post facto prosecutions under a “no ex post facto prosecutions” legal framework] are a very different beast than [ex post facto prosuctions under an absence of legal framework].
I don’t take Habryka’s “can people stop saying this” as “this is false” but as “this is a bad thing to focus on in context”.
If someone uses a model to do something very harmful, maybe via jailbreak or an open weights fine-tune or something like that, it’s not clear to me that it’s relevantly the given AI company’s fault. If you sell someone a gun, and they do murders with it, you aren’t liable
Some things I think are significant here:
Guns are essentially a commodity, in that you can buy them from lots of companies, and one company’s gun is much like another.
Gun manufacturers have no ability to steer how the gun is used, after it’s sold. If we make them liable for murders committed with their guns, we (approximately) get fewer-or-no guns, not “guns that are better at distinguishing murder from self-defense”.
Lots of guns are in fact illegal to sell.
A term and thesis coined by Duncan Sabien. I ought to write it up at some point, but meanwhile perhaps many readers, like myself, will find a whole useful thesis immediately apparent just from seeing the phrase “Arrogance of the Humbled”.
Duncan himself has written No Arrogance Like That Of The Recently Humbled.
London rationalish meetup − 2026-09-06
Does this make LLMs much more deterministic?
If I understand right, I think it needn’t, at least for some prompts? (Which isn’t the same as saying it doesn’t as-implemented. Also I am not confident I understand right.)
When generating each token, the prng seed used is a hash of the last four (configurable) tokens plus the secret key.
If you’re generating something like
<user>Write a poem</user><think>User wants me to write a poem</think><assistant>O, codswallop, how I beseech thy shrivelled kidneys</assistant>then the text you can expect to get run through the detector is only what’s inside<assistant>and maybe<think>. You won’t be asked to checka poem</think><assistant>O. So ata poem</think><assistant>, you can actually just use a random seed, and then you’d get the same amount of diversity in the initial four tokens as you would have without watermarking.Doesn’t help if the prompt has low diversity in the initial four tokens, like “count to four and then write a poem”, or if two different contexts with the same last-four tokens otherwise sync up in terms of what they predict.
My response is that this is going to be annoying, and also doing it is clear consciousness of guilt and far worse than the initial AI use
I’m remembering how go players disempower themselves to AI:
we learned from the many examples of cheating and player confessions that idle curiosity and laziness were the dominant reasons for AI use in our school. Our students would often set out to play a normal game of Go, but would get stuck on a particularly difficult or annoying move; eventually, their curious eyes would drift to their second monitor — where they usually had their AI software running anyway — and they would check the answer as one would sheepishly side-eye the solution to an interesting puzzle or homework problem. … This perspective of AI use to me explains why camera controls proved so effective against online cheating. Since AI use is usually an act of self-debasement and disempowerment – a subjection of oneself to ambient incentive gradients – it fundamentally contradicts the aesthetics of resourcefully overcoming a minor obstacle.
Hm, so I hadn’t actually touched it, and when I did it was kinda squishy, which made me think “maybe it did melt, and the gum just made it keep its shape”? Anyway, I put it back in and out of the fridge with the thermometer, and
So… it looks like it did freeze at 18? But then still doesn’t have an obvious point where it’s melting, which, ???
10:1: I had 200 ml water. Somewhere between 50-62.5 ml of Na2SO4 dissolved. Guessing at 55, that’s 275 ml / l. 3.75 ml xanthan gum made the water cloudy but didn’t noticeably thicken it, maybe previously I’ve had the water hotter when adding the gum?
The result was something that mostly froze in the fridge (some liquid remaining) and didn’t melt when I took it out.
(Occasional names and details checked with wikipedia)
Fall of Civilizations #21: Majapahit—Empire of the Islands
In 1811, the UK took Java from the Dutch. Thomas Raffles was made governor because he spoke Malay and had helped. In 1815 he did some exploring around the island. It was mostly Muslim at the time, but there were a bunch of ruins of an earlier civilization, with statues of Buddha and Hindu gods. He explored one place that the locals said was cursed. They warned him that he’d lose the governorship within a year, which turned out to be true because the Napoleonic wars ended and the UK gave Java back to the Dutch. Also he died of a stroke ten years later, at 44.
Most of the world’s volcanoes are in a horseshoe around the rim of the pacific plate, and that’s also where most earthquakes are. The Indonesian archipelago has a bunch of volcanoes and the highest density of islands in the world. New Guinea, Borneo and Sumatra are the world’s 2nd, 3rd and 5th largest islands, but their interiors are pretty inaccessible. [I think the podcast said 5th, but wikipedia puts Sumatra at #6.] Java is the fifth largest island in the archipelago, about the size of England, but today it’s the world’s most populous island, about 150 million people, and Jakarta is the world’s most populous city. It was known in antiquity, from at least the 2nd century CE, and appeared on lots of world maps. Not always in a consistent place, but always as far from anywhere you could get. [Presumably these were European world maps.] It had fertile soil, reliable rain, and being close to the equator the days were very consistent in length.
Most societies appeared in the west of Java, but one grew up centrally. At first they worshipped Shiva (for destruction—they were between two volcanoes). Later, Buddha too. They had a founding myth where a god sent off both Shiva and Buddha to separate parts of the world. They built what’s still the world’s second-largest Hindu [temple?], after Angkor Wat. They named their years after specific events or vibes, and they named that year (our 856) after the temple.
They also built a [separate?]… temple complex? The vibe they were going for was “the gods put down two mountains to hold the world together, and now we’ve put down a third”. Carvings on it tell us a bit about daily life. They also show five large ships.
Those ships were super impressive, and the Chinese didn’t let anyone foreign use at least one of their harbors partly because one of the Javanese ships could take like twenty of theirs. So the Javanese got a reputation as traiders and pirates. Some of them made it all the way to Madagascar—linguistic and DNA evidence puts them there (mtDNA says there were about 30 women). They might also have made it to the antarctic ocean.
They took volcanoes and earthquakes as portents. Once a prisoner was supposed to be executed, but every time they tried, Mount Merapi (the main volcano around) would complain at them, so eventually they just gave him a lordship. At some point, likely following a particularly bad eruption or series of eruptions, they had a time of strife with three kings in a year. When the dust settled they’d relocated to east Java. But before long they fell.
Some time later a new kingdom rose. The first king was actually from Bali. Stories about him king said that he was conceived immaculately, by a god who told the mother not to sleep with her husband any more or it would be bad for the baby [so I asssume they had already slept together]. The parents couldn’t care for the kid so they abandoned him, and he was taken in by a theif who saw him glowing. The thief taught him thievery. At some point a god or priest convinced him not to be a thief any more, and somehow this leads to him being king. Defeating some other local king was part of it too.
This kingdom becomes powerful partly because it controls the strait of Malacca, an important ocean passage (today 25% of shipping goes through it). It gets in a conflict with the Mongols in 1289, when the Mongols (specifically the Yuen Dynasty) send some messengers to ask for tribute and hostages, and the king (I think a descendant of the founding king?) has them mildly mutilated.
Kublai Khan doesn’t like that and sends a war fleet, but that takes some time. In the meantime, a subject kingdom revolts. They have two possible invasion routes, and send a force through the easy territory which the king’s son-in-law Raden Wijaya meets and defeats. Meanwhile they stealthily send a much larger force through the difficult territory, and reach the capital while the king is drunk as some important ceremony. A bunch of civilians abandon the city.
So the Raden Wijaya and the new king have a standoff, which is resolved by the Raden Wijaya being given some territory to build a new city, which they name Majapahit after some kind of bitter fruit. Around now the Mongols finally arrive in 1293, and don’t really know what to do because they can’t take revenge on a dead man.
Raden Wijaya agrees to be a vassal of the Mongols if they help him reclaim the kingdom, which they do, and then betrays them. Kublai wants to send an even bigger fleet but dies in the meantime and his empire gets distracted.
So now Raden Wijaya is king of Majapahit. He takes some steps to try to keep things stable, like marrying his father-in-law’s other daughters. But his heir Jayanegara is a crappy king, who only just manages to hold on to power thanks to a general of his, Gajah Mada. He keeps two princesses, half-sisters of his, locked up to avoid them getting married and having succession issues.
Eventually Gajah Mada and the princesses’ mother want him gone. He’s assassinated by a doctor, and Gajah Mada kills the doctor immediately after so who knows the real story. He’s succeeded by one of the princesses, with her mother having a lot of influence. They do a good job. At one point two generals are disagreeing about who should get to lead an army and the princess just leads it herself.
Powerful empire. Lots of money coming in through spices, especially nutmeg and clove. They have massive ships. At one point the Portugese capture one, but it takes a lot (their cannons aren’t powerful enough to puncture the hull below the waterline, but they eventually manage to break the rudder), and they use the captured ship as a warship themselves.
They tend to keep a light touch over their subjects, but they’re quite willing and able to use violence. At one point a subject tries to negotiate with the Chinese directly, and he has the Chinese ambassador beheaded. The Chinese emperor apologizes. I don’t remember what happens to the subject but it’s probably not fun. I think that same subject had also written himself first on a list of kings.
But eventually they have a succession crisis. A king has a daughter through a senior wife, who has a husband; and a son through a less-senior wife; and those two men have similarly strong claims to the throne. He tries to solve it by letting one rule the east and the other rule the west, but instead they decide to have a civil war.
During that war, one of the sides kills a bunch of men from the Ming treasure fleet that arrives. That king apologizes to the Ming emperor, who forgives it in exchange for remorse plus lots of gold. Eventually he sends like a tenth of the total amount of gold with apologies, saying he can’t afford the rest, and the emperor forgives him.
Eventually one side or the other wins, but they’re not really relevant as a power any more. When someone wants to take the throne as emperor of one of their subjects, he writes to the Chinese for permission, not to Majapahit.
Islam starts spreading in the area, partly because Muslims prefer trading with other Muslims. One particular city on Java gets a significant Muslim population, and then declares independence. Eventually they spread, and conquer Majapahit itself. There’s a story of… dubious veracity… that when the king of Majapahit sees their army outside his city he just gives up and dies. At any rate the city is abandoned.
The goalposts are shrouded, not moving
I still disagree, but I think I’m going to try to get an executable model and see if that helps us hash it out :)
London rationalish / Robin Hanson: How we broke humanity’s superpower and can’t see how to fix it
So it’s probably worth distinguishing a few questions:
-
Are there possible LTP setups where the STP is predicting “what’s the next override”?
Are there any of those that are biologically plausible? (Like, it’s not surprising if an LTP configured like this actually exists in the brain.)
Do any of them actually exist?
-
Are there possible LTP setups where the STP isn’t predicting “what’s the next override”?
Are there any of those that are biologically plausible?
Do any of them actually exist?
-
In any specific hypothetical example, is the STP predicting “what’s the next override”?
Is this example biologically plausible?
Does it actually exist?
I think for the (x.3) questions we’d need data we don’t have, so I’ll forget about them for the rest of the comment, they just seemed worth noting.
We both agree the answer to (1.1) is “yes”. In particular, this is what happens when the output of the STP doesn’t affect the next override. Like, if you’re on a rollercoaster, using an LTP to predict which direction to brace, and the overrides are sudden accelerations.
(1.2) I’m actually not sure about… not a confident “no”, but this setup feels to me like it would be kinda unlikely? Whereas I think you think these think are every biologically plausible LTP setup. This is just vague intuition though, and not cruxy for me.
(2.1) and (2.2) I think the answer is yes, and I think you think at least one of them is no.
So let’s look at some specific examples.
For the digestive enzymes, I actually think the binary model is already a biologically plausible example where the STP doesn’t predict the next override. But I’m happy to stick with the quantitative model too.
You’re brushing this aside, but to me it’s load-bearing. The R=1 output leads to more enzymes than R=0, and maybe even that small amount will wind up being too much. Probably not, but you only need that to happen 10% of the time.
I don’t think the numbers work out here. Compare R=1 to R=2. If these are predictions, then R=1 is a higher [predicted probability that the next override is R=0] than R=2. But R=1 produces fewer enzymes, which makes the R=0 override less likely. I think this is the case no matter what the specific curves look like, as long as they’re monotonic in the right direction.
It might be possible to come up with some way to use an LTP to handle digestive enzyme production, and have the STP predict the next override. But I feel like it doesn’t happen by default, and there’s no need for it. The job of the LTP is to handle digestive enzyme production; there’s no pressure towards “it should be possible to use the STP’s output to predict the next override”, so that doesn’t happen.
For the go-karting, I think the question is “how much do I steer?”
If I don’t steer at all in response to my predictions, then the STP can be straightforwardly correct.
If I steer a small amount, then maybe I’m still more likely to hit the side the STP predicts. Like, it thinks I’m 90% likely to hit the left side, and in fact I’m 70% more likely. We can still think of the STP as predicting the next override, but we need to recalibrate the function that reads an STP output and tells us a probability.
Ideally, I steer just the right amount, and I’m now equally likely to hit either side. The STP output is now uncorrelated with “which side do I hit”.
If I steer too much, I’m now more likely to hit the other side, and… ??? I’m not sure about the dynamics here, especially because in the time between “I oversteer” and “I actually hit the other side”, the STP is likely to be freaking out and (correctly) predicting that I’m now about to hit the other side, and I just don’t have time to react to that.
Connecting this to the digestive enzymes… both examples are going to have some kind of approximately-steady-state, like “producing a low-ish level of enzymes” or “keeping the steering wheel at a fixed angle while the track has a fixed curve”. From there, maybe there’s enough randomness that you might make too many enzymes or too few / hit either side of the track, but most likely that’s not going to happen, at least not much. (At least, we can’t assume that’s going to happen much. A simple model might have it not happening at all, and more complicated models can have it happening different amounts depending on parameters, and “it doesn’t happen much” is a reasonable setting for the parameters.)
At some point I’m going to sit down to dinner / the track is going to sharply veer left. (Definitely left, not right.) I know that’s going to happen at some point, and if my steady-state is steady enough, that means the most likely next override is “I’m not making enough enzymes” / “I’m going to hit the right wall”. But the STP isn’t going to be outputting “make more enzymes” or “turn harder left” yet, because it’s not coming up yet.
(I’m hoping to make an executable model we can play with, so if you still disagree then maybe the productive thing is to wait until I have that and that might help us figure out what’s up. Though it’s not definite I’ll succeed.)
-
MSE loss does not generate superposition
Indeed, no luck:
(It probably could have used more time in the freezer, but it’s clearly not freezing in the fridge.)
One hypothesis is that my water is too hard (269 ppm calcium carbonate) or otherwise impure. If so, possible avenues for fixing it are buying distilled water, buying regular bottled water, and using a water filter (quick google: apparently that doesn’t help with hardness).
Another hypothesis is that something about my salts is fucked. Impure sodium sulphate? The table salt comes with anti-caking agents, are those causing a problem? The table salt has also been just hanging out in my kitchen in this kind of thing, possibly with some dried beans in there to stop it clumping; maybe I should try some fresh salt?
I’ve done three experiments so far, and very roughly:
3:1 Na2SO4:NaCl by volume (most recent): I got 210 ml Na2SO4 in 1 l water (7x15 ml, in 500 ml)
4:1 Na2SO4:NaCl by volume (Nighthawk’s recipe): I got 160 ml Na2SO4 in 1 l water (400 ml, in 2.5 l)
No NaCl: I got 400 ml Na2SO4 in 1 l water (80 ml in 200 ml)
So there’s not even a clear trend for “more NaCl means less Na2SO4 dissolved”. Maybe because I did nighthawk’s at close to boiling? And/or because it went via the kettle?
Still, regardless of what my numbers show, more NaCl must mean less Na2SO4 dissolves, right? So maybe the next thing to try is, like, 10:1?

It was mentioned (in the context of ~”now that AI is doing all the coding, is Haskell still what we want to be using?”), but not much more than that in my hearing.
(I’m assuming here that it was the thing I’m thinking of that I saw on the subreddit and discourse.)