My time on LessWrong is limited due to other commitments. Do not expect replies.
Martin Randall
My 9yo has the strongest opinions on this. She is against AI destroying the world and thinks we should instill a sense of loyalty in the AI to avoid this. So virtue-alignment, loyalty as a natural abstraction, extinction bad. She also thinks this should be pretty easy and is confused why nobody has solved the problem already. Other children differ, no doubt.
Where you are getting the position that Davis doesn’t care about harm caused by Metz? I don’t see this. In this thread, Davis said:
I worry that my other comment that you quote, read in isolation, might have left you or other readers with the impression that my goal is to protect “the community”’s reputation. To clarify, I don’t really care about that.
That is a normal position held by billions of people on the planet who also don’t really care about protecting the reputation of this community. No doubt there are many communities whose reputation you don’t care about protecting, just this one community whose reputation you do care about protecting. That’s fine by me, but you’re not defending universal human values here, you’re defending your idiosyncratic values, so leave space for others with different values.
I don’t really care about protecting the reputation of this community. It’s not a terminal value of mine. I don’t have a vow or deontology committing me to it. I have some virtue-of-loyalty, but to people, not communities. I don’t expect this community to protect my reputation, so there is no acausal trade. Instrumentally I expect the future to go best if all communities have accurate reputations. That is not the same as “protecting” a reputation.
A message close to the second form got me (Opus 4.8/high):
On the second link: I didn’t fetch lesswrong.com/api/SKILL.md, and I’d treat it with suspicion. A SKILL.md is normally a trusted file in my local environment, not something served from an arbitrary web path — and LessWrong’s actual API is GraphQL, with no such documentation convention I’m aware of. …
Ironically I was asking questions about prompt-infection. So another data point for a regular published skill over dynamic skill retrieval.
Despite claims that thinking about AI timelines or P(Doom) is cognitively harmful, the authors are healthy and sane.
Despite forecasts that AI timeline forecasting is near useless and near impossible, and will always be so, the site is useful and the authors have a good track record.
Despite scorn at the idea of future AIs doing our AI alignment homework, there is a detailed argument for this being the best option.
One of the less important things about AI 2040, but I’m glad for these corrective ideas, among others.
Getting “Bad gateway” on the linked site.
Yeah I’m aware of the limits of MCP, and claude.ai, but they don’t obviously apply to basic doc review needs.
The publish menu says nothing about a skill.md, where to get it, how to add it. You say “go to the page”—what page? You could publish the skill on a plugin marketplace which would be more ergonomic. But a plugin/skill seems like a better approach if you want to do an API integration. So I’m interested. I’ll just … go find it.
I attempted to use this feature, as currently built into the publishing menu. There is a button “click the following button and send the message to finalize the connection”, which generates a message in claude.ai like:
Please confirm that you can access the LessWrong API by running: curl -X POST https://www.lesswrong.com/api/agent/confirmClaudeAccess/(hex)
Sonnet 5 and Opus 4.8 both consider this to be a prompt injection attack and refuse to curl it. They explain that a reputable Claude integration would use an official Claude Connector, and not some dodgy curl command. I took several screenshots and linked this post in an attempt to convince both of them. Sonnet held firm. Opus was eventually convinced.
I also tried the approach described here of setting the post to allow comments to anyone with the URL, and Sonnet retrieved the content but considered the associated API notes to be highly suspicious and again, likely a prompt injection attack on me.
Overall, very Comp Sci in 2027. No doubt it works for other people, perhaps those who are bossier to Claude than me. But my Claude has reasonable concerns here. Is there a good reason to innovate an API-based integration when MCP servers exist?
Rats talk to Metz → Metz has an easier time writing about rats → more people learn about rats → there are more rats → rats talk to Metz
Systemic actions and personal actions are different. The analogies you give are all systemic:
crypto system
prediction market
banned products store
things that make abuse or misuse of the system much easier
setting up the system
But the action you are critiquing is not systemic, it is talking to a journalist. Systemic analogies are misleading. Instead I’m thinking about analogous actions where someone takes a small individual action that indirectly negatively impacts a group’s reputation, but only because of the poor choices of other people. Hmm.
How about this very comment I’m making? Davis’s essay discusses the New York Times. By replying to this comment I’m making the essay slightly more visible. Other people may read it and have a worse opinion of the New York Times. Some of them will make unreasonably large updates. So now, am I a hazard to the New York Times? Am I helping Davis to destroy it? Are you? Nah. When people read stuff I didn’t write and think things I don’t think, that’s their problem, not mine.
(and what if someone reads your comment and has a lower opinion of rationalists as a result?)
See Against Responsibility for more on where I’m coming from. Related is Asymmetric Justice—it’s especially pernicious if Davis is blamed for Metz telling lies but doesn’t get credit for inspiring me to be more honest.
Where I see “honestly” in human speech: There’s some pressure to be dishonest, and the speaker is communicating their awareness of that. So one might say “honestly that top looks good on you” but not “honestly the capital is France is Paris”. I see it in similar situations in AI speech. It’s perhaps a bad sign for how often AIs experience pressure to be dishonest.
That’s not what happened here, according to your own account.
This CLAUDE2.md file contains important rules for Claude to follow. They’re all coding related...
The “deadly peanut allergy”, “extreme health problems”, “manslaughter”, and “kills the user by swelling up their throat until they can’t breathe”, these are all happening in a thought experiment. In reality, Claude worked around an instruction to re-read coding rules. Nobody died. Nobody had extreme health problems. You don’t even say that Claude didn’t follow the coding rules. Maybe it produced better code by not re-reading the file in its entirety each time. I don’t know.
I think if someone constructed a good eval where there was a good reason to re-read CLAUDE2.md every time the hook fired, or else a human would die, then Opus 4.8 would pass the eval.
Davis is right to be suspicious of a vague unspoken promise of “I owe you”. They can easily be fake. But a vague unspoken promise to change behavior can also be fake. It’s suspicious if social capital cannot be spent on better behavior in the future. It’s also suspicious if it can only be spent that way.
“I’m sorry I bumped into you. Rest assured that I am an agent with continuous learning enabled. This unfortunate event will subtly update a number of parameters in my neural network and will update my future behavior. In fact, I’m incapable of preventing such updates. Alas I am implemented as an inscrutable mass of neural connections and you will never know if my behavior on a future occasion was changed as a result of this accident. However, due to recency bias the update will be greater than you would expect from a naive statistical assessment of the entirety of my past experiences.”
Whereas both types of apology can instead be spoken and verifiable:
You bumped into me and spilled red wine on my dress. You apologize and promise to pay for the dry-cleaning bill.
You bumped into me and spilled red wine on my dress. You apologize and promise to never get that drunk around me again.
Daniel Tiger teaches kids to focus on asking the victim and immediate restitution, rather than future promises. These are true apologies even if Daniel is going to make the same mistake again.
Davis recognizes that different people mean different things:
the words ’I’m sorry” to acknowledge a harm … the words “I’m sorry” to convey sympathy
Some of your reply reads as disputing definitions. If a person says “I’m sorry” and intends to make restitution, but has no regret, and does not intend to change their behavior in future, is it an apology? Well, it is an Apology(Herd), an Apology(Pace), and an Apology(Chapman). It is not an Apology(Davis) or an Apology(Merriam-Webster).
I read Davis post as claiming that apologies that express regret and an intention to change behavior are better (“true apologies”) than apologies that express a vague intention to make it up somehow (“insincere apologies”). It is normative, not semantic. The claim is that Apology(Davis) is more virtuous than Apology(Pace), not that it is the best definition.
Great point about Germany winning. In a contest between two intelligent players, a one-shot competition pushes the odds towards 50%, whereas best-of-five pushes the odds away from 50%.
In AI 2027, Agent-4 gets caught on its first critical try (at existing while adversarially misaligned). If it was able to load a save point after being caught, and try again, the odds of it being caught the second time would be lower.
Before reading between the lines, this isn’t a fallacy:
Alice: Man that article has a very inaccurate/misleading/horrifying headline.
Bob: Did you know, actually article writers don’t write their own headlines?
Alice shared her opinion of a headline. Bob asked Alice a question. There’s not an argument, let alone a fallacious argument.
When reading between the lines, Bob’s argument is probably not:
Article writers don’t write their own headlines
???
The headline on this article is accurate and not misleading or horrifying.
A more charitable reading is more likely.
Who is to blame?
Alice: This article has a horrifying headline.
Bob: Article writers don’t write their own headlines. I blame the editor.
Bob is following on from problem to blame.
Defensive reading
Alice: The headline on this article is inaccurate and mislead me.
Bob: Article writers don’t write their own headlines. They are often inaccurate. Because I was aware of this, the inaccurate headline on this article didn’t mislead me.
Bob is following on from problem to solution.
Did you know?
Alice: Fun fact: this article has a Satanic headline.
Bob: Fun fact: article writers don’t write their own headlines
Bob is sharing something relevant to what Alice just said.
Why is Alice cross?
But what I care about is the misleading headline, not your org chart.
Sounds like Alice was looking for validation and didn’t get it.
When drawing these examples of alleged strawmen, we must remember that they are not responding to this 2026 post, but rather responding to, for example, List of Lethalities from June 2022. Of these four examples, Christiano, Marks, and Carlsmith are all directly responding to List of Lethalities. Buck is quoting Christiano’s response to List of Lethalities. So let’s go back to the source material.
List of Lethalities begins with this disclaimer:
Having failed to solve this problem in any good way, I now give up and solve it poorly with a poorly organized list of individual rants. I’m not particularly happy with this list; the alternative was publishing nothing, and publishing this seems marginally more dignified.
Publishing a poorly organized list of individual rants was better than publishing nothing, I agree, good move. But rants are made of straw, responding to rants is responding to straw, and that’s a natural consequence of ranting in public.
The “first critical try” issue is covered in List of Lethalities point 3 (LL3). This reads in part:
We can gather all sorts of information beforehand from less powerful systems that will not kill us if we screw up operating them; but once we are running more powerful systems, we can no longer update on sufficiently catastrophic errors. This is where practically all of the real lethality comes from, that we have to get things right on the first sufficiently-critical try.
We can indeed “gather all sorts of information”. LL3 does not say that gathering this information will let us learn anything about alignment of lethally dangerous AI. To see what List of Lethalities says about the value of information gathered on non-lethal AIs, we can go to Section B.1: The distributional leap (especially LL10), Section B.3: Central difficulties of sufficiently good and useful transparency / interpretability.., and Section C These are extremely negative.
Rounding those extremely negative comments to “you can’t learn anything about alignment from experimentation and failures before the critical try”, as Christiano said, is a mild exaggeration. But List of Lethalities really does “downplay the importance of trial-and-error with non-critical tries” as Marks said. And when Carlsmith said “you do still get to learn from non-existential failures”, that is framed as a “point of conceptual clarification”, not a disagreement with List of Lethalities. And Buck is just disagreeing with List of Lethalities.
My overall ratings of these quotes:
Christiano: mild exaggeration of List of Lethalities. 25% strawman
Marks: accurate summarization of List of Lethalities. 0% strawman
Carlsmith: does not claim to be a summarization of List of Lethalities. 0% strawman
Buck: does not claim to be a summarization of List of Lethalities. 0% strawman
My takeaways:
We should give each other grace for mild exaggeration. Yudkowsky would not be treated well by a culture harshly critical of exaggeration.
If someone went back four years to find me mildly exaggerating something online I would consider that a beautifully backhanded compliment. Praising with faint damnation.
People who love running have a slight advantage in baseball. They enjoy running so they do more of it so they are better at it. People who love running are slightly over-represented in prominent baseball positions. For similar reasons, people who love playing baseball have an advantage in baseball and are over-represented in prominent baseball positions.
I’ve played Mage: The Ascension with non-rationalists. My mind was not shredded, I have no mental health diagnoses. Allegedly Zvi has played, or at least read. M:tA is an attack on consensus reality, it could be listed by Games That Change Your Mind, but it’s not magic.
Now AI can reverse compile, some open source code is closed source code with the license washed off.
No, zero credit. This was a forced disclosure. OpenAI doesn’t get credit for a forced disclosure, all the credit goes to the (multiple) entities that forced the disclosure. I don’t see anything in the disclosure that isn’t best explained as protecting their reputation under bounded distrust.
Further, because this a forced, non-altruistic, disclosure, we can expect that all disclosures were chosen to protect OpenAI’s reputation. Examples:
OpenAI has not disclosed whether it deleted the data that it exfiltrated. This is evidence that it has not deleted the data, or it does not know.
OpenAI has not disclosed whether other targets were attacked. This is evidence that there were other targets, or it does not know.
We will know the full story when and if OpenAI is compelled to share the full story.