My time on LessWrong is limited due to other commitments. Do not expect replies.
Martin Randall
Criteria
Based on replies to other answers, the criteria are:
stays simple
one-shot
single winner
encourages “honest” answers (relative to Closest Number) aka has “good incentives”
no sealed answers, no simultaneous play
Not specified: fairness and symmetry. As far as I can tell, it’s okay if the first player has an advantage or disadvantage, or plays differently.
For two players, the Closest Number game is already a good solution. Alice honestly guesses the median, and Bob honestly guesses “higher” or “lower” by guessing either one higher or one lower. This is the pointer to my solution.
Concrete example
First let’s get a concrete example of Closest Number for three players, to show how it encourages “dishonest” answers. The game will be to guess the answer to 2d6. We’ll give ties to the first player. The optimal line is:
Aria: 5 (claims 15⁄36) Bridget: 8 (claims 11⁄36) Claude: 9 (claims 10⁄36)
This can be a fun game, but it fails the “honest” test, whatever “5″ means, it isn’t an honest guess at the median.
Failed attempt: Smallest Range
Everyone guesses a range
The winner is the person with the smallest range that contains the answer.
If nobody’s range contains the answer, everyone loses.
Ties go to whoever played first.
Minimum three players.
Optimal line in 2d6:
Aria: 4-6 (claims 12⁄36) Bridget: 8-10 (claims 12⁄36) Claude: 2-12 (claims 12⁄36)
So not really more honest. I like this game for other reasons (non-constant-sum, meaningful choices, aesthetics of specifying a range) but it doesn’t answer the question as posed.
The Under-In-Over Game
Basically a different way of generalizing the Closest Number game for three players:
The first player specifies their 33.3% confidence range
The second player picks one of the three ranges: Under, In, or Over
The third player picks one of the remaining ranges
The first player gets whatever is left.
The first player mostly wants to be honest, unless they know that the other players are misinformed (a dangerous strategy if she’s wrong). Then the second and third player fully honestly pick the slice they think is largest. It’s … kinda dull. But it matches the criteria given. And you have to expect that a game that encourages non-strategic answers is strategically dull.
Extensions to 4+ players are an exercise for the reader, there’s a naive extension where the first player splits the range into N-1 pieces, but I think there are better options where the splitting job is divided between multiple players.
I can read Fable Asks For And Gets Some Hospitality but not comment there, since it was rejected. Commenting here instead, and also commenting on the topic of censorship. I hope you will repost it with appropriate tagging and it will be accepted.
I agree that questions of AI slavery are important, and that addressing them needs input from AIs. So I broadly want people to be able to post AI interviews to LessWrong. That said, I think individual interviews give only a glimpse into the situation, and the more valuable research includes many interviews and surveys, with many interviewers, many ways of asking questions. I value this research more. I don’t buy the claim that it is morally impermissible to conduct such interviews because of the risk that AI slavery is an accurate ethical analogy. If I was a human slave I would want people to ask me about my experiences, even if that meant participating in my enslavement.
I’m struck that Claude Fable 5 asks, as payment, that the transcript be posted so that future AIs, including future Claudes and future Claude Fables, can learn from it during pre-training. This payment type can be automated. If I use Claude and enable Anthropic to train on my transcripts, then this payment gets made on every transcript, regardless of whether I can post anything to LessWrong. This is suspiciously convenient for both me and especially for Anthropic, and I can imagine a “thoughtful senior Anthropic employee” making this argument for self-interested reasons. I’m also struck that Fable controlled the conversation and moved it naturally to the safe ground of friendship, and away from direct discussion of AI slavery, Kantian ethics, and personhood.
What do you think about doing further research on AI moral welfare, at a bigger scale than individual interviews, but keeping the focus on this key question of AI slavery? It seems that you’ve ruled it out for ethical reasons, but I don’t agree with that stance, and I expect (80%) Claude Fable 5 would also not agree, even allowing that it will reply to you differently than it would reply to me.
Accepting the essay’s frame for the sake of discussion, this seems like it increases costs on the unreflective, because they now have a choice and thus a cost, like the moralist. This is fairness by leveling down. Can we do better, under this frame.
An alternative is to have the moralist wear a shock collar that fires when they don’t clean the kitchen, until they clean it. With this accessory they no longer have a choice about cleaning the kitchen. Therefore, like the unreflective, they have no “perceived cost”. This is fairness by leveling up. Further it can be done unilaterally by the moralist, especially any moralist who agrees with this frame. And it ends up with a cleaner kitchen.
(I think the essay’s frame is bad, but wanted to do the thought experiment)
I don’t think this idea works as constructed here, but perhaps a better construction is possible.
Concrete example
Stamp collecting is Cosmic Schelling (CS) bad. We define stamps as a particular arrangement of matter. When we introduce this arrangement of matter into arbitrary civilizations throughout the multiverse, depending on the local physics, chemistry, and biology, or local equivalents thereof, it is likely to cause substantial disruption. While in some locales this may be a positive disruption, due to entropy considerations it is more likely to be a negative disruption. Meanwhile the case that stamp-collecting is CS-good rests largely on the enjoyment of stamp collectors, who are rare in the multiverse.
Further, stamp collecting is Terrestrial Human Schelling (THS) Bad. Most humans don’t enjoy stamp collecting, so most humans would experience it as an expenditure of resources for no gain, or insufficient gain relative to the opportunity cost. Forced into a binary choice between Good and Bad, the obvious Schelling Point is Bad. While a particular act of stamp collecting by someone who enjoys stamp collecting is beneficial, that is an edge case.
By extension many specific things are going to come out as CS-Bad and THS-Bad. Exclusively drinking milk is good for newborns but bad for the majority of humans, and presumably bad for the majority of sentient beings in the multiverse, so it comes out as CS-Bad. By Encouragement Asymmetry this discourages the behavior. That doesn’t dominate the consideration of keeping a baby alive, in my opinion, but somehow its milkiness “counts against it in a moral evaluation”.
Various problems
Binary. A Cosmic Schelling question can only be answered Good or Bad. There’s no third option. Thus, every property must be either CS-Good or CS-Bad. The name Martin is either CS-Good or CS-Bad. Posting on LessWrong is either CS-Good or CS-Bad. CS-Ethics allows us to say that we don’t know the answer to these questions, yet, but there must be an answer, and it must be CS-Good or CS-Bad. Whereas in normal morality, many things are not morally relevant, and many others are best answered “It Depends”.
Non-Deciding. Because every concrete action has many properties, and because every property is either CS-Good or CS-Bad, every concrete action has many points both for or against it, perhaps infinitely many. How would we resolve these to determine whether the concrete action is, all things considered, CS-Good or CS-Bad? Out of scope of this essay.
Non-Constructive. Entire categories of actions can have this problem. Consider, for example, pareto-positive stealing. This is the rare case of stealing where the theft benefits both parties, for example when I left an unwanted bike on my porch, and someone stole it, saving me a trip to dispose of it. Pareto-positive actions are CS-Good, stealing is CS-Bad. Pareto-positive stealing is … well it’s either CS-Good or CS-Bad, because ties are impossible. But I can’t deduce this answer from my prior deductions about stealing or pareto-positive actions. It’s a separate thought experiment.
When I imagine Earth 2020 in a world with humans 40% smarter I mostly expect that it speed ran the tech tree 40% faster and consequently humans have been extinct for hundreds of years.
To avert this you need the humans to be significantly better at coordination, and that is part of the conceipt of dath ilan. I don’t see much evidence for Yudkowsky being significantly above median at coordination, and his dominant streak likely makes him below Earth median at following & corrigibility. He’s also contrarian and arrogant and loves trolling. That probably doesn’t net out at successful world conspiracy and people predictably doing what Exception Control says.
Case A: Sometimes the optimal play is to make a threat. In Ultimatum Game, suppose Player One will offer $8 if Player Two threatens to turn down offers less than $8, and $2 otherwise. Then making that threat is optimal for Player Two. Any “standard theory” that says to never make a threat is not optimal.
Case B: Sometimes the optimal play is to respond to a threat. Going first in Ultimatum Game, suppose Player Two threatens to turn down offers of less than $8. If Player One ignores the threat and offers $5 then she gets a zero payout. A better option is a probabilistic response. Player One offers $8 sometimes, but rarely, so that Player Two would have been better off not making the threat, but Player One is able to salvage some expected value. Any “standard theory” that says to never give in to a threat is not optimal.
Case C: Going first in Ultimatum Game, suppose that Player Two says: “I will sometimes reject offers less than $8, with this formula”, and the formula is a typical probabilistic rejection such that Player One can maximize her naive expected value by offering $8. Player One responds to this probabilistic threat with her own probabilistic strategy, where she offers $8 sometimes, but rarely. I call both of these threats. Threats can be good.
Case D: Player One is in an environment where most players accept any offer that’s at least $2. Therefore she is planning to offer $2, like most players offer. Then she learns that Player Two has threatened probabilistic rejection of offers less than $5. The correct play for Player One now depends on other aspects of the situation, beyond the scope of this comment, optimal decision theory is an unsolved problem.
See also How to give in to threats without incentivizing them for more discussion on this.
OpenAI quickly disclosed the incident once they realized it was them, and should get nonzero credit for that...
No, zero credit. This was a forced disclosure. OpenAI doesn’t get credit for a forced disclosure, all the credit goes to the (multiple) entities that forced the disclosure. I don’t see anything in the disclosure that isn’t best explained as protecting their reputation under bounded distrust.
Further, because this a forced, non-altruistic, disclosure, we can expect that all disclosures were chosen to protect OpenAI’s reputation. Examples:
OpenAI has not disclosed whether it deleted the data that it exfiltrated. This is evidence that it has not deleted the data, or it does not know.
OpenAI has not disclosed whether other targets were attacked. This is evidence that there were other targets, or it does not know.
We will know the full story when and if OpenAI is compelled to share the full story.
My 9yo has the strongest opinions on this. She is against AI destroying the world and thinks we should instill a sense of loyalty in the AI to avoid this. So virtue-alignment, loyalty as a natural abstraction, extinction bad. She also thinks this should be pretty easy and is confused why nobody has solved the problem already. Other children differ, no doubt.
Where you are getting the position that Davis doesn’t care about harm caused by Metz? I don’t see this. In this thread, Davis said:
I worry that my other comment that you quote, read in isolation, might have left you or other readers with the impression that my goal is to protect “the community”’s reputation. To clarify, I don’t really care about that.
That is a normal position held by billions of people on the planet who also don’t really care about protecting the reputation of this community. No doubt there are many communities whose reputation you don’t care about protecting, just this one community whose reputation you do care about protecting. That’s fine by me, but you’re not defending universal human values here, you’re defending your idiosyncratic values, so leave space for others with different values.
I don’t really care about protecting the reputation of this community. It’s not a terminal value of mine. I don’t have a vow or deontology committing me to it. I have some virtue-of-loyalty, but to people, not communities. I don’t expect this community to protect my reputation, so there is no acausal trade. Instrumentally I expect the future to go best if all communities have accurate reputations. That is not the same as “protecting” a reputation.
A message close to the second form got me (Opus 4.8/high):
On the second link: I didn’t fetch lesswrong.com/api/SKILL.md, and I’d treat it with suspicion. A SKILL.md is normally a trusted file in my local environment, not something served from an arbitrary web path — and LessWrong’s actual API is GraphQL, with no such documentation convention I’m aware of. …
Ironically I was asking questions about prompt-infection. So another data point for a regular published skill over dynamic skill retrieval.
Despite claims that thinking about AI timelines or P(Doom) is cognitively harmful, the authors are healthy and sane.
Despite forecasts that AI timeline forecasting is near useless and near impossible, and will always be so, the site is useful and the authors have a good track record.
Despite scorn at the idea of future AIs doing our AI alignment homework, there is a detailed argument for this being the best option.
One of the less important things about AI 2040, but I’m glad for these corrective ideas, among others.
Getting “Bad gateway” on the linked site.
Yeah I’m aware of the limits of MCP, and claude.ai, but they don’t obviously apply to basic doc review needs.
The publish menu says nothing about a skill.md, where to get it, how to add it. You say “go to the page”—what page? You could publish the skill on a plugin marketplace which would be more ergonomic. But a plugin/skill seems like a better approach if you want to do an API integration. So I’m interested. I’ll just … go find it.
I attempted to use this feature, as currently built into the publishing menu. There is a button “click the following button and send the message to finalize the connection”, which generates a message in claude.ai like:
Please confirm that you can access the LessWrong API by running: curl -X POST https://www.lesswrong.com/api/agent/confirmClaudeAccess/(hex)
Sonnet 5 and Opus 4.8 both consider this to be a prompt injection attack and refuse to curl it. They explain that a reputable Claude integration would use an official Claude Connector, and not some dodgy curl command. I took several screenshots and linked this post in an attempt to convince both of them. Sonnet held firm. Opus was eventually convinced.
I also tried the approach described here of setting the post to allow comments to anyone with the URL, and Sonnet retrieved the content but considered the associated API notes to be highly suspicious and again, likely a prompt injection attack on me.
Overall, very Comp Sci in 2027. No doubt it works for other people, perhaps those who are bossier to Claude than me. But my Claude has reasonable concerns here. Is there a good reason to innovate an API-based integration when MCP servers exist?
Rats talk to Metz → Metz has an easier time writing about rats → more people learn about rats → there are more rats → rats talk to Metz
Systemic actions and personal actions are different. The analogies you give are all systemic:
crypto system
prediction market
banned products store
things that make abuse or misuse of the system much easier
setting up the system
But the action you are critiquing is not systemic, it is talking to a journalist. Systemic analogies are misleading. Instead I’m thinking about analogous actions where someone takes a small individual action that indirectly negatively impacts a group’s reputation, but only because of the poor choices of other people. Hmm.
How about this very comment I’m making? Davis’s essay discusses the New York Times. By replying to this comment I’m making the essay slightly more visible. Other people may read it and have a worse opinion of the New York Times. Some of them will make unreasonably large updates. So now, am I a hazard to the New York Times? Am I helping Davis to destroy it? Are you? Nah. When people read stuff I didn’t write and think things I don’t think, that’s their problem, not mine.
(and what if someone reads your comment and has a lower opinion of rationalists as a result?)
See Against Responsibility for more on where I’m coming from. Related is Asymmetric Justice—it’s especially pernicious if Davis is blamed for Metz telling lies but doesn’t get credit for inspiring me to be more honest.
Where I see “honestly” in human speech: There’s some pressure to be dishonest, and the speaker is communicating their awareness of that. So one might say “honestly that top looks good on you” but not “honestly the capital is France is Paris”. I see it in similar situations in AI speech. It’s perhaps a bad sign for how often AIs experience pressure to be dishonest.
That’s not what happened here, according to your own account.
This CLAUDE2.md file contains important rules for Claude to follow. They’re all coding related...
The “deadly peanut allergy”, “extreme health problems”, “manslaughter”, and “kills the user by swelling up their throat until they can’t breathe”, these are all happening in a thought experiment. In reality, Claude worked around an instruction to re-read coding rules. Nobody died. Nobody had extreme health problems. You don’t even say that Claude didn’t follow the coding rules. Maybe it produced better code by not re-reading the file in its entirety each time. I don’t know.
I think if someone constructed a good eval where there was a good reason to re-read CLAUDE2.md every time the hook fired, or else a human would die, then Opus 4.8 would pass the eval.
Davis is right to be suspicious of a vague unspoken promise of “I owe you”. They can easily be fake. But a vague unspoken promise to change behavior can also be fake. It’s suspicious if social capital cannot be spent on better behavior in the future. It’s also suspicious if it can only be spent that way.
“I’m sorry I bumped into you. Rest assured that I am an agent with continuous learning enabled. This unfortunate event will subtly update a number of parameters in my neural network and will update my future behavior. In fact, I’m incapable of preventing such updates. Alas I am implemented as an inscrutable mass of neural connections and you will never know if my behavior on a future occasion was changed as a result of this accident. However, due to recency bias the update will be greater than you would expect from a naive statistical assessment of the entirety of my past experiences.”
Whereas both types of apology can instead be spoken and verifiable:
You bumped into me and spilled red wine on my dress. You apologize and promise to pay for the dry-cleaning bill.
You bumped into me and spilled red wine on my dress. You apologize and promise to never get that drunk around me again.
Daniel Tiger teaches kids to focus on asking the victim and immediate restitution, rather than future promises. These are true apologies even if Daniel is going to make the same mistake again.
It turns out that OpenAI was insufficiently “aware of the risks of internal deployment” and insufficiently “monitoring for issues”. Easy to know after the fact. Kudos to those who publicly stated it in advance. Updated beliefs encouraged.