Steff
The bias still making some experts underestimate LLMs
Do your readers get annoyed if you post every day for 30 days?
Thanks for this reminder. I myself am someone who’s prone to this sort of thing, so I don’t recognize it in others’ writings (I just sort of mentally snicker, amused at the comeuppance) without someone else pointing it out.
A challenge: Can you make an LLM follow these instructions?
>> Indeed, I find that the magic words “ruthlessly maximize” are a very good brainstorming aid. For example, when you hear “The reward function triggers when the supervisor presses the ‘approve’ button”, then your mind might default to nice, friendly reward-increasing strategies, like being helpful. Whereas when you hear “Ruthlessly maximize how much the supervisor presses the ‘approve’ button”, then you’ll notice that the space of possible reward-increasing strategies is actually quite wide, and includes strategies like kidnapping the supervisor’s children to force the supervisor to press the button nonstop.
Just popping in to say that I love this bit. I mean, I hate the danger it represents, but I enjoy how effectively you described it.
The one name LLMs may fear
I think even with the most pessimistic understanding of the changes that the Superhappies would enforce, the True Ending is still vastly more evil.
Is it weird that I’d fear being changed to eat-my-own-non-sentient-babies as a form of death, the end of my existence, but I wouldn’t fear the change to not-feel-pains in the same way? Maybe that’s irrational.
However, something I don’t feel conflicted about is this: I loathe those who would kill themselves in this situation. You’re about to die anyway; suicide in this case is actually homicide.
Claude’s malicious compliance and normalization of deviance
X-risk is less viral than political tribal fear
The most striking part of the essay for me was this:
>> By making the child choose — and making them choose a specific answer — the PPP more fundamentally subverts the child’s autonomy. And crucially, they do so in a way that obscures the power structure. With the strongman father, the child can say “why did you make me do this?” or perhaps “what duty do I have and why?”. The child can say “I am not having fun” and the strongman father can say “that is unfortunate.” But the PPP makes the child self-inflict the wounds, while occluding the wider context of the imposition.
I’d never thought of corrigibility from this perspective—how difficult it can be to align our own children. I feel like there’s more here to mine, more parallels to make that might give us insights into how to better write LLM constitutions—but I’m not sure what direction to take it.
>> If coordination is possible, why did they choose to shut down unilaterally?
It makes sense to me: One company self-sacrificing in order to create a Schelling Point that all other companies could rally on.
Though I suppose it doesn’t matter what makes sense to me. Rather, what matters is whether this makes sense to the (very profit driven) board members for these companies.
What I want to see,
which I predict is maybe more realistic(EDIT: Thinking on it more, I’ve changed my prediction, I think convincing one frontier company to shut down is actually more likely!) --
US and China buying equal stakes in each others’ companies. We sell half of Anthropic and OpenAI and a spun-off Deepmind. They sell half of DeepSeek and spin offs of Alibaba and Baidu.
We work together for a slow down, gesture towards great partnership and profit sharing and economic growth between the two countries yadda yadda, and make a good example for the rest of the world so India et al will also join in.
Over many repetitions (to balance out bad luck streaks), it’s rational to play the game for prices of $20 or $19. If you’re saying that SB should reject those prices when presented it while waking, that’s irrational, because an SB that accepts those prices will make money.
This is assuming that if asked twice, she’ll always make the same decision because her decision-making is deterministic. (We could assume there’s a chance she makes a different decision on Tuesday than on Monday but I think that’s an unnecessary and irrelevant complication?)
If he cancels the invasion when Beauty is sleeping, then that weighs the probability towards worlds where she’s not asleep, i.e. Tails. Which is why Heads goes down from 3⁄4 to 2⁄3.
(Not sure if this is identical to the $30/6/6 EV form of the questions? Not sure which is more useful)
I take back one piece, when I said $18 doesn’t correspond to anything. I believe that corresponds to how much Beauty should expect to be making per day on average: (1/2)(30) + (1/2)(6).
If we’re talking about expectation per iteration, though, then it should be (1/2)(30) + (1/2)(12) = $21. If you’re using the $18 value there, then you can get money pumped.
I’m going to be sort of shitty and not give this comment the time it deserves, and I’ll only respond to parts of it. If there’s anything you’d like to focus on that I ignore, please hit me up with another comment and I’ll try to get to it when I can.
First, I’ll say that some of the confusion probably stems from the fact that the post is defending against an argument against Halferism, so I’m presenting a Halfer point and showing it to not result in any contradictions, not justifying the Halfer point directly. You being a Thirder, you wouldn’t think that P(obs Mon) = 3⁄4, since that’s not the Thirder position.
Secondly, I’d like to note my complete agreement with the statement “SB’s answer really depends on what she believes the Riddler used as an heuristic for invading the experiment”.
>> I notice that you seem to be ignoring that point (that your versions of SB problem that have 1⁄2 as a solution work exactly the same without the whole sleep and amnesia thing)
I’m not sure what you’re saying with this piece, but my guess-of-an-answer-to-what-you-might-be-saying is this: The solution does work exactly the same as asking about the coin without any sleep or amnesia, because it should be the same. That’s part and parcel to the whole Halfer position that in waking, Beauty hasn’t learned anything, and without learning anything, there’s nothing to update.
(Analogously: There’s no need to talk about “self locating” or “self indicating” in anthropics, because being-a-self gives no additional information on top of “there exists 1 of the thing that is me”.)
>> The proportion of time she would spend awake on Monday during many repetitions of the experiment is 2⁄3. She should answer 2⁄3. After finishing writing this reply, I think there’s a chance that the crux of our disagreement is here.
Damn, that’s my mistake. 2⁄3 of the time is spent awake on Monday, that’s true. I can see why that would seem to indicate the Thirder position. Probably most helpful would be to focus on my Riddler example, with the added condition that the Riddler wasn’t doing either of the actions you described. Instead, he didn’t know about Beauty at all until he happened to stumble upon her. (I’ve fixed the post up a bit.)
I’m guessing you’ll say that in this version, Beauty’s belief in Monday should not be 3/4?
>> What if the ASI itself protects liberal democracy?
>> Liberal democracy is not the true target
I think the only feasible solution is one and the same: ASI protects the government, and the government is good government because it replaces liberal democracy because AI BECOMES the new government. To some extent.
You could imagine a world where China still has Xi and America still has a president, but their actions are all constrained by what the ASI will allow.
Then of course, the greatest political matter of the day will be: Who gets to influence the ASI? Either corrigibly, or if the ASI can’t be influenced, then this question is reserved for the period of time leading up to the ASI’s escape-from-our-influence.
I’m thinking there’s a legit chance that the future will come down to one man (Asmodei) and all of our efforts should be focused on trying to teach him and encourage him to be able to make the best choices he can make.
I believe what you’ve described is just the original problem again with an extra possible waking of no question asked on Tuesday, but that doesn’t impact the credences of her waking and observing the question. The answer is still 1⁄2.
Agreed with the argument, however, I know I’m biased towards agreeing because the idea of WBEs terrified me because of Amazing Digital Circus style scenarios.