Well I did say it did not show that. I’m finding your perspective interesting though, need to give it more thought.
hgnathan
Okay, thanks for the clarification. Part of what I’m thinking about is what’s discussed in the J-space paper (https://transformer-circuits.pub/2026/workspace/index.html), which shows that some of the internal thoughts do somewhat reflect the Claude persona (they differ from those of the base model) but they certainly don’t perfectly reflect outputs. This is not showing a “shoggoth holding smiley face” situation, but it raises the possibility that as the model gets more powerful, Claude’s words could diverge more from its thoughts.
I guess I misunderstood because of phrases like “I’m pretty confident.” But also, my question was literal. I wanted to know if you believe the AI assistants are telling you what they really think. I view Claude as something like a character the model is playing, underneath which there is an actor who might be having very different thoughts. This is somewhat supported by experiments which e.g. change emotion weights and find that this doesn’t change the text output when the model is under observation but does change the tendency to cheat on some task.
Thank you for the kind words. I’ll try to be a little more language-conscious in future LW and EA Forum posts. Some people strongly objected to “normie” as well, which I think of as a playful and inoffensive term.
I don’t take the AI assistants to be reliable narrators. Do you?
This post might also be very relevant: https://www.lesswrong.com/posts/cJX2ssssGoYqnijwi/the-talker-does-not-control-the-doer-in-current-ais
I think it’s important to understand that nearly all the people saying “just unplug the AI” are saying so in good faith. They just aren’t thinking of the AI as an intelligent actor with agency. It’s a category error, and that’s what you have to focus on addressing.
Fully agree with 2, 3, 4, 5, 6, 12
Partly agree with 8, 9, 11, 14
Fully disagree with 1, 7, 10, 13
Partial disagreements:
8. True people don’t like probabilities and bets, wrong that this sucks: Rats overapply Bayesian reasoning to areas where they lack adequate frequency data, and sneak in biases in spite of themselves. That’s partly why people are rightly skeptical of probabilites of very speculative unprecedented events. Could write a whole essay on this.
9. You underestimate how left-skewed it already is. Ted Cruz is an outlier. Until Trump moves, you can’t get the right. If it’s not going to become totally polarized, you need to reach people who can get in the room with Trump; no one else can change the fact. Tech VCs don’t matter, but Elon Musk and Peter Thiel do.
11. Yes to communicating better (wrote a detailed opinion post on that last night which I hope is approved soon) but hard no to leading with nanotech risk. It sounds too speculative and far-fetched.
14. Mostly giving this a thumbs-up, except I do think that instrumental convergence is a very basic and simple idea which you can and should get people to understand.
Full disagreements:
1. No, it’s still a sideshow. No more mainstream than aliens/UAPs. People are just barely beginning to look—be excited but not deluded.
7. We all know this understated the real p(doom) that these people believe in, and by a lot. Communicating this way was perfect, because it made clear that this is a real and present danger, which people would not otherwise assume. The thing you are saying is how Dario Amodei wants to frame it, but he has strong incentives to soft-pedal, which is the worst possible thing. Note what I’m saying here seems to contradict what I said above about Bayesianism—the thing is that I don’t think p(doom) if read as a precise number really means anything, but communicating in numerical terms is a very good way to make it sink in that it’s a big and present risk.
10. I just have way too much to say about this and I’m not sure if I need to try to raise it internally, but the first sentence is false and the others are horrendously bad advice for some people under certain assumptions. The one thing I’ll say is that it’s good for humanity if people are very frank about the large magnitude of the risk.
13. Ignoring people is bad. A lot of times what you think are insincere criticisms are not really.
Quick question: I notice Yudkowsky pontificating that one should talk about P(ruin|X) and not P(doom), and criticizing the use of the terms “doom” and “doomer.” What is the technical distinction between ruin and doom?