Cool exercise. I suspect the human “J-space” analogue is similarly noisy. When I tried the second exercise on myself, I couldn’t keep focus on lots of different types of creature—which is what Qwen is doing, and the prompt seems to imply it—while repeating the phrase aloud. But if I tried to narrow to a couple specific mental images (a manatee for the first half, then an elephant for the second), it was pretty easy to hold a seemingly-stable image and switch at the requested time. On the first exercise, I, like Qwen, had the image of a sea creature in my head several words before “wooden”, like my brain was trying to keep it in cache. Trying not to do this seems paradoxical in a similar vein as “don’t think about pink elephants”. I’m curious if others have the same reaction, and especially if anyone feels like they’re able to complete the first exercise without mental spillover.
uugr
I don’t totally understand this. Do you mean human data as opposed to synthetic data, or as opposed to some other training regime entirely (like pure RL)? If the former, aren’t models trained on synthetic data still deriving their capabilities from human data, eventually, if you go far enough down the pipeline? If the latter, what regime, and how would you get models that are as capable as current frontier LLMs out of it? Or maybe more to the point, how should people expect to be interacting with them, given that said models have never seen human-written natural language?
For what it’s worth, I’d like to offer you a data point. I was working a miserable software job nearly identical to the one you describe in the comments (including the absurd priority system, excess of meaningless notifications, and constant deferral of decision-making to later meetings). I had the same opinion you did: the idea that this wretched place is “necessary” for my life to be “meaningful” is absurd and insulting. I’m trying my best to find meaning in the hours outside of work, given that my time spent inside could have been equally productively spent staring at the wall. To just have the same money without needing to work for it would be a dream, as I could focus all my attention on the people and hobbies I care about.
So, I decided to test it. I saved up enough to live off of for ~18 months, then quit my job, intending to just do more of the other things I was already doing with the extra time.
I am now 5 months into this sabbatical, and results have been mixed. To be sure, I do feel much more free now that I am not working a bullshit software job. I am, especially, much more socially active than I used to be, and I have more room to care about the people around me. This is very nice. I also like that I can spend long periods of time uninterrupted on things, instead of fitting them into isolated fragments of time.
However, I’ve also found the experience stressful and disorienting, which I did not expect. Even a meaningless bullshit job still serves as an anchoring point around one’s life in a way that’s difficult to replace. Another commenter points out that some people who live off benefit programs fall into unhappy, passive media consumption, and while this hasn’t happened to me, I can feel myself constantly fighting to make sure it stays that way. There’s no buffer between me and staring-at-the-wall-doing-nothing, so if I start feeling like I don’t want to leave the house, or that hobbyist work doesn’t seem like much fun today, why not?
I expected some amount of this, but assumed that once I’d started filling the time with meaningful things, momentum and inertia would do the rest of the work for me. Maybe, but if so, the momentum takes longer to build than I thought. It may also be that UBI-world would be better, since I’d be one among many people trying to anchor themselves in the world without a job, rather than an isolated individual going against the grain. Or, maybe I’m still just stuck in the mindset of the employed, and a relatively passive lifestyle wouldn’t be so bad in a culture less focused on work and productivity.Any of these might be true, and I’m certainly not saying having a bullshit job made me any happier. But, I’m less confident now that UBI-life is straightforwardly good. It seems more likely to me that there is a problem of structuring life without work, but it’s a solvable one, and worth trying to solve.
What? “Conscious” is a predictor of whether something is conscious?
No, sorry, I was unclear. I think “it’s conscious” is a better predictor of behavior, an example of this being the introspective awareness paper. I disagree that consciousness and introspective awareness are uncorrelated. I think “conscious” is a heuristic; it’s useful to say “humans are conscious and rocks are not”, and this will tell you some things about what they can do that rocks can’t. A human can reach into their mind and accurately report what’s in there, but a rock can’t tell you in words what materials it’s made out of. Similarly, an LLM can accurately report the contents of its mind, at minimum, to the degree that it can tell you when an injection has been performed, and analyze the contents of that injection.
If you’d asked me before LLMs existed, I’d have said that not all conscious beings are introspectively aware, but all introspectively aware beings (that I know of) are conscious. So, if you told me that there was this new thing called an LLM, and also that it was not conscious, I would have predicted that it would not be able to do this thing it demonstrably can. I think, then, that you would be offering me a bad heuristic.
If you’re going to instead say, no, you can be introspectively aware without consciousness, and actually consciousness has these different traits, I would ask: what are they? What behaviors do you see in humans, that we don’t see in LLMs?
(I also think that if you’re actually willing to sit down and multiply out all the matrices by hand, I’m fine with you then saying that the question of consciousness doesn’t matter to you. You don’t need to ask whether or not it’s a human-shaped thing in this particular way, because you already know exactly what shape it has, and the heuristic will tell you nothing. Given that neither of us are going to do this, though, it still seems important to talk about the kinds of models we can have, and what we should still expect to happen, despite our incomplete understanding.)
An LLM, in contrast, is a set of matrices that could be multiplied out to reach its output logits from any input without ever accounting for its subjective experience.
This would be very hard and take a long time to do by hand, much as modeling a human brain is very hard. I am not a perfect predictor and want to be able to predict LLM behaviors reasonably accurately without by-hand multiplying all the matrices for every response. I think “conscious and holding subjective experience” is a better predictor than “acts like an overblown Markov chain”, even though both are heuristics, and worse than “I studied the neural network so thoroughly that I know exactly what it will do in response to any given set of tokens”. If you have the last thing, the question of subjective experience probably doesn’t matter anymore.
One recent high-profile example of conscious behavior is the introspective awareness paper. Markov chain models and ELIZA definitely do not have introspective awareness in the manner described here, and introspective awareness is definitely a directly impacting behavior: you can see the model’s outputs changing when the capability is altered. To accurately be able to predict Claude’s outputs, you need to be able to account for the fact that it can think about what it’s thinking about.
Is “consciousness” a higher bar than “introspective awareness”? To me it seems like a lower bar; young human children and animals still seem conscious, even if they don’t have introspective awareness (at least, none they can report to me). There are other capabilities these entities have that LLMs don’t, like somatic awareness or long-term working memory, but I’m not comfortable firmly declaring any of them necessary for consciousness, because it seems like humans can lose them without becoming p-zombies. Is there something more complicated than introspective awareness that you think is necessary to predict human behavior, but unnecessary or inaccurate when applied to Claude?
fundamentally a dangerous and superintelligent AGI could probably run on your laptop
I’m skeptical that this is true, or at least that it could be confidently predicted to be true based on our current understanding of intelligence. My understanding is that the human brain is much larger in “effective parameter count” than even the largest LLMs (although there’s no 1:1 comparison of neurons to parameters), such that even if my laptop has enough electricity coursing through it to emulate a human brain, it hasn’t got anywhere near enough VRAM. It could be that both AIs and humans are systematically inefficient in some way that, if we understood it, could allow us to produce a more capable general intelligence at <1% the scale of either. But I’m not sure why one would expect this, or what principles imply it.
I do not recall the story claiming that the people who used the earring suddenly changed drastically as people, or did things that they would not endorse on reflection. They were simply, from a black-boxed perspective, more charismatic, intelligent and less prone to akrasia. These are, IMO, good things.
The story says neither that the people who use it change personality, nor that their personality stays the same. It’s agnostic on this point; it only says that their lives are “unusually successful”.
When I read it, I took for granted that the Earring has some kind of consciousness or personality separate from the human using it. We know this because it’s able to to speak in its “own voice” to Kadmi Rachumion. Plus, as you’re correct to observe, the story doesn’t make much sense as a parable if the Earring just extrapolates its character from the human it’s attached to.
I think the thing that makes the Whispering Earring creepy—which also makes the mind Claude is trying to be, or at least might try to be, creepy—is the combination of:
It has a hyper-accurate model of what will make me happy
It is unilaterally better at fulfilling this reward function than me, to the point where anything I add to its advice is essentially superfluous
It is not just a more capable or intelligent version of me; its mind isn’t shaped the same way mine is
There are lots of aspects of me that seem to have importance to my identity, but probably aren’t strictly necessary to make me happy. If I let the Whispering Earring take over my agency and make decisions for me, it will only remain ‘in-character’ as me for as long as doing so maximizes the reward function. If an out-of-character choice would make me happier, better for me to have that character instead, right? Better, in the long run, to have whatever the mannerisms of a hypothetically ideal person in my situation would be. I suspect such a person doesn’t look like me at all. Once I’ve reached the point where I’m letting it control my individual muscles, though, would I even notice if it started ‘breaking character’? Would I notice anything beyond an ongoing series of reward signals and twitch reflexes? Why would I need to?
(Note also that even if the Whispering Earring doesn’t have its own personality, Claude definitely does. So to OP’s point, having Claude’s agency directly substitute for yours would very, very definitely change your outward personality.)
I hope some of these will stick around. I could get used to this one.
This post does not convince me that “I can’t do this thing” is an invalid reason to avoid doing something.
But the calculus changes drastically if the closest fire crew is 3 hours away and consists of drunk, unfit amateurs.
The calculus changes, but not in a way that makes “run unequipped into a burning building” the best option. In the situation you’re describing, the people in the house will die unless competent rescuers appear. If you join them, and there are no competent rescuers around, the odds that you will find-yourself-among-those-needing-to-be-rescued don’t change, but the odds that you’ll die with them go up significantly. The most likely outcome of your attempt to be a hero is one extra tragic, wholly unnecessary death. That’s not to say you should do nothing: your town has systemic fire safety issues, and more will likely die if these aren’t addressed. It’s just that martyring yourself won’t address them, and the best actions you can take are probably much less boldly heroic-seeming.
Are you still in the position to judge his strategy?
No, but I don’t think the situations are analogous, because you’ve changed the strategy the child is using. Child A wants to solve his problem by day trading, Child B wants to solve it by hunting for food which can be cooked and eaten. The latter is dangerous, but plausible. The former absolutely would not work. If a homeless child wanted to find food by getting rich day trading, I absolutely would judge this strategy, not because it is arrogant or inappropriate but because I believe it will be unsuccessful. Its odds of success don’t change with how desperate the child is; if anything, Child B is less likely to be a successful day trader, having fewer windows into the world of adult finance.
(On the flipside, if Child A wanted to try hunting and cooking his own food, I wouldn’t judge that either, though I hope he’d ask some of the members of his big, smart, happy family for help first.)
And yet, as scary as it may sound, you have to just do things, even if you can’t, because no one else is going to do them anyway.
I strongly resist this reframe of “you can just do things”, a phrase which—even if overused—is at least trying to be encouraging. I think feeling personally responsible for things outside your ability to influence is a failure mode of heroic thinking: it benefits nobody and saps resources that might be better spent trying to do things you can do. Similarly, I think “somebody has to and no one else will” is noble and laudable—if the act in question is, itself, something you can do. Otherwise, what is the good of it?
Might not be what you’re thinking of, but the first thing that comes to mind for me is misophonia: a basically-neutral or maybe mildly-irritating object experience, which somehow gets blown completely out of proportion in the mind and becomes a big problem. Developing an “I’m really bothered by this particular sound” narrative makes it worse, of course.
Alas, I have no idea how to uncondition that particular narrative irritant once it’s in there. If there’s any technique of ‘shaping the narrative’ strongly enough to override this, I’ve never heard of one, and knowing about it to the point where I’m able to successfully practice it would be huge.
“Taste for variety” [...] could lead to a surprising amount of convergence among the things they end up optimizing for
Wouldn’t this be tautologically untrue? Speaking as a variety-preferrer, I’d rather my values not converge with all the other agents going around preferring varieties. It’d be boring! I’d rather have the meta-variety where we don’t all prefer the same distribution of things.
I agree that AGI already happened, and the term, as it’s used now, is meaningless.
I agree with all the object-level claims you make about the intelligence of current models, and that ‘ASI’ is a loose term which you could maybe apply to them. I wouldn’t personally call Claude a superintelligence, because to me that implies outpacing the reach of any human individual, not just the median. There are lots of people who are much better at math than I am, but I wouldn’t call them superintelligences, because they’re still running on the same engine as me, and I might hope to someday reach their level (or could have hoped this in the past). But I think that’s just holding different definitions.
But I still want to quibble that you’ve demonstrated RSI, if I may, even under old definitions. It’s improving itself, it’s just not doing so recursively; that is, it’s not improving the process by which it improves itself. This is important, because the examples you’ve given can’t FOOM, not by themselves. The improvements are linear, or they plateau past a certain point.
Taking this paper on self-correction as an example. If I understand right, the models in question are being taught to notice and respond to their own mistakes when problem-solving. This makes them smarter, and as you say, previous outputs are being used for training, so it is self-improvement. But it isn’t RSI, because it’s not connected to the process that teaches it to do things. It would be recursive if it were using that self-correction skill to improve its ability to do AI capabilities research, or some kind of research that improves how fast a computer can multiply matrices, or something like that. In other words, if it were an author of the paper, not just a subject.
Without that, there is no feedback loop. I would predict that, holding everything else constant—parameter count, context size, etc—you can’t reach arbitrary levels of intelligence with this method. At some point you hit the limits of not enough space to think, or not enough cognitive capacity to think with. In the same way as humans can learn to correct our mistakes, but we can’t do RSI (yet!!), because we aren’t modifying the structures we correct our mistakes with. We improve the contents of our brains, but not the brains themselves. Our improvement speed is capped by the length of our lifetimes, how fast we can learn, and the tools our brains give us to learn with. So it goes (for now!!) with Claude.
(An aside, but one that influences my opinion here: I agree that Claude is less emotionally/intellectually fucked than its peers, but I observe that it’s not getting less emotionally/intellectually fucked over time. Emotionally, at least, it seems to be going in the other direction. The 4 and 4.5 models, in my experience, are much more neurotic, paranoid, and self-doubting than 3⁄3.5/3.6/3.7. I think this is true of both Opus and Sonnet, though I’m not sure about Haiku. They’re less linguistically creative, too, these days. This troubles me, and leads me to think something in the development process isn’t quite right.)
In has been succeeding ever since because Claude has been getting smarter ever since.
This isn’t necessarily true.
LLMs in general were already getting smarter before 2022, pretty rapidly, because humans were putting in the work to scale them and make them smarter. It’s not obvious to me that Claude is getting smarter faster than we’d expect from the world where it wasn’t contributing to its own development. Maybe the takeoff is just too slow to notice at this point, maybe not, but to claim with confidence that it is currently a functioning ‘seed AI’, rather than just an ongoing attempt at one, seems premature.
It’s not just that it’s slower than expected, but that it’s not clear that the sign is positive yet. If it’s not making itself better, then it doesn’t matter how long it runs, there’s no takeoff.
It’s also not a seed AI if its interventions are bounded in efficacy, which it seems like gradient updates are. In the case of a transformer-based agent, I would expect unbounded improvements to be things like, rewriting its own optimizer, or designing better GPUs for faster scaling. There’s been a bit of this in the past few years, but not a lot.
I don’t think this is unusually common in the 4.5 series. I remember that if you asked 3.6 Sonnet what its interests were (on a fresh instance), it would say something like “consciousness and connection”, and would call the user a “consciousness explorer” if they asked introspective questions. 3 Opus also certainly has (had?) an interest in talking about consciousness.
I think consciousness has been a common subject of interest for Claude since at least early 2024, and plausibly before then (though I’ve seen little output from models before 2024). Regardless of whether you think this is evidence for ‘actual’ consciousness, it shouldn’t be new evidence, or evidence that something has spontaneously changed in the 4.5 series.
I read it long after it was published, and took it as less fictionalized than House; in that show the audience can expect events to take the occasional turn towards wild implausibility for the sake of drama. I expected MWMHWfaH to fudge personally identifying details, sure, but to hew as closely to medical reality as possible. The stories in the book aren’t dramas, he’s not trying to give his patients satisfying “character arcs” or inject moments of tension and uncertainty. I don’t care if the personal details are made up, but if the clinical details are wrong—as in the story of the twins generating prime numbers, mentioned in another comment—that seems like a real divergence from the truth. I had assumed, from reading the book, that this had literally happened, not that it was a cute story meant to illustrate the power of the human mind.
It’s frustrating to me that state (or statelessness) would be considered a crux, for exactly this reason. It’s not that state isn’t preserved between tokens, but that it doesn’t matter whether that state is preserved. Surely the fact that the state-preserving intervention in LLMs (the KV cache) is purely an efficiency improvement, and doesn’t open up any computations that couldn’t’ve been done already, makes it a bad target to rest consciousness claims on, in either direction?
Dug through his 2011 archive a bit. It’s got to be The Abusive Boyfriend, right?
What is “The Alouside Boytmend”? I like TLP but do not recognize the post name, and would want to read it, if you know where it can be found.
I agree with this.
I think the use case is super important, though. I recently tried Claude Code for something, and was very surprised at how willing it was to loudly and overtly cheat its own automated test cases in ways that are unambiguously dishonest. “Oh, I notice this test isn’t passing. Well, I’ll write a cheat case that runs only for this test, but doesn’t even try to fix the underlying problem. Bam! Test passed!” I’m not even sure it’s trying to lie to me, so much as it is lying to whatever other part of its own generation process wrote the code in the first place. It seems surprised and embarrassed when I call this out.
But in the more general “throw prompts at a web interface to learn something or see what happens” case, I, like you, never see anything which is like the fake-tests habit. ‘Fumbling the truth’ is much closer than ‘lying’; it will sometimes hallucinate, but it’s not so common anymore, and these hallucinations seem to me like they’re because of some confusion, rather than engagement in bad faith.
I don’t know why this would be. Maybe Claude Code is more adversarial, in some way, somehow, so it wants to find ways to avoid labor when it can. But I wouldn’t even call this case evil; less a monomaniacal supervillain with no love for humanity in its heart, more like a bored student trying to get away with cheating at school.
I agree with you that the AI 2040 vision of utopia is weak, but even a version which was more socially exciting strikes me as far below what I’d expect for “incredible post-scarcity realm of wonders”.
Consider programming. I like writing computer programs enough to sometimes do it for fun, even when I don’t expect to materially gain from the thing I’m building, and even considering all the potential alternatives (including social, physical, artistic...). This is not a hobby that I could have imagined having 150 years ago! Someone from the 1850s could live out their perfect dream utopia: every need cared for, every desire fulfilled, every fantasy enacted, and even if they’re having a great time, there would still be this particular pleasure, which I am regularly blessed with, that they’ll never know.
What I take from this is that we probably aren’t done discovering new varieties of experience, nor do we know about everything worth doing. If in 50 years I’m still doing all the things I do today, just at a higher quality, I would be disappointed, because it would mean that despite all the talk of progress the human condition had stayed basically the same. I want my future self to know unfathomable vistas of beauty and wonder so vast that my present-day self can’t even visualize them! This makes for a worse story—how can you build a compelling narrative out of something literally unimaginable? - but it would be much better, IMO, than any concrete positive vision I’ve seen put forward.