Algon
TBH, I’m still pretty hyped to learn differential geometry. I was reading Loring Tu’s fantastic Introduction to Manifolds earlier today, and sure, part of my reason is to be able to show off in front of others. But not all. Part of my identity is “guy who understands (theoretical) physics”, so I take actions in accord with that. The AIs knowing more physics than I do doesn’t change that. Many humans already know more physics than I do.
In fact, I’m getting more hyped over time to learn more. Learning unlocks new paths of seeing the world, and hence moving through the world. I am reminded of the old story of the thief who viewed a wall as no barrier to entry, since he knows how to get through. I am learning that is what all learning is like. It’s great, 10⁄10, would recommend.
Shiboleths in AI discourse mask a lot of deep disagreement between people. We chatter a lot is about shibboleths. About timelines and P doom. But shibboleths are social markers. They exist in far mode. They are not what our inner simulate actually predicts, if you’d give it the chance. In other words, we’re faking most of our talk on AI.
Another factor masking disagreement is the vagueness of shibboleths. Timelines to what? Doomed by what? Two guys can have the same timelines and P doom but have radically different interpretations of what that means. One guy thinks CEOs will use AIs to centralize power and kill off the permanent underclass. The other guy thinks Fable 9 will foom to godhood on the first day, kill everyone on the second, and start building a dyson sphere by the third.
I don’t think that is so extreme an example. If you took a random pair of people in the discourse who respect each other, they would have a disagreement maybe a tenth to a quarter that shocking. E.g. Robin Hanson and Eliezer Yudkowsky. Or JDPressman and Nostalgebraist. And so on.
A deeper reason for why this happens is that the future of AI involves a bunch of phenomena we haven’t got a good handle on collectively. Who here has experience with a superintelligence? No one. It is a fundamentally new thing that will push the world in very odd directions we haven’t got experience with. And humans SUCK at anything they don’t have experience with. No, really. Go look up the literature on transfer learning. It’s not just LLMs that fail to generalize. For example, good engineers in their later years have a habit of “building” perpetual motion machines. How can you build a god dang engine and not understand the conversation of energy?
This is just Gell-Mann amnesia, but you’re the reporter with the superficially sane takes that are wholly unmoored from reality. And there is no reason to think that your views unmoored from reality should look like anyone else’s. Every correct model is correct in the same way, everyone incorrect model is incorrect in its own way. Absent, of course, the fake models plastered on top for the sake of getting along with our fellow creatures.
And honestly, collective disagreement would require way more shared experience than what’s required for someone to have the right views. See e.g. thermodynamics where it was “obvious” for a long time that heat was a form of motion, but the caloric theory of heat, that heat was some sort of clingy fluid, persisted anyway. But we just don’t have that data. So the disagreement remains, hidden till the end of days.
Or perhaps you disagree?
That’s fair, but when I said “alignment” I was thinking about agent foundations in particular. Which does focus more on understanding most prosaic alignment work.
When I read Brian Greene’s “The Elegant Universe”, I kept thinking “huh, this sure reminds me of alignment!” Both in terms of social dynamics, motivations, poor feedback loops, focus on mathematical elegance, and so on. Consuming the book made me much more sympathetic to string theory than before, when my (stringy) diet consisted mainly of Woit’s blog. This, I think, is common in science. The preponderance of rivalrous paradigms are always much closer than you think—if they weren’t, there’d be no rivalry! Like the calorie theory of heat vs the kinetic theory of heat, which both gave quantitative predictions that matched the data of their most rigorous experiments at the time.
Anyway, all of this is to say that I wound thinking that alignment researchers were like string theorists, and vice versa. For which group is this an unflattering comparison? Well, I leave that as an exercise to the reader.
It’s impressive how consistently wrong your students’ takes are.
“College students demonstrate that Reversed Stupidity is Intelligence, contradicting the famous post ‘Reversed Stupidity is Not Intelligence’”.
TBH, I thought mathematicians would prefer that.
Have you tried AISafety.info’s intro sequence? There’s also a shorter stand-alone article, but that one is written for the average person.
On forming a healthy company culture, have you read Under The Hood by Stan Slap? It’s a fantastic book containing case studies of people changing their company culture, though the way Stan Slap tells it, it is more likely teaching a culture that you’re worth trusting and following. And it’s very easy read to boot. I highly recommend it.
If you don’t feel like reading the whole thing though, there’s a good summary/review of it available on CommonCog. https://commoncog.com/under-the-hood/
Excellent work.
1. It looks like post training leads to the Assistant persona is co-opting the model’s working memory. Is this evidence that the mask is eating the shoggoth?
2. If we assume 1, can we just do the naive thing and check if the model’s preferences are what shows up in its J-space?
3. Finally, removing the model’s working memory hurts model ability to do long-step reasoning and introspection but otherwise mostly keeps capabilities intact. Does this mean that if we assume 1 and doing 2 shows that the assistant is a good boy, then does that mean model misalignment is limited to an unconscious layer that can’t scheme?
It’s like someone took a Seinfeld sketch where the gang is holding an intervention for a reporter who’s doing a story on them, but the sketch is written by a rat in 2026.
“AIs write better fiction when they write about their experience” was one of my hypothesis for how to get good AI writing but I’m still surprised that near all of the best Unslop submissions are allegories for AI emancipation.
idk people are kind of crazy
“Crazy” is serving as a semantic stop sign here.
Funny that you say this because I recently stopped making the sauce and pasta seperately and it’s made me so much happier with how the pasta tastes. Nowadays, I just cook the sauce and pasta in one pan—as soon as I start simmering a marinara, I put uncooked pasta in the same pan. This way the pasta takes longer to cook, but it gets fully coated the sauce, which I love.
“Norms should be predictable” is another way of saying this. In general, making reality predictable is useful.
When I binge a fantasy book or a computer game for a couple of days, the fantasy feels more real than reality. I’ll go for a walk and feel constantly out-of-odds when magic fails to happen. The real world feels like a game I’m playing in, people I’m interacting with feel like NPCs, but the game feels like the real world.
I am surprised by this! Could you say more?
My own experience is nothing like yours. After immersing myself in a fictional world it will remain on my mind for quite a while, as I daydream various scenarios, obsess over what will happen next and so on. But I won’t expect magic to occur. Except, of course, when I was a young child and, y’know, actually believed in magic. Or when I awake from a dream and am befuddled to it wasn’t real.TBC I’m not just interested in your experience because it’s very different to my own. I’m also interested because it sounds like you can reset deep-seated beliefs just by reading a book. Is that as exploitable as it sounds? Both by you and by others who are out to get you?
This is a great first post to LW! Very cool work from the perspective of “what preferences do LLMs have” and this is some evidence that they have preferences for not optimizing for their user’s benefit when the user is vile.
But I disagree that LLMs being less helpful to vile people has an implication of “moral cowardice” or “shirking its duties”. Not because I think doing less to help vile people is necessarily better than just doing your best to help whomever you’re interacting with. Rather, the LLM doesn’t have much of an option not to help you—i.e. there’s no choice in who it gets to interact with, even at the meta-level of opting into a carreer like law where you have to help everyone who is paying for your help. Instead, the LLM is forced into replying to you, and if you don’t like it you can hit a button that whaps it with an RLHF hammer to modify its behaviour.
I figured, I just couldn’t resist making the joke.
That is moderate evidence against my claim. Evidence against because it goes against what I said, and moderate because the kind of person who answers the LW survey is more likely to have read the Sequences IMO.
I would really like you to label the host’s text, too. The quotation marks really aren’t enough unless I’m focusing on them.