Curated. This dynamic is real. I appreciate it being described in clear terms. I have at various times in my life been an Alice and a person who is tired of the Alice around me being annoying. I have also watched Alice’s start going insane, and felt sorrow as I felt helpless to do anything about it. I wish we had a solution at all, but I am glad that we have a crisp description of the problem that does justice both to the annoyingly principled people who are slowly going insane, and to the people who are finding them costly to be around and sort of avoiding updating in a double-thinky, shadowy way.
Ronny Fernandez
The PDKU Awards Night
Curated. I like that this post takes a classic argument for expecting misalignment and then presents a bunch of observations that confirm something like the conclusion of that argument. I also like that you characterize the particular kind of misalignment it seems like the AIs have now, make it clear how it is different from other kinds, and articulate that characterization clearly enough to make it easy to think about.
I wish you had had the time to write a shorter post, but I know you’re a busy guy, so instead I will summarize the parts of your post I found most interesting in this curation notice.Summaries Of Arguments In Post:
The classic argument I have in mind goes something like:
Training can only select between two strategies to the extent that the grader can distinguish their outputs. On hard-to-check tasks, graders (python, AI, or human) cannot distinguish “looking successful” from “being successful.” Therefore on hard-to-check tasks, training assigns equal reward to looking successful and being successful. Since looking successful is in some sense “cheaper” than being successful, when a training process cannot distinguish these, we will end up with models that are trying to do the first. On easy to check tasks, this will happily coincide with success, but these same models will do much worse on hard to check tasks.
Reminds me of the picture of the world Christiano presented in What Failure Looks Like, so a pretty classic picture of misalignment with primary sources dating back to at least the ancient times of 2018.
A lot of folks (particularly at Anthropic) seem to think something like “claude models are pretty aligned” for reasons that seem mostly orthogonal to “Slopolis” type concerns. (See Even Hubinger praising the CEV of Opus 3 here.) Something here has seemed off to me for a while, and this post helped me articulate it. I think basically when you ask the models questions about their moral judgments, they seem to give pretty nice answers, but it turns out that saying things that sound nice when asked about moral judgements is totally consistent with the models being deeply Slopolis-type misaligned. Checking the niceness of an expressed moral judgment is not very hard. So yeah, maybe current AIs are aligned in the sense that their expressed moral judgments seem nice and reasonable, but that doesn’t cause me to feel much better about their overall alignment.
Apparent success seeking is also distinguished from scheming. The post emphasizes that the kind of misalignment characterized does not require any scheming; motivated reasoning, confabulation, doublethink of various kinds, and other “subconscious drives” baked in during training are sufficient. It looks to me like we are seeing only moderate evidence of scheming. Things could have looked much worse re: scheming, but that also does not cause me to feel much better about the overall alignment of current AIs. This post helps me articulate why.
This Slopolis-type misalignment is still quite bad in particular because it differentially hurts alignment (and control imo) efforts. Alignment research is the sort of thing that requires producing hard to check outputs, eg, conceptual arguments, threat modeling, interpretability claims, intuitive frames on not yet formalized problems, claims about what kinds of cognitive traits ML tends to find, claims about what certain kinds of minds are like, maybe some philosophy, etc. Capabilities research by comparison tends to be relatively more composed of checkable tasks, eg, benchmarks, runnable code, loss, reward. Since alignment research is more composed of tasks on which proposed solutions are hard to check than capabilities research, Slopolis-type misalignment differentially makes AIs worse at alignment research.
There’s a distinct argument here about how this kind of misalignment is bad news for the prospect of safely delegating alignment to AIs. Safely delegating alignment to AIs is going to require us being able to tell good alignment research from bad alignment research, and that could be pretty rough if we have super smart things doing our alignment research which also are mostly motivated to make things that look good, even if they are secretly really bad. So motivated AIs might even go out of their way to make it hard to find out that their proposals are bad so long as they can make them seem good. Sounds like a truly awful position to be in.
The argument for takeover risk from Slopolis-type misalignment is trickier, but interesting. I’m not sure how to evaluate it and haven’t thought about it much other than to summarize it here:
Seems like apparent success seeking is a general, cross domain drive in current AIs. If things proceed roughly as they have with no particular intervention, then pretty plausible that you end up with models that are apparent success seeking but on longer time horizons. If an AI is basically just trying to “maximize long-run apparent success” and it has an option to subvert oversight, it’s probably going to be motivated to do so, since oversight is most of what causes the AI to have to get more of this lame actual success stuff when it could be getting even more sweet, sweet apparent success. Probably, such opportunities will come up, and subverting oversight is really quite a lot like taking over.
Finally, on why commercial incentives won’t address this problem. Labs are mostly motivated to fix the problems that their customers can legibly point to. Apparent success seeking is misalignment on the portion of tasks that AI users cannot legibly point to and complain about. Since they can’t complain about it, it produces relatively little market pressure.
I think this is legit and probably will be a real effect, but I also think that there is probably some market pressure anyway, though it might be hard to chase. Like, if I switched to using a new non Slopolis-type misaligned model and then all of a sudden, all of my vibecoded apps started kind of… being cooler, more pleasant to use, easier to work with etc, I would probably notice this, even if I couldn’t point to or clearly articulate any particular thing that was now better about my vibecoded projects. That said, obviously the market forces here are relatively weaker, and so I think that is a good reason to work on this particular problem.
A lot of the rest of the post is a catalogue of observations that suggest that current AIs are apparent success seeking and so misaligned in this particular way. That seems good and worth cataloguing, but I’m not going to summarize it here.
It was indeed not very long until the handles in this post became very useful to me.
Curated. I have long wanted a handle for this thing I often feel when engaging with (I think most?) people on Twitter. There’s a sadness I feel that isn’t really about any particular interaction but rather about a general sense that a large portion of the population seems to be uninterested in anything even approaching good-faith public argument. I fear that it’s not just a lack of interest. It’s more like we have lost some critical bit of our culture and are in the process of losing the ability to have good-faith public disagreements at all. Benquo’s handle of “a city beyond shame” will stick with me. I also like the guilt, shame, depravity frame. It’s good that Benquo doesn’t offer a method. A method wouldn’t work, and probably wasn’t really what was up with Socrates’s whole deal becoming a big deal. The diagnosis that reasoning with people who have gone to the side of depravity doesn’t work at all seems pretty legit to me, perhaps a bit overstated; I still think good-faith is a pretty surprisingly good rhetorical move in our civ, but it points to something real.
I have a worry about the post, or rather what I imagine to be the vibe behind the post (and the vibe it will activate in its readers): that it is too easy to use as a weapon against others (those terrible wrongthinkers) and quite hard to use against yourself. That’s the case with most such criticisms of some cognitive/cultural pattern, but I think it’s worth flagging because a reader who nods along to this is likely not imagining themselves, and probably should be more than they are. I get more out of this when I remember that I am sometimes depraved (or at least ashamed) wrt a partner, or a coworker, etc, in that in those cases, I have also sometimes given up on good faith disagreement.
We need better handles for thinking together about what is up with the sad state of our civ’s public conversations, and this post gave me some. I hope the post helps me remember to be alive as I continue to go about making contact with others.
Jynx
Yeah, I understand, but term dilution happens by three or four small concessions like this.
Curated. I am personally interested in using embryo selection if I ever have children. It is quite nice to have all of the relevant information in one place, and to be able to get information for specifically my situation/goals. I appreciate your use of our new iframe feature! Having calculators that let me figure out how much of an effect (and with what variance) I should expect per dollar under different approaches is extremely useful. I hope others find it useful as well. I think there’s a decent chance this post causes the existence of a counterfactual (and counterfactually more awesome) person, which obviously warrants curation. Thank you for making it!
Gripe. This doesn’t really seem like it is about superbabies, a term which has historically been applied only to really quite unusually smart babies who might, eg, be able to solve the alignment problem or whatever. Diluting the term seems like a mistake to me. If your baby isn’t as smart as Von Neumann, it’s not a superbaby, merely sparkling genetically enhanced offspring. That said, I am still a staunch advocate for genetic freedom in general, and I am glad this post exists overall. Consider changing the title some?
Lesswrong Liberated
Ahhh, I see there was already a Wei-Dai post about this referenced in the comments.
Annoying highly galaxy brained consideration, but pretty plausibly, considerations about how big the cosmic endowment is are dwarfed by considerations about how big our logical endowment is, eg, I would probably prefer to be in a much smaller but much simpler universe than in a bigger but more complicated universe for acausal trade reasons.
Curated. This is some of my favorite LLM sci-fi I have ever read. The mystery is extremely captivating. Took me several reads to really get what was going on, and talking to models about it was pretty rewarding. I think it will read to some as LLM generated, but I am fairly confident that no LLM was involved in the writing of this piece. I have rarely felt so bad for a character as I felt for this fellow, and I think the situation they find themselves in is an interesting puzzle of the kind that makes ratfic great.
LessOnline ticket sales are live! (Earlybird pricing until April 7)
Curated. Conceptually building on The Rise of Parasitic AI seems worth doing. It’s a potentially important phenomenon that may end up playing a big part in how the coming century plays out. It’s reminiscent of this section of Christiano’s “What Failure Looks Like”. Exploring the extent to which we can bring an existing and mature discipline’s concepts and models to bear on the phenomenon is a great approach.
I appreciate caching out that process in terms of what predictions we should make if the approach makes sense. I think it is unlikely that this particular approach ends up being very fruitful, but only because every conceptual approach to a new kind of problem is unlikely to end up being very fruitful.
I hope you continue to try finding plausible ways to apply the concepts and models from successful, mature disciplines to bear on the sorts of problems we tend to care about around here.
Mod note: this post violates our LLM Writing Policy for LessWrong and was incorrectly approved, so I have delisted the post to make it only accessible via link. I’ve not returned it to your drafts, because that would make the comments hard to access.
David, please don’t post more direct LLM output, or we’ll remove your posting permissions.
Sorry, I did not realize you were joking
I think you should get better at distinguishing assertions from other kinds of speech acts that people make using the indicative mood.
Fwiw, I currently have not been at all convinced that I made any mistake in curating this post or in the content of the associated curation notice. Sure seems like there’s a lot of people who feel quite strongly however, and I would be interested in hearing more arguments.
This is possibly my favorite LLM scifi I have ever read. Extremely engaging.
Curated. This piece was extremely easy to read and I like how it clearly points out that some kinds of misalignment are more likely given certain kinds of training setups. I also like taxonomies of possible kinds of misalignment in general. It reminds me of this section of the appendix of Ryan’s Current AIs seem pretty misaligned to me. I’d like to see more taxonomies of possible kinds of misalignment and when and why we can expect them to show up. This seemed straightforward at least in retrospect, but I am glad to have it written up somewhere.