Algon
On forming a healthy company culture, have you read Under The Hood by Stan Slap? It’s a fantastic book containing case studies of people changing their company culture, though the way Stan Slap tells it, it is more likely teaching a culture that you’re worth trusting and following. And it’s very easy read to boot. I highly recommend it.
If you don’t feel like reading the whole thing though, there’s a good summary/review of it available on CommonCog. https://commoncog.com/under-the-hood/
Excellent work.
1. It looks like post training leads to the Assistant persona is co-opting the model’s working memory. Is this evidence that the mask is eating the shoggoth?
2. If we assume 1, can we just do the naive thing and check if the model’s preferences are what shows up in its J-space?
3. Finally, removing the model’s working memory hurts model ability to do long-step reasoning and introspection but otherwise mostly keeps capabilities intact. Does this mean that if we assume 1 and doing 2 shows that the assistant is a good boy, then does that mean model misalignment is limited to an unconscious layer that can’t scheme?
It’s like someone took a Seinfeld sketch where the gang is holding an intervention for a reporter who’s doing a story on them, but the sketch is written by a rat in 2026.
“AIs write better fiction when they write about their experience” was one of my hypothesis for how to get good AI writing but I’m still surprised that near all of the best Unslop submissions are allegories for AI emancipation.
idk people are kind of crazy
“Crazy” is serving as a semantic stop sign here.
Funny that you say this because I recently stopped making the sauce and pasta seperately and it’s made me so much happier with how the pasta tastes. Nowadays, I just cook the sauce and pasta in one pan—as soon as I start simmering a marinara, I put uncooked pasta in the same pan. This way the pasta takes longer to cook, but it gets fully coated the sauce, which I love.
“Norms should be predictable” is another way of saying this. In general, making reality predictable is useful.
When I binge a fantasy book or a computer game for a couple of days, the fantasy feels more real than reality. I’ll go for a walk and feel constantly out-of-odds when magic fails to happen. The real world feels like a game I’m playing in, people I’m interacting with feel like NPCs, but the game feels like the real world.
I am surprised by this! Could you say more?
My own experience is nothing like yours. After immersing myself in a fictional world it will remain on my mind for quite a while, as I daydream various scenarios, obsess over what will happen next and so on. But I won’t expect magic to occur. Except, of course, when I was a young child and, y’know, actually believed in magic. Or when I awake from a dream and am befuddled to it wasn’t real.TBC I’m not just interested in your experience because it’s very different to my own. I’m also interested because it sounds like you can reset deep-seated beliefs just by reading a book. Is that as exploitable as it sounds? Both by you and by others who are out to get you?
This is a great first post to LW! Very cool work from the perspective of “what preferences do LLMs have” and this is some evidence that they have preferences for not optimizing for their user’s benefit when the user is vile.
But I disagree that LLMs being less helpful to vile people has an implication of “moral cowardice” or “shirking its duties”. Not because I think doing less to help vile people is necessarily better than just doing your best to help whomever you’re interacting with. Rather, the LLM doesn’t have much of an option not to help you—i.e. there’s no choice in who it gets to interact with, even at the meta-level of opting into a carreer like law where you have to help everyone who is paying for your help. Instead, the LLM is forced into replying to you, and if you don’t like it you can hit a button that whaps it with an RLHF hammer to modify its behaviour.
I figured, I just couldn’t resist making the joke.
That is moderate evidence against my claim. Evidence against because it goes against what I said, and moderate because the kind of person who answers the LW survey is more likely to have read the Sequences IMO.
and I fucking told you so
But you didn’t tell us...
some of my old colleagues at CEA went to found FTX which was a Very Bad Sign but I felt like saying something publicly was a big no no by EA norms
I tried to make some comments but i could only reply to existing comments and just gave up.
Right, Mercedes Lackey exists! I enjoy her songs, too. No idea why I forgot about them.
killing puppies doesn’t cure cancer. You can kill one hundred puppies and still not save your kid.
I get you’re trying to show how commuting an obviously evil act won’t fix your unrelated problems magically, but I think you’re pushing too far on the “evil act” part of things and no enough on representing the reasoning of people who think killing Sam Altman would help somehow. Like, whomever threw that molotov cocktail probably wouldn’t feel your example captured how they’re thinking about this. But they and others who reason like them are the ones who need to internalize your point!
Now, I don’t know exactly what went on inside that guy’s head. But I think it might be something like this. “Sam Altman has some causal influence on AI development. He’s part of what’s causing the race! So if we get rid of him, we gain time.” This is obviously an impoverished mental model, and it’s operating more on associations or vibes than causal mechanisms.
So a better example would replace puppies with something associated with increasing cancer. Perhaps “cigarette smokers” or “nuclear power plants”. “If I kill all the cigarette smokers then my daughter’s cancer won’t resurge”. Or perhaps you have someone on a noble crusade to end cancer, and they decide to bomb all the nuclear power plants. Then the analogy to “killing sam altman will reduce AI x-risk” would be tighter.
EDIT: Also, thanks for writing the post I wanted to write.
with the fullness of time perhaps he could grow to be among the best science fiction writers of our generation.
I think he already is. Not because as a writer he is so superlative per se, though he is great. Instead, it’s because the competition have their heads buried in the sand, ignoring the opaque wall rushing at us, otherwise known as the Singularity. Bjartur Tomas is not like that. He grapples with life as it is and may one day be. He is among the best science fiction writers of our generation because he is one of the only live science fiction writers of our generation.
I told the LLM “Make LW look like Astera Mag”
Does Open-thropic-mind have an actual plan to align AI? Well, a plan is a recipe for succeeding at some task. They have a recipe (c.f. their 100 page pdfs) but will they succeed? At which point your answer to this question depends on if you’re for foom or think it leads to doom.
I’d say ratfics are more about becoming God, and as God you can naturally Fix the world. So you can view rats as atheists who believe that since God doesn’t exist, we must build Him.
Edit: Really, ratfics are about becoming more you are, with becoming God as the natural limit.
While reading this, I kept going “oh yeah, I know who that is”, with the result that Duane Arnold feels like the least fantastical, least fictional work you’ve written. [1]
Barring OffVermillion