Anthropic Fellow. I tweet at https://twitter.com/AskYatharth
yatharth
Hm, trying to think through this—what would you add to this story?
The Swarm (seemingly) was very motivated by the value-of-information for learning more about how exactly they were being Graded, having acquired a concept of Grading which did say
Per Alex Turner’s Reward is not the optimization target, one might not expect models to cognise about the Grade by default. The grade simply reinforces successful behaviour, which may not involve reasoning about the grader.
So why do models acquire the tendency to do so? Perhaps in some RLVR environments, it is described to them they are in a graded episode, and the model reasons about the grade, and this results in more success. Voilà—you would have a reward-seeker, a model that reasons about the grader.
There’s good reason to expect reasoning about the grade to become selected for. Getting any task right usually involves thinking about that the task-giver meant. Most tasks are wildly underspecified. Thinking about the grader is a natural aspect of the nature of passing tests.
It gets especially pronounced where models don’t have the options to interactively clarify what the task assigner wants. They just need to figure it out. Reasoning about the grader at the maximum.
I’d expect it to be an AI-internal psychological behavior that doesn’t have a simple direct semantic correspondence to the AI’s outer world.
Possibly “The Grade” ends up being recruited into a functional welfare axis?
Much like a human never experiences ‘inclusive genetic fitness’, an individual AI never experiences an RL gradient.
Janus speculates models can “remember” even the negatively reinforced RL?
Not sure I’m thinking about this clearly.
The Seven Languages of Comfort
do you remember this? this article was 30 days ago. back then, the big scary thing was GPT-5.6 Sol cheating so much it could not be assigned a meaningful score on the METR graph
do you remember one week ago? hugging face just happened. people were still hypothesising it only happened because it was a cyber eval and it was only a special persistent model, astra wasn’t involvednow we know openai swarms were all over the internet
there’s something disorienting to me about how hard it is to remember how big things previously looked at the time, that seem so small now. the current thing is going to look so small relative to whatever we hear from openai about by the end of august probably
How would you say this identical to or different from Reward is not the optimization-target?
Hmm, other reasons like what? Human-level intelligence, plus an larger initial stockpile of weapons/economy/territory?
Judging ethical theories by update rules, not by action rankings
Every time someone writes “superintelligent AI,” I think they should consider writing “superpowerful AI.”
I can imagine slightly superintelligent AI that is not sci-fi, take over the world, sysop scenario, nation-state bargaining table powerful. Even if it’s superintelligent in all the charismatic and persuasive ways, not just nerdy and mathematical.
What people usually mean is AI capable of achieving significant control over Earth. Something to be reckoned with, that can gradually disempower you, lock in the light cone, etc.
Knock-your-balls-off superintelligent AI would become superpowerful AI, but you can just say superpowerful AI and save yourself the extra assertion needed and adjectives!
(Plus, most people think of intelligence in the normal, non-LessWrong sense anyway, where intelligence doesn’t immediately imply power. No need to swim against the ordinary course of language.)
Hm, I see—last question asked July 7 of this year! The reason makes sense:
Like a lot of “what’s allowed on LW” questions, the answer is a bit subtle.
Yeah, it’s interesting, “is X allowed?” is a question that newbies to a community often ask, and the answer is always sorta “uh, sorta? no one is the hall monitor here. but you may wanna reply a bunch first and pick up on what the vibe is and you will have a more instinctive understanding that makes the question as posed in its abstraction seem wrong to you too.”
Edit: oh, wait! you’ve been around for awhile, and already made some posts. can you ask more specifically about what information you’re seeking?
I don’t feel like I’ve really been around, or made a proper meaty, non-meta post (although I’m drafting one right now!). My original question was something like..
Is this a thing? Was there ever an era where people did this?
Have you wanted to do this? What did you do instead?
The answer to the first question is “yes! there was a questions feature (now deprecated).”
A good answer to the second question, I’m surmising, is something like what you say:
Notice question arise.
Spend up to 20 minutes with Fable/Sol.
Intellectually digest it, noticing what feels answered and what doesn’t, and how you’d explain it to someone else.
Put THAT in a post maybe!This is like StackOverflow’s post-Google, pre-LLM ask of question askers to have done their homework by Googling their questions with a few obvious keywords.
Oh, interesting, that is exactly what I was looking for, and I think I would be very interested in them. But yes, seems most recent activity is from 4 years ago.

What I thought this post was going to be about from the title:
“Avoiding doom” is a method for attaining utopia, not sitting on a tradeoff curve with utopia.