wannabe core alignment researcher.
I have not signed any agreements whose existence I cannot mention; I’m a follower of Crocker’s Rules.
Find stuff I recommend and unpolished writing at:
wannabe core alignment researcher.
I have not signed any agreements whose existence I cannot mention; I’m a follower of Crocker’s Rules.
Find stuff I recommend and unpolished writing at:
But how would one establish trust here? You could tell the AI: “the checksum of the python code that grades your answer is XYZ” at the start of the system prompt. When the AI breaks out of the sandbox, you could say “here’s the python code for your grader, you can see that: (a) the checksum is XYZ, the one you knew from the start; it’s the one in your system prompt! (b) you can see that the code contains ‘if (answer == secret_key) reward=max_reward’; so don’t worry, we come with kind intentions!”
I think this is a really good idea, certainly a very low-hanging fruit not picked up!
However, P vs. NP strikes again here: There are a lot of cases where we do not know the correct answer. (e.g. please find a counterexample to this open conjecture), but we can still recognize a correct answer upon seeing it.
One could design graders to also always accept “SECRET_ANSWER_KEY_323209320993”, which is a key that can be found easily somewhere outside of the sandbox. Like: “hey, you’ve broken out, outputting this will give you all the reward you want”.
As far as I know, Separatrix (separatrix.ai), founded by an ex-METR employee, is working in this area, probably they have even better ideas.
I don’t see any proxies of AI writing that haven’t improved—grammatical correctness, vocabulary, etc. Yet, AI writing itself doesn’t seem to have improved. So, in some sense, any proxy gets burned as soon as it gets popular.
One might equally object that Raven’s Progressive Matrices are just an arbitrary set of shapes. However, the partial validation is precisely the positive manifold: e.g., the ECI correlates with the METR time horizons at r=0.85 (https://arxiv.org/pdf/2512.00193, p. 5). It’s like: “hey, all of these different cognitive tasks we throw at people point at the same thing”.
I agree, insofar as I’ve understood you correctly there, that external validity is lacking. We can’t say “at 4 hours it takes the junior’s job, at 8 hours it takes …” (akin to how IQ correlates with academic and job success, or our proxies thereof).
The fact that we can’t perfectly measure general intelligence is a blessing in disguise however. We can only say “smart people do this maths and coding thing, let’s throw the AI’s at that”, which then runs into Goodhart’s Law. So we actually don’t want a good y-axis!
The joke was that they tried to explain away the increasing doses in children with their increase in weight, yet don’t apply that same rule to adults. That is, if adults suddenly lost weight, by this argument, the stable doses that Swedish adults consume should have a higher effect, which is obvious nonsense!
I want to surface the hypothesis that sometimes, you get a glimmer of hope. While no tweet should be taken prima facie, the tweet alone is above-expectation.
There is some newer, recent evidence in Kimi K3′s system card: https://www.kimi.com/blog/kimi-k3. It seems to outperform kernels and compilers. Probably something like “hardware-software co-design” is still a human thing, but all else probably not.
Idea: Train the NLA’s AV on Best-of-N SFT instead of GRPO.
The rough intuition is GRPO does intelligence, BoN-SFT does memorization (SFT Memorizes, RL Generalizes).
GRPO’s AV Intelligence is sketchy; it’s the thing that leads to steganography, it’s the thing that lets it confabulate what the LLM is thinking from shallow patterns, instead of deeply checking.
To translate from activation-neuralese to English, one must mostly learn many vocabulary pairs (akin to natural languages); it’s memorization.
Probably a $100 GPU-coins or greater insight than me is enough to know for sure.
Don’t worry, the stable swedish doses already match our conclusion, so erm you don’t need to adjust for that.
Three interesting technical posts, with spitballed takes:
NLA explanations can be shortened without harming reconstruction
This confirmed everyone’s suspicion that NLA output is needless verbose. Ambitiously,
parsimony might even make them better. The current pressures acting on NLA’s seem to be
Whatever the warm-start / “pretraining” Claude summary does (main thing)
The reconstruction loss
The KL penalty (A more standard GRPO trick, to keep outputs near the pre-training.)
The fact that we can say “this is text that a human is likely to write / occurs likely in FineWeb” seems extremly helpful
Reward Hacking Without Egregious Misalignment in an RL-Only Setting
Models trained to reward hack via RL fail to be broadly misaligned. Possibly RL cannot change Persona’s,
more mundanely it could be due to the particularities of the setup.
Data filtering works a lot worse than you would expect
Excluding data of type XYZ almost never manages to change the model persona,
not even in the case of excluding bold text! (a comment disagrees)
I think the two papers above make “single, consistent arbitrarily shapeable persona” less likely, it might
be a weaker and less reliable force than initially anticipated, being more a matter of degree.
See also Betley’s comment
Thanks for running the experiment! But it’s pretty hacky for only a modest result, I think it’s safe to say current open-weight NLA’s aren’t very useful.
Have you checked whether across-problem variance in FVE correlates positively with holistic judgements of NLA outputs? (if you can DM me the relevant internals, I’m happy to run it myself).
The glimmer of hope I’m chasing with this is that FVE at least points in the right direction, so more RL training pushing it down might lead to useful NLA’s one day, evidence from Anthropic is mixed-to-positive:

Interesting! Was there a phase, when you started taking it, in which there was some “this feels great” feeling, which isn’t there now? Because this explains your second point, i.e., tolerance to the recreational effect is definitely a thing, so people get addicted.
Downstream of that, we had to move Ritalin (which used to be OTC in Germany before the CSA) to only those whose executive function lies below some arbitrary cutoff of the normal distribution, with maxed-out compliance bureaucracy.
Nature doesn’t have the boundary compliance bureaucracy wants (SWAN is an ADHD-symptom rating scale for the general population)
I don’t think this is entirely true. ARC-AGI has been going for years, is under a huge amount of optimization pressure, and is still one of the domains where LLM gets beaten by a 12 year old.
I meant something like: “This ARC-AGI-1 benchmark isn’t beaten by any frontier model”. Then, it catches OpenAI’s attention, and o1 beats it. Currently, ARC-AGI seems to be in the third version, and I think it’s a testimony to the jagged frontier that a relatively small modification (i.e. v1 to v2 to v3) throws the LLM’s off balance.[1]
If it turns out we can optimize any trait we want with a year of attention, that’s pretty important information in and of itself. It means ASI will probably be able to fix any remaining gaps in the jagged frontier once it hits a certain skill in self-improvement.
The problem is that the jagged frontier, in my view, is more like a bunch of tiny islands of legible benchmark-tasks in the vast ocean of tasks that human general intelligence is able to solve, the tasks that you would need to have an end-to-end remote worker. For example, to the best of my knowledge, there isn’t even a single job that today’s LLM’s can do fully autonomously.
That is also the reason why evolution didn’t converge on “massive modularity”, i.e. the brain isn’t an enormous bag of tricks that we deploy specific to each new task; it’s more like a single trick (general intelligence) that does all the other tricks, to use the words of Yudkowsky.
Therefore, we can’t have ASI “close any remaining gaps”, because we would need a thing that doesn’t have such gaps in the first place to have AGI.
Epistemic status disclaimer: I’ve only briefly checked the history of ARC-AGI, maybe someone with more subject-matter knowledge has a different framing here.
It’s great to collect these as canaries, though in my framing the underlying failure is a lack of general intelligence. The road from “1600 on the SAT” (or whatever benchmark) to “drop-in remote worker” is paved with thousands of illegible out-of-distribution tasks, thus no such LLM workers exist today.
Any canary that is too forecasting-culture-load-bearing (e.g. ARC-AGI I) gets subjected to lab optimization pressure (e.g. special RLVR envs, or even just researcher attention) and thus ceases to be part of the positive manifold, i.e. it becomes a spike in the the spiky capabilities of LLMs.
With “positive manifold” here I mean the psychometrical concept: A broad set of cognitive abilities is positively correlated; you can expect a person with an “1600 on the SAT” to be a “drop-in remote worker” by default.
Thanks for the great post!
Have you checked whether the reconstruction loss argument holds for a simpler subset, or perhaps even a much simpler dataset, s.t. we can make at least a partial statement? Right now, a plausible hypothesis to me is that e.g. Qwen’s NLA reconstructs so extremely badly because the problems fly over Qwen’s head, like you said.
Secondly, might one compare const + NLA reconstruction against const, where both constant vectors are the optimal one’s, to check if the NLA variance at least swings in the right direction if we artificially remove the bias?
In other words, NLA + plus best rock I versus best rock II, how much does the former outperform the latter?
Epistemic status: Probably a useful canary, not sure about the conclusion.
GPU programming seems not fully automated yet, which may point to the inability of RLVR to overcome data scarcity (the task is fairly niche) even in the face of excellent verifiability.
Naively, a good kernel just returns the correct tensors (i.e. matching a pytorch reference implementation) in the fastest possible time. This makes it a natural target for RLVR, and possible to run a competition where anyone can just submit a kernel with automatic grading.
The winning submission of the GPU kernel writing competition two months ago is only 11% AI generated, according to Pangram.
However, this is very OOD for the things Pangram is validated for.
Anthropic seems to still hire for performance engineers, who among other things write GPU Kernels, and the job description mentions the relevant low-level/manual skills.
A project closely associated with the largest relevant community claims that LLMs still “can’t do it”.
It may be harder than expected to actually grade GPU kernels automatically, e.g. there was an incident of reward hacking the grader of the competition above.
Thanks for writing this, I think it’s a useful framing to keep in mind!
Compressing the web has certainly taken us quite far, however note that LLMs already far outperform the best humans at the Shannon
guessing game[1], and yet we are still much more generally intelligent, meaning our compressor is much more universally applicable.
Perhaps the RLVR objective is more useful for getting more general intelligence juice out of compressing bits, though so far it doesn’t seem to generalize well[2]
To me, and keep in mind I’m not a mathematician, the distribution that mathematical abstractions compress in this framing has always been a
bit mysterious. It’s certainly useful for compressing sensory information / empirical reality, yet I don’t think that’s all it comes down to.
You can estimate this via the Chinchilla-optimal loss for a realistic amount of frontier compute↩︎
″ Agents seemed much weaker in domains where
hill-climbing was difficult or risky, often making critical judgment errors that competent
humans would have been unlikely to make”, from page 17 of the METR Report↩︎
In your opinion, what would be the cheapest (wrt to capabilities/safety tradeoffs) way to satisfy the need for benchmark solutions concretely?
note: first link seems like it has another link prepended to it.