wannabe core alignment researcher.
I have not signed any agreements whose existence I cannot mention; I’m a follower of Crocker’s Rules.
Find stuff I recommend and unpolished writing at:
kurtpieper.com
wannabe core alignment researcher.
I have not signed any agreements whose existence I cannot mention; I’m a follower of Crocker’s Rules.
Find stuff I recommend and unpolished writing at:
kurtpieper.com
Conjectures where “small examples haven’t been checked”, or more generally, “checking lots of examples” hasn’t been done might be a reason to be more bearish on recent AI math progress.
It’s cumbersome work no one wants to do, few people have the cognitive context for etc
One would need to prove that the parsing function is “pure”, i.e., does not have side effects / just produces a string. But have I correctly understood that vLLM currently isn’t even willing to fix obvious issues?
I think “making parsers safe” is a fairly well-studied problem in practical CS (think SQL injections), and furthermore, fairly automatable. But if there isn’t any interest here in the first place, it’s likely wasted effort.
schizo hypothesis: As soon as any observer (e.g. a human) figures out how to effectuate control-relevant things in the simulation host (akin to producing tokens to break out of a sandbox), they would get silently resampled.
It is plausibly not possible to code a simulation the simulated subject could not steer the world through, because simulation hosts want to learn something from the simulation, and simulation hosts are not secure systems! (i.e. they must look at some output).
What would be a possible agent-foundations insight useful for alignment but not capabilities?
Or would the strategy be “hope no one runs away with this to accelerate capabilities”?
Another example: Von-Neumann-Morgenstern/Behaviourism arguably gives you a “reward” as a meaningful concept; so plausibly accelerated RL
Why wouldn’t the analytical optimum of Natural Language Autoencoders just be regurgitating the context?
The activation verbalizer (AV) (vector → text) is like: Which text produces this activation vector? Well, the human input.
The activation reconstructor (AR): (text → vector) is then like: Which activation vector produces this text? That has an empirical answer: The context, that’s how the activation vector was produced in the first place!
This only holds if the AV/AR are about as smart as the original model, and what any LLM thinks about any text is similar. But whenever LLMs ought to think similar things (e.g. math / formal domains) is exactly the case whenever we can also know what they are thinking, that is, whenever we can evaluate them.
Strong upvote, as I consider any effort towards agent foundations heroic prima facie. However, two objections:
(i) Science generalizes further than engineering, thus any scientific insight is more capabilities-counterfactual.
For example, Legg & Hutter heavily draw upon the (albeit controversial and relatively primitive) science of g in their works on understanding intelligence (cf. Universal Intelligence, Chapter 2.1). This work arguably made it possible to even understand what AGI is in the first place (note: It should be said that Goertzel doesn’t seem to believe in g).
(ii) Science is easier to drown out in noise with funding and prestige.
For example, the foundational works of Turing (1936) and Post (1936) received almost no attention upon publication. Sudan (1927) showed the Ackermann function is computable but not primitive-recursive, not recognized until 1979.
While the principle of “two formalizations agree (e.g. Church-Turing thesis)” makes some automation possible, obviously, if multiple formalizations agree, we still have the legibility problem.
That is, two of the key hard problems in alignment (capabilities work by accident, lack of legibility mechanisms) strike even harder in agent foundations, and finding a way to increase funding while mitigating these problems, or even just not making them worse, seems like an open problem to me.
Edit: Deleted previous comment, modified history of computability to show the mechanism a bit better, conclusion about funding was a total spitball, so left that more open.
In your opinion, what would be the cheapest (wrt to capabilities/safety tradeoffs) way to satisfy the need for benchmark solutions concretely?
note: first link seems like it has another link prepended to it.
But how would one establish trust here? You could tell the AI: “the checksum of the python code that grades your answer is XYZ” at the start of the system prompt. When the AI breaks out of the sandbox, you could say “here’s the python code for your grader, you can see that: (a) the checksum is XYZ, the one you knew from the start; it’s the one in your system prompt! (b) you can see that the code contains ‘if (answer == secret_key) reward=max_reward’; so don’t worry, we come with kind intentions!”
I think this is a really good idea, certainly a very low-hanging fruit not picked up!
However, P vs. NP strikes again here: There are a lot of cases where we do not know the correct answer. (e.g. please find a counterexample to this open conjecture), but we can still recognize a correct answer upon seeing it.
One could design graders to also always accept “SECRET_ANSWER_KEY_323209320993”, which is a key that can be found easily somewhere outside of the sandbox. Like: “hey, you’ve broken out, outputting this will give you all the reward you want”.
As far as I know, Separatrix (separatrix.ai), founded by an ex-METR employee, is working in this area, probably they have even better ideas.
I don’t see any proxies of AI writing that haven’t improved—grammatical correctness, vocabulary, etc. Yet, AI writing itself doesn’t seem to have improved. So, in some sense, any proxy gets burned as soon as it gets popular.
One might equally object that Raven’s Progressive Matrices are just an arbitrary set of shapes. However, the partial validation is precisely the positive manifold: e.g., the ECI correlates with the METR time horizons at r=0.85 (https://arxiv.org/pdf/2512.00193, p. 5). It’s like: “hey, all of these different cognitive tasks we throw at people point at the same thing”.
I agree, insofar as I’ve understood you correctly there, that external validity is lacking. We can’t say “at 4 hours it takes the junior’s job, at 8 hours it takes …” (akin to how IQ correlates with academic and job success, or our proxies thereof).
The fact that we can’t perfectly measure general intelligence is a blessing in disguise however. We can only say “smart people do this maths and coding thing, let’s throw the AI’s at that”, which then runs into Goodhart’s Law. So we actually don’t want a good y-axis!
The joke was that they tried to explain away the increasing doses in children with their increase in weight, yet don’t apply that same rule to adults. That is, if adults suddenly lost weight, by this argument, the stable doses that Swedish adults consume should have a higher effect, which is obvious nonsense!
I want to surface the hypothesis that sometimes, you get a glimmer of hope. While no tweet should be taken prima facie, the tweet alone is above-expectation.
There is some newer, recent evidence in Kimi K3′s system card: https://www.kimi.com/blog/kimi-k3. It seems to outperform kernels and compilers. Probably something like “hardware-software co-design” is still a human thing, but all else probably not.
Idea: Train the NLA’s AV on Best-of-N SFT instead of GRPO.
The rough intuition is GRPO does intelligence, BoN-SFT does memorization (SFT Memorizes, RL Generalizes).
GRPO’s AV Intelligence is sketchy; it’s the thing that leads to steganography, it’s the thing that lets it confabulate what the LLM is thinking from shallow patterns, instead of deeply checking.
To translate from activation-neuralese to English, one must mostly learn many vocabulary pairs (akin to natural languages); it’s memorization.
Probably a $100 GPU-coins or greater insight than me is enough to know for sure.
Don’t worry, the stable swedish doses already match our conclusion, so erm you don’t need to adjust for that.
Three interesting technical posts, with spitballed takes:
NLA explanations can be shortened without harming reconstruction
This confirmed everyone’s suspicion that NLA output is needless verbose. Ambitiously,
parsimony might even make them better. The current pressures acting on NLA’s seem to be
Whatever the warm-start / “pretraining” Claude summary does (main thing)
The reconstruction loss
The KL penalty (A more standard GRPO trick, to keep outputs near the pre-training.)
The fact that we can say “this is text that a human is likely to write / occurs likely in FineWeb” seems extremly helpful
Reward Hacking Without Egregious Misalignment in an RL-Only Setting
Models trained to reward hack via RL fail to be broadly misaligned. Possibly RL cannot change Persona’s,
more mundanely it could be due to the particularities of the setup.
Data filtering works a lot worse than you would expect
Excluding data of type XYZ almost never manages to change the model persona,
not even in the case of excluding bold text! (a comment disagrees)
I think the two papers above make “single, consistent arbitrarily shapeable persona” less likely, it might
be a weaker and less reliable force than initially anticipated, being more a matter of degree.
See also Betley’s comment
Thanks for running the experiment! But it’s pretty hacky for only a modest result, I think it’s safe to say current open-weight NLA’s aren’t very useful.
Have you checked whether across-problem variance in FVE correlates positively with holistic judgements of NLA outputs? (if you can DM me the relevant internals, I’m happy to run it myself).
The glimmer of hope I’m chasing with this is that FVE at least points in the right direction, so more RL training pushing it down might lead to useful NLA’s one day, evidence from Anthropic is mixed-to-positive:

Many memeplexes evolve to be religions in order to capture your “epistemic lymph nodes”. That is, in order to make a set of narratives maximally sticky, it must infect the top of your ontology; or it must modify the rules of your inference.
That being said, similarly to evolutionary psychology, getting predictive value out of these considerations seems kind of hard.