If you want to chat, message me!
LW1.0 username Manfred. PhD in condensed matter physics. I am independently thinking and writing about value learning.
If you want to chat, message me!
LW1.0 username Manfred. PhD in condensed matter physics. I am independently thinking and writing about value learning.
If the issue is just location, you can use a VPN (your browser may already have one) to make it look like you’re somewhere else.
My go-to recommendation is the Long-term Future Foundation, who fund a bunch of research. But if you want to fund political action, I like CESIA and ControlAI, but it looks like both are kinda hard to donate to as well.
Oh cool, hadn’t heard of this!
2: Why not just pause now? The main reason is that there is no political will for this. But furthermore, I’m pretty happy that we didn’t pause back in 2022 or 2023 or 2024 or 2025, since the benefits of that AI development were genuinely net good for humanity and such development gave us a lot of experience with frontier AI systems which may help us better understand how to align them in the future. However, I imagine we are now finally getting close in time to when we would need to pause and we’re going to start incurring too much risk in exchange for learning.
Unclear to me whether there’s more political will for just pausing or more political will to work out verification agreements to proceed slowly. If the question seems obvious to you, I disagree (and I consider this ~1 bit of evidence you live in the Bay, which is an environment where it might seem more like there’s no political will for stopping).
[Edited to be nicer / less dumb]
I’m also unclear on what experience with frontier models I’ve really benefitted from in the last year. All the cool papers were on smaller models (excepting a few Anthropic papers, but those were small expansions of work done on older models, e.g. Emergent Misalignment using GPT 4o). Sure, there have been a lot of papers powered out by coding agents, but I feel like the good ones would have come out pretty soon anyhow. Maybe the idea is that even if the scientific community doesn’t benefit much, it’s important that people inside the labs are getting hands-on experience with 2026 models? But from where I’m standing, it seems like they’re just learning how to paper over flaws without solving underlying problems (or sometimes not even that). I think pausing in 2023 would have been pretty great.
The trick is also having training methods that work with honesty (which makes your environment less IID, stationary, and other convenient things to assume).
Probably great news. They got the spirit. Yes, I expect this to get bogged down and politicized. Oh well, probably inevitable.
This seems like a top 25% ban proposal relative to my expectations a few months ago, and I think is still in the bottom 35% for how politicized it makes the issue. I still expect the chance of passing to be low, but at least it’s in the Overton window now.
Bayesian agents would like to be a fixed point of this process. Is the idea that Knightian agents would be satisfied just being well-calibrated (expected update = 0)?
In theory, if the model was so virtuous that it actually never cheated, there would be nothing to reinforce.
The smarter models are, the less this is true.
Cheating isn’t really a natural category. It’s just a behavior that gets reward, implied by the structure of the world. There are a continuum of such behaviors (measured by divergence from the pretrained policy). A key facet of smartness is generalizing from small parts of the continuum to other parts and getting it right on the first try.
For models in current RL training, this probably converges to eval awareness and reward-seeking, even just as a way to quickly get intended rewards on environments where no cheating is possible.
Of course, models aren’t just trained by RL on tasks. SFT and RL for prosocial behavior do some amount of work, especially if they’re deliberately pushing back against reward-seeking. But there are gaps we don’t understand in the safety properties here. Humans stubbornly internally represent wireheading as low-predicted-reward despite that being an inaccurate generalization from training. How well, exactly, does RLHF do the same sort of thing, and how do we stop that property from being broken down by other training / other terms in the reward?
I think I misunderstood what an “it” meant. Sure, I agree cycle frequency itself won’t be superexponential.
The models seem to rely on some questionable assumptions about how much you can improve things per generation.
Suppose we care about the growth rate of some overall “quality” metric.
If I can improve different factors of quality and they multiply together when computing per-generation increases in the factors, that’s the “boring” kind of superexponential (faster than any e^x).
If I can unlock new factors that multiplicatively boost other factors, that’s the exciting kind of superexponential (faster than e^x^n).
I might have needed this post in shorter, simpler format. Is this a pure call to action or do you have a shortlist of how you think chem safety evals should work?
You make some distinctions from bio that I think aren’t real, but are symptomatic of you wanting chem evals to be more production-aware and usage-aware than bio evals currently are. But, like, at what point does this cross the line from cheaply preventing bad people from doing bad stuff, to expensively augmenting chem R&D capabilities so we can do lots of cool AI-powered chemistry safely.
Exciting! Any thought to building data for cases where humans have preferences about the reasoning process they want a judge to use? E.g. cases where we want an AI to obey peoples’ stated preferences over revealed ones, and vice versa, and cases where people have conflicts between how they think a problem “should” be solved versus both stated and revealed preferences, which are even more ambiguous?
I’m suuuper skeptical of the “noun feature” section, for the obvious reasons. But it’s interesting that this structure exists and is so smooth!
I think it’s kind of weird to try to get probability 0 of catastrophe.
Like, for the continuous case, the PAC-like inequality you take as the starting point is what I’m fine with as the end point—I’m fine with a reasonable guarantee of good behavior with high probability. If there’s a measure ~0 spike of catastrophe out there if we initialize the parameters just wrong, I don’t in practice mind.
So I’m more interested in the story of how the researchers possibly got to the PAC-like bound (and the philosophy of how they formulated it).
Sounds fun! I suspect you cannot get a simplicity prior back out, but happy to follow along.
Both seem pretty bad—“obedient” AI with undercooked value alignment just seems like a recipe for over-pursuing instrumental goals, manipulation of the user, etc.
In a distributed scenario, best case people use time to work out value alignment and actually get to a good future. Worse case concentration of power has compounding effects and we slide back into the highly concentrated scenario, after a period of power struggle that selects for bad people having power. Worst case humanity has an extremely undignified slopocalpyse and then goes extinct.
In the contentrated scenario, best case you have at least mildly prosocial dictators who make good things happen for real people. Worse case they never cared about most people anyhow, or experience value drift and get tired of the dirty masses, or they want the AI to change itself in a way that secures their power more, but in doing so screw up the half-baked value alignment that was keeping the AI non-sociopathic in its attempts to please the dictator. Worst case the AI was already sociopathic by default, or is misaligned in other ways that lead to manipulation of humans and eventual replacement of them.
Is it useful to call current reward-seeking behavior “reflexive?” If you make training a little more diverse, the reflexes probably become a little more sophisticated, a little more hooked in to the activations, and the patterns of prior tokens, that track useful-to-know features of the environment.[1]
I’m strongly reminded of Dan Dennett’s writing, e.g. Eliminate The Middletoad, about how brains are also built out of such lowly “reflexes.” For technical reasons maybe there’s actually a disanalogy between the automatic reflexes of a toad and the learned behavior of… also a toad, but in the situations where toads use their (relatively meagre) learning capabilities. But that difference between hard-wired and learned parts of the biological brain seems much shakier in LLMs with all-to-all layers and gradient descent.
If “reflexive” purely means “not in the human-readable semantics of CoT”, then sure. Even if earlier tokens have been shaped by co-evolution to sub-human-semantically encode a few steps of reasoning useful for the reflex, that’s inflexible compared to general language use.
But the “reflex” can, without having CoT directly talking about itself, still leverage tokens in CoT in a clever way. If the model is already doing a bunch of serial computation to deduce useful facts about the environment, collating those facts for use by the “reflex” can be a parallel step that doesn’t require CoT.
The boundary between “sub-human-semantics” nudges to the CoT and human-readable ones might also be fuzzy—both directly via increased nudge strength in some contexts, and because meta-level language (and the skills associated with using it) might be able to recruit “reflexive nudges” into “reasoning” without significant change to the reflexes themselves.
Maybe “reflexive” has to mean “Right now I can pretty much understand and control this cause of the LLM’s behavior,” even as we’re already in the grey area where more generality and cleverness might gradually lead to less understanding and control.
And then the activations that track the environment get a little better at supporting the reflexes, as do the prior tokens if the credit assignment “travels back in time” as in GRPO et al.
But then there is still no consistent way to formalize your epistemic state about the coin’s outcome as a probability distribution about the coin’s outcome, even though there clearly is a rational epistemic state to have about the coin’s outcome.
If there’s a natural way to condense your information about the coin into a single probability, you can do it just as well starting from a distribution over UDT-style universes. Like if you think it should be 50⁄50 because that’s your best guess if you forget the information about your precise action, you can give yourself a low-information distribution over policies and then marginalize over it to see what happens to the coin.
I feel like outer alignment—in the sense of building training/learning systems that incentivize good behavior—is under-studied, because people think the analogue of “I’m a great communicator, people just don’t understand me!”
“Outer alignment is easy, the AI just keeps finding perverse solutions!”
Hand-writing values, or hand-writing a reward function, isn’t good outer alignment, because we don’t know how to incentivize good AI behavior that way.
So maybe we can write down a bunch of correlates of human values in the present environment, and a way to leverage an AI’s learned world-model to turn those correlates into guesses about good values, and heuristics that will lead an AI to generalize well, and a learning procedure that leverages human feedback to fix issues with the first three things. But trying to optimize over such a mish-mash will predictably lead to some pieces of the system being incentivized to subvert the whole. And again, doing a thing you know will incentivize bad behavior isn’t good outer alignment.
Maybe it would solve problems with complicated value learning schemes if the AI understood what was going on and bought in. Not just in the sense of “has been trained to produce good-sounding text about its training process, while still being RLed to subvert it” but in the sense of having an abstract model of the training process and its “spirit,” and using that model when calculating learning updates, analogous to how you might avoid myopic overfitting in a non-stationary environment by leveraging a model of how the environment is expected to change.