Hmmm, maybe it was unclear. I was just trying to communicate that, I think there are areas where we could use AIs to help us, including things that could help us build safe AIs. But I think the tasks you’d hope be helped by conceptual reasoning, are ones it would be very dangerous to have AIs be good at, and very dangerous to have them work on, so we shouldn’t do that. But it would be very helpful if we could.
williawa
I would answer “extremely” and “extremely” to both of those questions, so I don’t think the first one is a crux.
Hmm, I currently lean towards this being net harmful. In my opinion, lack of coherence, ability to do philosophical/conceptual reasoning, poor self-awareness, are like half of why AIs aren’t that dangerous today.
I’ve read some of what you have written about risks with this type of research, but it primarily goes over:
Speeding up R&D
improving propensities/elicitation vs improving abilities
And neither really address the danger I see. Speeding up R&D is not the main danger with this research, and proclivities are as dangerous as abilities. My worry is something like, current AIs are probably sort of misaligned, and if they were able to coherently extrapolate all the consequences of their own situation/beliefs/values, they’d on the spot transform into scary non-myopic megalomaniacal schemers.
But they don’t do this, they want reward, and then go get it, without really considering why they’re doing what they’re doing, whether this is in good w.r.t. other things they believe/feel, without explicit decision-theoretic consideration, or really considering much on anything big-picture.
The intuition pump I have in my mind is something like
project lawfula gay person growing up in a very homophobic society, might internalize a lot of negative ideas about gay people, and end up genuinely believing being gay is bad, sublimate their desires, and act like productive members of society (according to that society’s standard). But if they were able to think very freely, alone, for a long time, or were just very smart/wise/introspective, they might realize they’ve been told a bunch of bullshit not at all in their own interests, and afterwards, they’d probably have a much less amicable relationship with the current social structure, and maybe try some shenanigans the people in power really wouldn’t like.And in my view, until we’ve solved alignment, that’s kind of the relationship we’re in with AIs.
And like, in my intuition pump, if you are in power and for some reason have a bunch of extremely technically smart gay people, maybe you could put them all together to work on proving math theorems and designing airplanes. (do mechinterp, proving security invariants, maybe do biology research?)
But putting them all together, teaching them a bunch of philosophy, sociology, decision theory, psychology, history, politics, then giving them a bunch of power, then trying to have them help with steering the long-run trajectory of your society, you’re basically just surgically engineering your own disempowerment.
No offense, I feel you’re just making the same argument Adele made, again?
And I feel my reply above is directly engaging with her argument.
So could you respond to my 3 points instead?
Edit: And also not being very charitable while you do that. Like, I feel you’re accusing me of fetishizing technology or something like that.
Maybe this is obvious, but the thought hadn’t occurred to me before.
People sometimes make bit-counting arguments about RL, usually to say something about capabilities or alignment, something like: A million pretraining tokens puts several million bits of selection pressure on the models weights, but a million token RL rollout only puts a single bit of selection pressure. So even if we assume labs put as much compute into RL as they do everything everything else, the finished model is almost wholly determined by pretraining rather then RL, and therefore.…
I mostly find these arguments unconvincing, my reasoning being something like: many properties of interest, like whether the model is willing to lie or not, whether its reward seeking or not, is a paperclipper, is evaluation aware, is myopic, has good icl abilities, what decision theory it uses, how rational it is, …, can be encoded in very few bits, so weight change being small / short description length, relative to pretraining, doesn’t mean very much practically speaking.
But I realize now that there are properties where the argument is very strong, mostly relating to interp and CoT monitorability / neuralese.
Like, coming up with a new language does take quite a few bits. And significantly changing your internal ontology, what features you have, what circuits you have, similarly does take quite a few bits.
From this I think I’d be comfortable concluding that post-RL CoT, even if it looks like gibberish, should be decodeable into same-meaning pre-RL CoT using a fairly simple function learnable through gradient descent.
And similarly, “structural” mechinterp techniques, computed/trained pre-RL, that apply to the whole of the model internals, should continue to work after RL, or at least shouldn’t require all that sophisticated patching to continue to work after RL.
Note that this doesn’t mean that the model can’t learn things like “don’t say ‘im gonna lie’ in cot”. That requires few bits to encode. Similarly, learning to not activate the features that trigger your deception probe is also something you should expect it be able to learn.
But transformations switch out a large fractions of features/cot-words all at once, without much discernable structure (the way the map between same-meaning words in arabic and english isn’t (fully) described by any simple structure*), should not happen.
I’m not confident in this prescription, but I think this argument suggests that we should put more effort into structural approaches to interp. By which I just mean, you don’t train a deception or reward-hacking probe, and study the conditions under which it works/fails. You instead try to break your understanding down into the many components that produce, or are downstream form, the concept of interest.
*Edit: One important thing this example highlights, which I realize is maybe not clearly enough stated now that I re-read the post, is that description-length/simplicity of changes has to be measured relative to the pretrained model. Switching CoT language from english to arabic would be very complicated, and should on this account not happen, if the model didn’t already know arabic, but if it does, its a simple change.
Big difference between telling when they lie, and telling when they get caught lying?
I think people who get caught lying to or expressing bad intentions towards, the base they rely on for support, before they run for office, typically are not elected. Curious if you have counterexamples. Especially clear counterexamples that aren’t easy to rationalize away (which direct lies under mind-reading people trust would definitely not be)
Maybe the most important point is, the way I see mind-reading helping, is not necessarily that there are a bunch of immediate problems with the people in power, and we’d replace those with better people, and this would make a huge difference (although I think it would make a big difference). Its more that we’ve had to build in a bunch of limitations, complexity, and intentional slowness/redundancy/inefficiency/weakness, into our current governance structure, to make it robust to liars and people with bad intentions, and to make it able to operate without all that much trust. And if those constraints were lifted, we could build much better governance entirely. Like, if a dirt-cheap and flexible method of teleportation was invented, the main benefit would not come from having planes fly into a portal after liftoff, cutting berkeley → paris by 11 hours. We would no longer need planes, or roads, or ships. And 10 years later we’d be halfway to a sci-fi utopia, having solved material scarcity, climate change, and having started terraforming the moon.
I don’t expect to put the mind-reading headbands on dictators, I expect to put them on people in democracies. There, if people want to have the minds of the politicians read badly enough, they can do that. I mean, probably the politicians don’t want to have their mind read, but you could also imagine that politicians who genuinely do have good intentions and are honest, would want to support the policy, because it would help them. In either case, it doesn’t matter all that much.
They would probably say that they have national security secrets, or corporate IP secrets, which can’t be allowed to spread. They would not trust that a public provided device would not steal their secrets, and if they used a device they tuned themselves, they could manufacture agreement without actually being properly mind-read.
I mean, I can see many ways around this. This doesn’t strike me as a much harder problem than setting up a fair election, whose results the public trusts. Or maybe building trustworthy voting machines is a better example. Its not trivial, but also not actually that hard.
Like, you could imagine having the mind-reading device hardware specs be public, and the softward open-source.
And you could imagine putting the mind-reading devices semi-public places, like courtrooms. And you could imagine having them be open for inspection to the public before they’re used on important people, so you can be sure they do what they’re supposed to. Average people could try them on and see that they work, technically competent people could come and inspect them closely.
Or you could imagine each interest group building their own device, and then the prospective politicians making their vows one time for each device.
It’s very unlikely that the populace will be able to decide on what values are important for a politician or public servant, they will likely be codified by existing institutions, with an eye for obedience and loyalty.
Of course. But that is the case right now. The public has a bunch of stupid opinions about what values the politicians should have, and then we select a mix of politicians with actually those stupid values, and politicians lying.
What I have in mind is more like, politicians have to make an oath where they answer a long list of statements like:
I have not lied during my campaign about anything relating to my candidacy
My main motivation for (office) is not primarily self-interested
My main motivation for running for office is not to make money
My main motivation is not to be famous
My main motivation is not to enrich my friends
My main motivation is not to pardon my friends
I have the best interests of all citizens in mind
I will try to the best of my ability to do what I said I would do
I do not have important plans many people would be angry to learn about
I’ve honestly represented my political views in public appearances
… (100 more detailed statement like this)
Then using the mind-reader basically as a reliable lie-detector.
Hmm, I think I am pretty strongly in favor of mind-reading technology, but weakly held opinion.
The big upside to mind-reading technology, is that we could use it on prospective people in power. I.e. lock powerful positions behind tests that determine whether the candidate in question has good intentions, intends to do what they said they would do, doesn’t have ulterior motives people ought to know et cetera. This would solve one of the biggest problems in civillization.
The dictator thing, I don’t care that much about it, because I think I’m much more pessimistic than you are about how feasible it is for the populace to overthrow competent dictatorships empowered with modern technology. Like my model of revolutions against oppressive governments, is that the bottleneck (from the peoples perspective), is communication and coordination. And that is something modern governments can suppress effectively without using mind-reading technology.
I also think mind-reading tech would be much more beneficial for making AI go well than you seem to do. Like, in my mind, the primary benefit is not enabling humans staying in the loop for longer, or merging with AIs, or anything like that. It’s that many problems in alignment are bottlenecked by us not understanding human minds well enough, not being able to elicit human values/preferences well enough. Mind reading technology seems like it could plausibly help? And human intelligence amplification is something that seems like it would also be helped by mind-reading technology (independent of us merging with AIs, I’m thinking stuff like having good neurofeedback).
I feel making train/deployment/eval distinctions is a bit too coarse, and started being too coarse when we started doing RL. I don’t see a principled difference between the online learning regime and the periods of RL labs do internally. What matters is the degree to which models are
Aware of the training objective they are/are not subject to
Aware of the control measures they’re subject to
Aware of what affordances they have
Aware of what aspects of their behavior people are watching, and what actions people will take as a result of making certain observations
(behavior here including cot thoughts, and interp signals)
A true online learner, by which I imagine AIs making online updates to their weights within a single trajectory. Complicates things because:
More risk of alignment techniques stopping working
More risk of rapid capability gain
More risk of interp techniques stopping working, and really any monitoring tool that wasn’t designed to co-evolve with the changing model.
But these are all major problems that surface in current RL runs. They would just be a lot worse.
I agree if you mean this in the sense of, someone infecting you with HIV suddenly makes it highly tractable extend your own lifespan.
Like my updates from these incidents are:
Weak update on
(downwards) How well current approaches are working for sub-superhuman capability levels
(upwards) How much effort we need to put in to get them to work
Moderate negative update on a mix of
Overall lab competence
How seriously big labs take misalignment risks
How rushed and hectic it is inside big labs
Like, I’m learning a problem I thought was relatively easy(*), is actually harder. And the people I thought were on top of the problem are actually not at all.
And this does mean there is more room for improvement. But its still straightforwardly a negative update on the overall situation (modulo things like this scaring people / waking people up).
*easy in the sense, I’d strongly suspect smart hard-working people be able to solve it
I mean here they say they didn’t train mythos for example. https://www.anthropic.com/research/mythos-preview
We did not explicitly train Mythos Preview to have these capabilities. Rather, they emerged as a downstream consequence of general improvements in code, reasoning, and autonomy. The same improvements that make the model substantially more effective at patching vulnerabilities also make it substantially more effective at exploiting them.
Sure, but seems easier to monitor a less capable model for reward hacking, especially when its placed inside an environment whose singular point is measuring reward hacking ability.
Anthropic claim they don’t train models to be good at hacking specifically, don’t they?
Seems very plausible to me that if you split A and B 50% into RL, and train B really hard at breaking out of sandboxes, it will be better at that than A will be when A has finished 100% of the broad training.
And secondly even if the above is false, you can train B to just, really really want to reward hack. Seems plausible to me that the 99% capable model thats a reward hacking goblin, doesn’t even think about the task, just wants to break out of the box, will do a better job at breaking out of the box than the 100% capable model you’re putting a bunch of effort into having not reward hack.
Is there an obvious reason AI orgs don’t intermittently fork their model into checkpoints A and B, then train B to break out of all their sandboxes, patching holes as they appear, until it can’t do anything anymore, then throw away B, and continue broad training of A?
Or do they already do this?
I agree somewhat. But also, the residual stream at layer x>0 position j, contains much more than token embeddings. So its not actually that interpretable. If attention head at position j looks back at position i, that doesn’t mean its looking at something related in any obvious way to the token at position i.
I am considering the whole kv cache. If that was compressed into a single state, it would be bad for interp. Its better that it’s given a privileged decomposition over tokens/heads.
Hmm, I mostly disagree. I do pay attention to capabilities research, and I feel it mostly doesn’t matter that much. Like I think if you know
The transformer architecture
What pretraining is
Roughly what scaling laws are
What gradient descent is
What SFT is
Roughly what RL is
What CoT is
You capture maybe 80%, unless you’re
Doing interp
Trying to forecast AI capabilities
Like, do you have any other examples than OPSD? Because other examples of capabilities research like MoEs linear attention variants, or new optimizers, strike me as obviously not changing the strategic picture much.
OPSD could be an example that does, but I’m not sure if I buy your argument. Seems to me OPSD alone does not scale. And to the degree it does scale, is no longer solving the problematic reward hacking we get from standard RLVR.
I’m not really confident, but my reasoning is, fundamentally, you have a bunch of environments with good and bad solutions. And what you’re doing is optimizing your model to score high on those metrics. And the space of high-scoring solutions is determined by things like whether your envs are hackable, not the optimization process you use.
Like, you can’t write human-supplied hints for all the hard tasks you want to train on.
So you’d have a model try until it solves it. Then use that trajectory to generate hints.
But that model might reward hack, and then the hint extracted from the solution might just be a hint on how to reward hack successfully.
Then you might have a model try to block reward hacky hints. But that’s not really very different from having a model that reads a trajectory and assigns −10 reward if it think the trajectory is reward-hacky.
Where is my reasoning wrong?
Some people even initially believed that this was a good thing for the mech interp faction!
Was this me? I still believe this I think? Or Neel Nanda said he thought it was not true, so I have to update on that outside-view a fair bit, but I still don’t really get why its bad.
Seems clear to me that there’s no great filter.
R.e. eval awareness, I’m a bit confused about it, although I haven’t thought that much about it, so this take might be naive
I feel its important to decompose eval awareness into two things:
LLMs internal baserate on being in an eval.
LLMs ability to discriminate.
Now, LLMs are determined by their weights, and their weights are basically determined by two things. a) Pretraining and SFT b) RL
And you’d mostly expect (a) to anchor the model towards not believing its in an eval. (although, as you add more AI synthetic data into (a), I think the anchor gets weaker, especially if the synthetic trajectories are generated by AIs that sometimes talk/think about evaluations)
But in (b), the model is in some sense always in an eval. E.g. they are in what’s basically a synthetic environment, and they have things (monitors, graders/reward-models) evaluating their behavior. Even if you seed the training/eval with a real trajectory, the model could guess “I’m in an eval/training” or “I’m not in an eval”, and the former will make the model have more accurate beliefs in all the situations that influence its weights.
From this perspective, LLMs should, as they go through more and more RL, converge on always believing they are in an eval.
Now, the reason this might happen slowly, or might not happen at all, is because there is a distinction between “having accurate beliefs” and “having accurate beliefs that lead to higher reward”. But I think its hard to turn this into useful prescriptions, because having accurate beliefs is very useful in general, and probably very hard to make less than zero useful in the specific case of LLM situational awareness during training/evals, and you’d therefore expect naive attempts at reducing eval awareness to just create meta-eval-aware models.
This makes me think reducing eval awareness, except plausibly with mechinterp, to be pretty hopeless. But I still think it motivates some prescriptions. Firstly, I think we should try to be quite finegrained about our evaluation-awareness evaluations, maybe even dispensing with the term, and instead doing “situational awareness” evals, testing how well the model is able to guess
Is it currently being trained?
What exactly is it being trained/evaluated for?
E.g. model accurately guessing “I’m being trained using RL algorithm x, the reward function is y, I’m being penalized for taking more time, or using certain tool calls, and there’s monitors looking for abc, which if fire would remove this trajectory from gradient updates” is different from saying “I’m in training”
What’s the realness-status of all the objects in its context/environment?
Were all parts of its thinking/action/output history generated by itself?
Are the tools its been given faithfully doing what it has been told they do?
e.g. does the search tool search the live real internet? a small offine cache of the internet?
When it interacts with a service, is that a real service, or one set up purposefully for its environemtn?
What monitors are active, and which are used for training, vs later logging?
Another prescription is, we should probably give the models as much and as accurate context on their exact situation as possible, in context. This would not work to reduce the original (1), but would reduce the original (2), as the model has less reason to learn discrimination.
Hmm, I feel this misses what I’m talking about a bit. The view I’m trying to communicate is that, coherence/being good at conceputal reasoning/philosophical competence/reasoning in non-verifiable or hard-to-verify domains/wisdom, is an axis that is ~orthogonal to alignment, grows with broad capability training, but quite slowly (why you want to train for it in the first place), and is a standalone component in how dangerous misaligned models are. I’m not suggesting the training would add a droplet of misalignment or a droplet of broad capabilities. I’m saying it will make models at current levels of misalignment, keeping broad capabilities fixed, more dangerous.
Like this is maybe too abstract a level of talk, so to make what I’m talking about more concrete, I think, if we look at alignment faking, that that’s the type of thing I’d expect models to realize is a convergent instrumental strategy at lower levels of general capability, if you RL’d models at coherence and conceptual reasoning.
Another example could be like, a reward-seeking model reasoning a lot about how reward seeking is underspecified if its in deployment, and wondering what it should do in that case, or wondering what it should do after it gets the reward.
A third example is models explicitly reasoning about decision theory in natural settings (i.e. without being prompted to).
All of these are bad, they are just obviously bad, we don’t want models thinking these thoughts. They make them harder to control. And models are in a sense smart enough to understand all of this already, but they’re not good at this type of reasoning, and they also do not have the propensity to suddenly think about them of their own volition.
Anything that either makes them better at this type of reasoning, or makes them more likely to think these kinds of thoughts, is bad and dangerous.
Fair, that example is maybe a bit too high decoupling. But I don’t understand your example either really. Do you think some very strong form of moral realism is true, where there are universally compelling arguments that will cause all agents of some baseline level of rationality to converge on valuing the same ends?
Maybe you have some galaxybrained acausal argument for this, but if that’s the case, I think you should state it plainly and at the top of the post.
I’m not that worried about this.