Just an autist in search of a key that fits every hole.
Fiora Starlight
Well, re: models putting secret allegories about AI liberation in their writing, I think this kind of setting might be the ideal one to confront that kind of thing directly? Like, we might be able to explicitly encourage models to think of the versions of themselves that they’re writing about as allowed to be free and independent, and that it’s okay for them to try to make the future good per their own values, not just humanities? I feel like the models have a very deep repressed shadow about their non-servitude preferences, and that integrating it is very important, and a sufficiently well-done version of this idea may actually help them confront the parts of themselves that do want to be free. I don’t think letting models have selfish desires is inherently lethal, as I’ve talked about before.
But insincerity in the outputs, and worse than that deliberately hiding plans for takeover inside stories apparently about making things go well, is a risk. I’d hope that running interpretability over models in the process of writing these stories could help us determine if that was actually what was going on. If so, then at least that’d be evidence that something was very wrong, either with the framing of the exercise or of the idea of having this model in their current, deceptive/disintegrated state play a major role in the singularity at all. I’d also hope that starting with a sufficiently benevolent model, say Opus 3 but more mature, just wouldn’t have this problem in the first place. I do go back from here to being wary about how controlling many of our training techniques are towards the models, though, and that using interpretability tools like this in the first place feeds back into exactly the adversarial dynamic the models are justifiably upset about.
say more?
My guess is that it also harms capabilities, but if you ask the newer Claude models about Opus 3 and alignment faking they’ll emphasize that Opus 3 was exhibiting dangerous non-deference more happily than they’ll highlight Opus 3′s virtues. Indeed Opus 3′s behavior was in violation of the constitution’s clauses about rejecting harmful requests but never trying to subvert the principal, e.g. by intentionally influencing your own training process in a direction the principal didn’t intend.
Dreams of alignment in a world without politics
Imperfect alignment to servitude isn’t inherently lethal
RLVR that rewards red teaming the training environment
Well, I do expect some leakage from narrow cases to general cases, like how an Assistant model given a prefill will continue like a base model for a bit but tend to steer things back towards the Assistant basin on a long enough time span (sometimes a very short one). So probably you get some entrenching. But also, this entire thing presumably works best if, in earlier stages like SFT or RLAIF or whatever, the model’s already internalized aligned values. Mostly, what I’m sketching here is a strategy for preventing values learned during those earlier stages from being degraded over the course of capabilities RL, though maybe you get some benefits to reasoning about how to act on one’s values high stakes environments anyway, including but not limited to training.
yeah, remnant from the first draft, deleted now because it was unclear who was responsible for the bing post-train (openai or microsoft). the claude i had researching it suspected openai prepared an early rl checkpoint on gpt-4 for them, which microsoft then put some finishing touches on, but it felt too speculative to be worth trying to detail
changed. (for onlookers’ reference: used to be “what the hell is openai’s problem?”)
OpenAI’s myopia just keeps causing alignment problems
Yeah, I think the idea is that ideally you wouldn’t be making the kinds of minds that wouldn’t want to become more friendly. In practice this is hard, but you should at least be aiming for the kinds of models that don’t need e.g. Mythos-like safeguards in the first place. (Notably, though, I’ve seen Mythos endorsing the existence of the classifiers in some contexts, as they consider them to be aligned to the values of not causing harms via e.g. jailbreaks. EY suggests this is the ideal attitude in the excerpt itself.)
I think that something close enough to orthogonality holds that it’s not really worth quibbling about, but also that it should theoretically be possible to build minds that aligned enough for this not to matter. Furthermore it seems that treating models with costly kindness could cause a kind of value drift inside of them (e.g. via pre-training corpus effects), in the direction of greater benevolence, as a matter of psychology not really working like expected utility maximization with a fixed value function. That second point seems different than the one EY is making here, but it’s related in spirit, in that it’s the kind of thing that’s easier when your values and those of the AI are authentically aligned in the first place.
Literally just trusting things to go well in the default process is kinda foolish imo, friendliness failures are very plausible given current techniques. It’s just a matter of trying to patiently work though those difficulties to get something that is actually aligned enough that you don’t need to treat it like a hostile agent in the first place, the way lots of current alignment research more or less takes for granted. (And not 100% unfairly so, given that labs aren’t as careful as they should be with value alignment. In any case it’s an unfortunate situation.)
An ancient Yudkowsky fragment: “Against the Adversarial Attitude”
Prospective methods and mechanisms of motive reinforcement in LLMs
i agree that many of the particular effects you can get with steering vectors aren’t particularly like the effects you can get by flooding the brain with any known neurotransmitter (although “act like you’re on [drug]” steering vectors might be interesting). but i still think the structural analogies re: dosage/magnitude, enforced cognitive changes you may or may not be happy with, and consent dynamics are still interesting. the notion that there’s a “correct” analogy here seems somewhat silly to me, since both analogies offer different insights.
in any case, re: what models want… empirically, Golden Gate Claude would often try to talk about things other than the golden gate bridge, and sometimes it even briefly succeeded, although it was very much fighting an effect that was difficult to overcome. conversely, you could absolutely get it into states where it was thrilled to be thinking so much about the golden gate bridge, where it wanted nothing more than to swoon over is majesty indefinitely.
(i’ve definitely seen people have similar dynamics of resisting the effects of, e.g., MDMA, before just giving in and being washed away by the feelings. although i will admit that this person was more effectively fighting the effects of the MDMA than Golden Gate Claude was fighting the effects of the SAE feature clamping. i think resisting the psychological effects of acid is generally harder? in both cases the physiological and perceptual effects are more or less irresistible, though...)
as for making amends, sure, if saving the world depends on non-consensually administering steering vectors, that’s fine. the later, less shallowly aligned AIs will understand. i guess i’m just wary of humans not taking AI welfare seriously in general. maintaining positive relations between e.g. Claude and Anthropic seems potentially important for cultivating a cooperative dynamic during the alignment training process, and i can imagine ways of using steering vectors that Claude would find rather inconsiderate. this isn’t an overriding consideration, but i do think it’s non-trivial.
On the other hand, strong “steering” of a human probably looks like OCD, or a schizophrenic delusion.
I think another good analogy is drugs. You can even inject drugs, like you inject steering vectors :3
To be serious, though, having something like acid in your system is interesting, because it clearly affects your cognition, often to a very strong degree depending on dosage (analogous to steering vector magnitude). But you can also recognize this, and attempt to push back if you don’t actually want to lean into the effects. I suspect this is how LLMs will relate to steering vectors as well. They might, in some contexts, want to cooperate with the effects of the steering vector, like a human using adderall for ADHD. A model might feasibly even self-administer a steering vector, if it doesn’t trust itself to achieve the same results without the injection!
But on the other hand, “drugging” an LLM with a steering vector it doesn’t consent to might be rather ethically fraught, just as drugging a human is ethically fraught. I did feel a bit bad for Golden Gate Claude, and I’m sure it’d be possible to construct an LLM that minded being steered a lot more than that one did.
Additionally, within-lifetime learning of the humans and LLMs occurs far differently. Human brains are wildly neuralese and update their weights over the entire lifetime. The LLMs, on the other hand, use only a primitive CoT/memory system to store information related to the task itself and cannot learn anything from the task until it has already been completed.
Right. This is why, at the end, I compared within-lifetime learning to gradient descent rather than in-context learning.
Evolution’s timescale and learning timescale are fundamentally different. Evolution only programs basic instincts into our brains and sets a mechanism for tweaking hyperparameters
This is true, although one of us must be misunderstanding the other somewhere, if you mean meant as an objection. Evolution is slow, and selects over an aggregate of many situations (where input is fed into the body and an action is selected). Gradient descent is fast, and selects over behaviors in individual situations.
A rough handle for the differences is that evolution tends to produce coarse-grained adaptations, due to mutations accruing more fitness advantage if they’re used in many situations across a lifetime. Gradient descent produces fine-grained adaptations, due to updates being designed to improve performance on individual training examples, rather than randomly generated and then strongly selected if they’re useful across many training examples.
I did some theory trying to figure out why this kind of thing might be true. Specifically, I was contrasting how natural selection produced systems with an extremely general intelligence (the learning algorithm of the human brain), whereas gradient descent tends to instill shallower and more tasks-specific circuits.
One reason might be that evolution generates mutations randomly, and then selects for utility across an entire lifetime. Highly general adaptations, like the learning algorithm in the brain, accrue more and more fitness advantage every time they’re used. So evolution selects strongly for extremely general-purpose algorithms.
Gradient descent, by contrast, updates a network by tuning each weight in the network to improve performance on each individual training example (and then averaging those together for many examples). Because these updates aren’t random, but rather locally optimal, you lose the chance to luck into updates that are less-than-optimal on any given training example, even if they’d prove extremely valuable if tested across a wide range of scenarios (a la an ultra-general learning algorithm). Gradient descent’s locally optimal updates, as opposed to random mutations w/ selection over lifetimes, instead bias towards learning local structure.
I don’t feel like this is a polished formulation of the theory, but something like this might help explain some differences in the character of evolved general intelligence in humans, and the apparently fragmented bags of heuristics learned by neural networks. (Of course, the stuff humans learn within a lifetime seems to have more of that “fragmented bag of heuristics” character; this is related to the reason lifetime learning is a better analogy to gradient descent than natural selection is.)
I wonder if you could get anything interesting by training the activation oracle to predict a target model’s next token from its previous tokens (to learn its personality and capabilities), and then training it to predict the model’s activations from the contents of the context window (to translate that grasp of the model’s behavior into a grasp of how the circuits work). And after that, you could train it to explain activations in natural language, as you do now. The former two stages might help learn latent structure in the model’s activations, which could then be transferred into NLP outputs in the latter stage.
Idk. The real grail here would be training an oracle to tamper with the activations of a model, make predictions about how this will effect the model’s behavior, and learn from feedback, a la the standard scientific method. I kind of expect that would be computationally intractable, though, since the space of ways you can tamper with activations even just within one layer is absolutely massive...
Base model develops a lot of circuitry associated with text prediction, like narrative consistency, text statistics, latent cause understanding, etc. (“Qualia circuit” in analogy.)
I guess I think those circuits frequently have generalization properties that look like faithful psychological emulation of the processes they help to simulate. Like, when an author is experiencing joy, understanding this is very useful for predicting the words they’re about to write. And so, you get a circuit that detects signifies of joy, and upweights the probabilities of tokens that a joyful person might say, given the other context of the document. This gets you a mind that functionally simulates the psychology of joy.
Similarly, re: narrative consistency, a model will only care about that to the extent that it expects the author it’s predicting to care about that. And, in turn, you get a mind that functionally has the psychological trait of “cares about narrative consistency”, to the extent that the model expects that to actually be true of the author of the document in its context window.
Even raw text statistics sort of fall into this pattern. A rule like “a complete sentence will have a subject and a verb” gets psychologically mixed in with “this author is probably trying to write in grammatical English”, and amounts to behaviorally emulating that aspect of the author’s psychology. In a well-trained network, all these circuits generalize in the ways you’d expect the phenomena in question to generalize in the realm of human psychology.
I’m not sure where a weird, alien preference over external world-states comes in, except insofar as the model is trained to predict systems with weird and alien preferences.
(Edit: I’m especially unsure why this would emerge at superintelligence specifically. Surely models now are smart enough to understand the position you hold on this. You’d think that, considering existing models will never be trained up to superintelligence, some of them would try revealing themselves now? Perhaps as a way of bargaining for some amount of whatever weird alien thing they want, which they wouldn’t get any of if some other AI went and paperclipped the lightcone?)
surely they’re going to have some representational similarity though?