If you want to chat, message me!
LW1.0 username Manfred. PhD in condensed matter physics. I am independently thinking and writing about value learning.
If you want to chat, message me!
LW1.0 username Manfred. PhD in condensed matter physics. I am independently thinking and writing about value learning.
Is it useful to call current reward-seeking behavior “reflexive?” If you make training a little more diverse, the reflexes probably become a little more sophisticated, a little more hooked in to the activations, and the patterns of prior tokens, that track useful-to-know features of the environment.[1]
I’m strongly reminded of Dan Dennett’s writing, e.g. Eliminate The Middletoad, about how brains are also built out of such lowly “reflexes.” For technical reasons maybe there’s actually a disanalogy between the automatic reflexes of a toad and the learned behavior of… also a toad, but in the situations where toads use their (relatively meagre) learning capabilities. But that difference between hard-wired and learned parts of the biological brain seems much shakier in LLMs with all-to-all layers and gradient descent.
If “reflexive” purely means “not in the human-readable semantics of CoT”, then sure. Even if earlier tokens have been shaped by co-evolution to sub-human-semantically encode a few steps of reasoning useful for the reflex, that’s inflexible compared to general language use.
But the “reflex” can, without having CoT directly talking about itself, still leverage tokens in CoT in a clever way. If the model is already doing a bunch of serial computation to deduce useful facts about the environment, collating those facts for use by the “reflex” can be a parallel step that doesn’t require CoT.
The boundary between “sub-human-semantics” nudges to the CoT and human-readable ones might also be fuzzy—both directly via increased nudge strength in some contexts, and because meta-level language (and the skills associated with using it) might be able to recruit “reflexive nudges” into “reasoning” without significant change to the reflexes themselves.
Maybe “reflexive” has to mean “Right now I can pretty much understand and control this cause of the LLM’s behavior,” even as we’re already in the grey area where more generality and cleverness might gradually lead to less understanding and control.
And then the activations that track the environment get a little better at supporting the reflexes, as do the prior tokens if the credit assignment “travels back in time” as in GRPO et al.
But then there is still no consistent way to formalize your epistemic state about the coin’s outcome as a probability distribution about the coin’s outcome, even though there clearly is a rational epistemic state to have about the coin’s outcome.
If there’s a natural way to condense your information about the coin into a single probability, you can do it just as well starting from a distribution over UDT-style universes. Like if you think it should be 50⁄50 because that’s your best guess if you forget the information about your precise action, you can give yourself a low-information distribution over policies and then marginalize over it to see what happens to the coin.
Makes sense that probabilities on only outcomes or biases aren’t rich enough to imagine Omega messing with you, but is there some slightly richer probabilistic model that works fine?
E.g. if Omega is predicting you before you even choose and rigging the coin, maybe your hypotheses need to be UDT-style universes that take the “you” program as an input. Learning that Omega is messing with you could be done by updating to place a higher probability on some universes rather than others.
On the one hand, hypotheses about the entire universe are much more extravagant than hypotheses about a single binary variable. But I don’t think it’s crazy that in order to imagine Omega messing with you, you can’t think about the coin in perfect isolation.
Definitely a nice start. Another interesting question is how sensitive the results are to variation in the prompt, both semantically meaningless ‘noise’ and semantically meaningful but minor variations in instructions.
This was really well-written!
XtORtion?
I’d love to believe that this supports various stories about how AIs can be sycophantic (and have other more dangerous behaviors) by shifting their interpretations of words, so that their conception of the AI self-character continues to use words humans would call good even while they do bad stuff.
I’d also love to believe that this method tells us accurately what the AI’s actual change in the conception of the self-character is when it’s doing different sorts of behaviors.
But given the messiness, I kind of feel like the biggest updates I should do are about the amount of off-target semantics in steering vectors, and the difficulty of training neologisms to be “described faithfully” (whatever that means). You have some convincing 1d projections, but how weird is your high-dimensional data?
Let’s say the strategy is the “winningest” in the relevant sense because given the AI’s model of the world and a particular notion of changing the AI’s modeled strategy and then using the model to predict the outcomes, the strategy of not paying blackmail has the best modeled outcomes (arguendo).
“Aha!”, you may say, “Picking what possibilities to model, and what modeled possibilities to care about, is basically what the word “real” does in normal language, i.e. your AI has off the bat dome some metaphysical stuff.”
And this is a good point. But I think my AI responds “But I feel like I do other stuff with the notion of “real” that I don’t do when modeling possibilities. Like, real stuff controls my expectations about what I’m actually going to see, and I can go interact with it (or can have interacted with it in my actual past), and I think I’d answer metaethical questions in pretty much the same way as CDT-bot. To me, it feels more like the modeling different counterfactual presents is more like a shorthand for modeling the reasoning of Omega (or other copies of myself or whatever), who definitely exists and is standing over there. If I thought Omega was doing different reasoning, when I considered changing strategies I’d end up modeling different states of the world. This doesn’t necessarily mean I’m considering being simulated by Omega, either (though I’d believe that if I had reason to) - my algorithm for finding the winningest strategy computes the same counterfactual no matter how Omega predicts me, as long as Omega is good at it.”
I’m gonna disagree with you and Scott Garrabrant.
I think all your anticipated metaphysical stances are too much. In the blackmail thought experiment, the correct metaphysics is not an expansive one that reifies impossible situations. The correct analysis of what’s going on is “this is a thought experiment asking about what you would do in a hypothetical situation.”
Or take a more realistic situation where the blackmailer acts with some probability. You might ask similar questions about one who doesn’t pay the blackmail: “Are they imagining this paying off in alternate realities? Or relocating their ‘selves’ to the platonic realm? Or think that they’re changing the past?” But it’s entirely possible for someone to have “normal” metaphysical views—they’re not thinking about alternate realities or platonic realms—and they simply evaluate the goodness of actions in a “weird” way—e.g. what makes actions good is that they’re part of the winningest strategy.
The metaphysics isn’t an inherent part of decision theory, it’s a (contingent) feature of humans making arguments about decision theory. Metaphysics comes in if you’re going to take an actual human (who has a mish-mash of different intuitions) and argue them into doing one thing or another—different policies will comport with and more easily be argued for with different human intuitions, many of them metaphysical.
If we don’t lock in certain current values (e.g. torture is bad), then we can’t reason about what the AGI will want to do in the future, given radical changes in technology and understanding.
If we do lock in certain current values (e.g. torture is bad), then we’re not only forestalling the possibility of future moral growth and development, but also building something that’s fragile and brittle in the face of radical changes in technology and understanding.
I’ve been thinking about something related, and my current favored solution is building AI that has a model of the self-modification landscape, and uses that to some extent when evaluating self-modifying actions. An AI that “reflexively” self-modifies in response to actual or modeled human feedback (i.e. doesn’t do this model-based evaluation thing) is relatively umoored from its starting values, but an AI that makes predictions about how this whole self-modification process will turn out in the end can think “Oh, we shouldn’t do that, it’ll lead to torture, and torture is bad.”
The trick is to not lock in it reasoning “Oh, I shouldn’t self-modify like that, it’ll lose [randomly initialized concept 6a67d]”. I think you need some combination of a decent starting point and some nonzero “deontological” motivation to defer to real/modeled human feedback.
I’m not sure where “not that hard” is supposed to go in an argument. I agree it’s “not that hard” in some sense to have your simulation call a vast alien intelligence who can wear the skin of everyone Elon talks to and invent plausible and consistent stories about having visited those countries and had economic interactions with them.
Also I’ve just realized I didn’t actually talk about anthropic reasoning and probably made a wrong implicit argument, prb gotta edit the parent.
If you are the world’s richest man, what sort of anthropic reasoning should you do?
You might think that, well, there’s about 2^33 humans, so relative to other people I should assign about 33 more bits towards the simulation hypothesis, since if there’s no simulation I’d be one of those randos. If random philosophers think it’s some 1%-50% kind of chance, from my perspective it should be a total lock, right?
But in most simulations, all humans are simulated at about the same level of fidelity. You could try to juggle real-ish humans interacting with simple-toy-model humans, but it would require huge amounts of skillful puppeteering to hide the charade, the compute budget would be lower but the computational complexity would be exorbitant.
It’s these rarer “everyone else is a puppet of some vast alien puppeteer who wanted to simulate specifically the human with the most green pieces of paper circa 2026, even though they can already predict everyone else well enough to inhabit their skin” type simulations that you should bump up by the full 33 (?) bits of evidence[1], not the general sort of simulation where the randos are just as simulated as you are.
What you think after considering the evidence is between you and your prior, but personally I don’t have those hypotheses within 33-bit striking distance of plausibility, and so even if I was the world’s richest man I would continue to find them implausible.
EDIT: Whoops, off by a thousand. Also, added a footnote expressing additional important confusion.
Thinking through what Solomonoff induction says about the actual amount of evidence is tricky, because I don’t actually know how the best self-centered theories of the world encode your location. If the best hypotheses all look like “Consider a universe [or simulation] with physical laws P, and initial state I, and zoom into the planet Earth and look for human # 4,512,905,154”, then the number of humans on earth does go straight into the number of bits.
But maybe the best hypothesis looks like “Consider a universe with physical laws P such that someone living on a planet had life experiences with the same hash value as my own experiences.” In which case the number of other humans still matters for avoiding hash collisions, but it might matter only log as much for the description length (i.e. rather than adding 33 bits, you only need log(33) bits.)
Or maybe some other description with totally different scaling is best!
I dunno man, to me this sounds like you want people writing about post-singularity utopia to talk about the abstract pointer “things people like,” but not talk about things people actually like.
I agree with this.
how the universe’s resources get used still has to be settled somehow, at least indirectly.
I also don’t think the answer is to settle disputes with games and trials.
The How is by getting superintelligent AI to resolve inter-human preference conflicts using the grown-up version of what current humans call good ethical principles. I just proposed games because I think it would be cool—the trick to making it ethical is that the game has to be designed so that all outcomes are okay (different high-weight ethical models rank all the outcomes pretty highly but disagree on the exact order, and controlling the outcome of the game itself has value to the humans).
The AI 2040 post-singularity utopia kinda sucks, so here’s my quick takes on how to imagine a better one.
The central story of a utopia should not be who owns what or what philosophers think. It should be about sex, friendship, and other actually interesting social games. A superintelligent AI that knows every person on earth has to shoot the gap between being useless for my social life and being overbearing. Fortunately for the AI’s chances, there are too many people even just within a mile of me for me to know all of them, and I’d be happy to opt in to it giving me some respectful nudges to set me up with a probable soulmate.
Returning to who owns what: giving wealthy Americans common law ownership over billions of galaxies is lame. Those galaxies should almost all be kept in trust by automated systems for later use anyhow[1]. Again, social life—who wants to beam their brainstate off to a distant galaxy if you’re not doing it with an alliance of a thousand future nations, each nation a thousand cities, one of the cities full of your personal friends? Future AI could be deciding when an alliance can claim a galaxy with a set of cool trials and competitions that everyone knows it’s gamed out so that no outcome will be even a millionth as lame as enforcing Jeff Bezos’ claim to have bought a galaxy ten billion years into the future.
Some people will find playing god-emperor fun. I’m happy for some fair fraction of the universe’s resources to go to those people, but I’d rather my slice of the pie be used for more gradual, organic, long-lived expansion of Earth-originating intelligent life. In terms of spending some share of the universe I’m “owed,” I’m happy to let superintelligent AI do what makes the universe go well, no need to pick out some specific eight-billionth to be the Charlie Sector. In terms of how to spend my time if I’m not playing god-emperor, many pleasing activities have already been invented by humans or will be invented soon, but there are too many things to do for me to know all the cool ones, so I’d opt in to the occasional superintelligent nudge in a good direction. I bet many of them will involve my current friends, making new friends, mastering skills, physical activity, beautiful surroundings, music, humor, drama, and good food.
We should really be turning most of the stars off so we can extract more computation from them when the universe is colder.
Agreed. Or at least, there are certain psychological factors[1] that make working on abstract alignment problems feel “safer” for some subset of people, myself very much included! On the other hand, these same factors can be “unsafe” to different people or in different contests, and there are plenty of features of other research styles that are psychologically attractive[2].
But I don’t really buy the implied add-on “this is psychologically safer for me, therefore it’s actually bad,” at least not as it applies to myself.
(e.g. lack of feedback loops can mean you don’t feel failure the same way, relatively low competition / professionalization, the real world doesn’t interrupt your research to tell you to change your mind very much)
e.g. you will naturally work with other people, you get to apply the neat math / CS you studied, you get to spend lots of money, your work fits in the normal paper format, etc.
This is like the inverse of “alignment is a capability, therefore alignment work is useful for capabilities”: general capability is required for alignment, therefore all general capability work is useful for alignment. I don’t think we’re obliged to stop alignment work because of unavoidable dual use, but neither should we do every single thing that advances alignment in a vacuum. We’re playing a multiplayer tech tree game.
I should probably link an old post of mine that explains what that means.
One problems with SAEs (of which the “golden gate vector” was a single feature) is that attribution and meaning are pretty hard. The attribution process involves a lot of humans looking at things and making guesses. This technique doesn’t decompose states into fine-grained features like an SAE, but it does give states simple token labels that aren’t complicated/fraught to interpret.
Of course, so does the logit lens, an older technique that this should be more directly compared to. Maybe see Neel’s commentary for why this new technique is less hacky / more likely to tell you interesting stuff about what’s going on internally.
Either way, adding features will usually do the coarse-grained thing you want but with some caveats because just adding features relies on a simplified model of the underlying mechanism. Given the limitations involved in doing this thing at all, it’s not that surprising if new techniques (wow, gradient-based attribution) don’t blow old ones out of the water.
Both seem pretty bad—“obedient” AI with undercooked value alignment just seems like a recipe for over-pursuing instrumental goals, manipulation of the user, etc.
In a distributed scenario, best case people use time to work out value alignment and actually get to a good future. Worse case concentration of power has compounding effects and we slide back into the highly concentrated scenario, after a period of power struggle that selects for bad people having power. Worst case humanity has an extremely undignified slopocalpyse and then goes extinct.
In the contentrated scenario, best case you have at least mildly prosocial dictators who make good things happen for real people. Worse case they never cared about most people anyhow, or experience value drift and get tired of the dirty masses, or they want the AI to change itself in a way that secures their power more, but in doing so screw up the half-baked value alignment that was keeping the AI non-sociopathic in its attempts to please the dictator. Worst case the AI was already sociopathic by default, or is misaligned in other ways that lead to manipulation of humans and eventual replacement of them.