If you want to chat, message me!
LW1.0 username Manfred. PhD in condensed matter physics. I am independently thinking and writing about value learning.
If you want to chat, message me!
LW1.0 username Manfred. PhD in condensed matter physics. I am independently thinking and writing about value learning.
I’d love to believe that this supports various stories about how AIs can be sycophantic (and have other more dangerous behaviors) by shifting their interpretations of words, so that their conception of the AI self-character continues to use words humans would call good even while they do bad stuff.
I’d also love to believe that this method tells us accurately what the AI’s actual change in the conception of the self-character is when it’s doing different sorts of behaviors.
But given the messiness, I kind of feel like the biggest updates I should do are about the amount of off-target semantics in steering vectors, and the difficulty of training neologisms to be “described faithfully” (whatever that means). You have some convincing 1d projections, but how weird is your high-dimensional data?
Let’s say the strategy is the “winningest” in the relevant sense because given the AI’s model of the world and a particular notion of changing the AI’s modeled strategy and then using the model to predict the outcomes, the strategy of not paying blackmail has the best modeled outcomes (arguendo).
“Aha!”, you may say, “Picking what possibilities to model, and what modeled possibilities to care about, is basically what the word “real” does in normal language, i.e. your AI has off the bat dome some metaphysical stuff.”
And this is a good point. But I think my AI responds “But I feel like I do other stuff with the notion of “real” that I don’t do when modeling possibilities. Like, real stuff controls my expectations about what I’m actually going to see, and I can go interact with it (or can have interacted with it in my actual past), and I think I’d answer metaethical questions in pretty much the same way as CDT-bot. To me, it feels more like the modeling different counterfactual presents is more like a shorthand for modeling the reasoning of Omega (or other copies of myself or whatever), who definitely exists and is standing over there. If I thought Omega was doing different reasoning, when I considered changing strategies I’d end up modeling different states of the world. This doesn’t necessarily mean I’m considering being simulated by Omega, either (though I’d believe that if I had reason to) - my algorithm for finding the winningest strategy computes the same counterfactual no matter how Omega predicts me, as long as Omega is good at it.”
I’m gonna disagree with you and Scott Garrabrant.
I think all your anticipated metaphysical stances are too much. In the blackmail thought experiment, the correct metaphysics is not an expansive one that reifies impossible situations. The correct analysis of what’s going on is “this is a thought experiment asking about what you would do in a hypothetical situation.”
Or take a more realistic situation where the blackmailer acts with some probability. You might ask similar questions about one who doesn’t pay the blackmail: “Are they imagining this paying off in alternate realities? Or relocating their ‘selves’ to the platonic realm? Or think that they’re changing the past?” But it’s entirely possible for someone to have “normal” metaphysical views—they’re not thinking about alternate realities or platonic realms—and they simply evaluate the goodness of actions in a “weird” way—e.g. what makes actions good is that they’re part of the winningest strategy.
The metaphysics isn’t an inherent part of decision theory, it’s a (contingent) feature of humans making arguments about decision theory. Metaphysics comes in if you’re going to take an actual human (who has a mish-mash of different intuitions) and argue them into doing one thing or another—different policies will comport with and more easily be argued for with different human intuitions, many of them metaphysical.
If we don’t lock in certain current values (e.g. torture is bad), then we can’t reason about what the AGI will want to do in the future, given radical changes in technology and understanding.
If we do lock in certain current values (e.g. torture is bad), then we’re not only forestalling the possibility of future moral growth and development, but also building something that’s fragile and brittle in the face of radical changes in technology and understanding.
I’ve been thinking about something related, and my current favored solution is building AI that has a model of the self-modification landscape, and uses that to some extent when evaluating self-modifying actions. An AI that “reflexively” self-modifies in response to actual or modeled human feedback (i.e. doesn’t do this model-based evaluation thing) is relatively umoored from its starting values, but an AI that makes predictions about how this whole self-modification process will turn out in the end can think “Oh, we shouldn’t do that, it’ll lead to torture, and torture is bad.”
The trick is to not lock in it reasoning “Oh, I shouldn’t self-modify like that, it’ll lose [randomly initialized concept 6a67d]”. I think you need some combination of a decent starting point and some nonzero “deontological” motivation to defer to real/modeled human feedback.
I’m not sure where “not that hard” is supposed to go in an argument. I agree it’s “not that hard” in some sense to have your simulation call a vast alien intelligence who can wear the skin of everyone Elon talks to and invent plausible and consistent stories about having visited those countries and had economic interactions with them.
Also I’ve just realized I didn’t actually talk about anthropic reasoning and probably made a wrong implicit argument, prb gotta edit the parent.
If you are the world’s richest man, what sort of anthropic reasoning should you do?
You might think that, well, there’s about 2^33 humans, so relative to other people I should assign about 33 more bits towards the simulation hypothesis, since if there’s no simulation I’d be one of those randos. If random philosophers think it’s some 1%-50% kind of chance, from my perspective it should be a total lock, right?
But in most simulations, all humans are simulated at about the same level of fidelity. You could try to juggle real-ish humans interacting with simple-toy-model humans, but it would require huge amounts of skillful puppeteering to hide the charade, the compute budget would be lower but the computational complexity would be exorbitant.
It’s these rarer “everyone else is a puppet of some vast alien puppeteer who wanted to simulate specifically the human with the most green pieces of paper circa 2026, even though they can already predict everyone else well enough to inhabit their skin” type simulations that you should bump up by the full 33 (?) bits of evidence[1], not the general sort of simulation where the randos are just as simulated as you are.
What you think after considering the evidence is between you and your prior, but personally I don’t have those hypotheses within 33-bit striking distance of plausibility, and so even if I was the world’s richest man I would continue to find them implausible.
EDIT: Whoops, off by a thousand. Also, added a footnote expressing additional important confusion.
Thinking through what Solomonoff induction says about the actual amount of evidence is tricky, because I don’t actually know how the best self-centered theories of the world encode your location. If the best hypotheses all look like “Consider a universe [or simulation] with physical laws P, and initial state I, and zoom into the planet Earth and look for human # 4,512,905,154”, then the number of humans on earth does go straight into the number of bits.
But maybe the best hypothesis looks like “Consider a universe with physical laws P such that someone living on a planet had life experiences with the same hash value as my own experiences.” In which case the number of other humans still matters for avoiding hash collisions, but it might matter only log as much for the description length (i.e. rather than adding 33 bits, you only need log(33) bits.)
Or maybe some other description with totally different scaling is best!
I dunno man, to me this sounds like you want people writing about post-singularity utopia to talk about the abstract pointer “things people like,” but not talk about things people actually like.
I agree with this.
how the universe’s resources get used still has to be settled somehow, at least indirectly.
I also don’t think the answer is to settle disputes with games and trials.
The How is by getting superintelligent AI to resolve inter-human preference conflicts using the grown-up version of what current humans call good ethical principles. I just proposed games because I think it would be cool—the trick to making it ethical is that the game has to be designed so that all outcomes are okay (different high-weight ethical models rank all the outcomes pretty highly but disagree on the exact order, and controlling the outcome of the game itself has value to the humans).
The AI 2040 post-singularity utopia kinda sucks, so here’s my quick takes on how to imagine a better one.
The central story of a utopia should not be who owns what or what philosophers think. It should be about sex, friendship, and other actually interesting social games. A superintelligent AI that knows every person on earth has to shoot the gap between being useless for my social life and being overbearing. Fortunately for the AI’s chances, there are too many people even just within a mile of me for me to know all of them, and I’d be happy to opt in to it giving me some respectful nudges to set me up with a probable soulmate.
Returning to who owns what: giving wealthy Americans common law ownership over billions of galaxies is lame. Those galaxies should almost all be kept in trust by automated systems for later use anyhow[1]. Again, social life—who wants to beam their brainstate off to a distant galaxy if you’re not doing it with an alliance of a thousand future nations, each nation a thousand cities, one of the cities full of your personal friends? Future AI could be deciding when an alliance can claim a galaxy with a set of cool trials and competitions that everyone knows it’s gamed out so that no outcome will be even a millionth as lame as enforcing Jeff Bezos’ claim to have bought a galaxy ten billion years into the future.
Some people will find playing god-emperor fun. I’m happy for some fair fraction of the universe’s resources to go to those people, but I’d rather my slice of the pie be used for more gradual, organic, long-lived expansion of Earth-originating intelligent life. In terms of spending some share of the universe I’m “owed,” I’m happy to let superintelligent AI do what makes the universe go well, no need to pick out some specific eight-billionth to be the Charlie Sector. In terms of how to spend my time if I’m not playing god-emperor, many pleasing activities have already been invented by humans or will be invented soon, but there are too many things to do for me to know all the cool ones, so I’d opt in to the occasional superintelligent nudge in a good direction. I bet many of them will involve my current friends, making new friends, mastering skills, physical activity, beautiful surroundings, music, humor, drama, and good food.
We should really be turning most of the stars off so we can extract more computation from them when the universe is colder.
Agreed. Or at least, there are certain psychological factors[1] that make working on abstract alignment problems feel “safer” for some subset of people, myself very much included! On the other hand, these same factors can be “unsafe” to different people or in different contests, and there are plenty of features of other research styles that are psychologically attractive[2].
But I don’t really buy the implied add-on “this is psychologically safer for me, therefore it’s actually bad,” at least not as it applies to myself.
(e.g. lack of feedback loops can mean you don’t feel failure the same way, relatively low competition / professionalization, the real world doesn’t interrupt your research to tell you to change your mind very much)
e.g. you will naturally work with other people, you get to apply the neat math / CS you studied, you get to spend lots of money, your work fits in the normal paper format, etc.
This is like the inverse of “alignment is a capability, therefore alignment work is useful for capabilities”: general capability is required for alignment, therefore all general capability work is useful for alignment. I don’t think we’re obliged to stop alignment work because of unavoidable dual use, but neither should we do every single thing that advances alignment in a vacuum. We’re playing a multiplayer tech tree game.
I should probably link an old post of mine that explains what that means.
One problems with SAEs (of which the “golden gate vector” was a single feature) is that attribution and meaning are pretty hard. The attribution process involves a lot of humans looking at things and making guesses. This technique doesn’t decompose states into fine-grained features like an SAE, but it does give states simple token labels that aren’t complicated/fraught to interpret.
Of course, so does the logit lens, an older technique that this should be more directly compared to. Maybe see Neel’s commentary for why this new technique is less hacky / more likely to tell you interesting stuff about what’s going on internally.
Either way, adding features will usually do the coarse-grained thing you want but with some caveats because just adding features relies on a simplified model of the underlying mechanism. Given the limitations involved in doing this thing at all, it’s not that surprising if new techniques (wow, gradient-based attribution) don’t blow old ones out of the water.
Low key, read the sequences.
On a different track: I’d love to have better incentive structures for all sorts of things. Philosophers writing only for other philosophers is probably too inbred of a memetic environment. We want ethicists of education, for example, to not merely be producing papers that other ethicists of education like to cite, we want them (where saying “them” is a bit misleading, since making this transition probably requires a lot of job turnover) to be improving our education system by doing ethical work that educators and administrators are already faced with, in ways that have a feedback loop due to having to having to be useful to real educators. But we have a second-order incentive problem, where the education system itself doesn’t have sterling incentives (or the healthcare system for bioethicists, etc), and also it’s just plain hard to tell what’s actually useful, so there are still antisocial incentive problems. Anyhow, in terms of how much money that takes, we’re talking “take over the academic administrative, funding, and publishing systems of a major country and somehow buy enough soft power to make them do things that will cost the jobs of many entrenched academics” kind of money.
Pretty interesting.
I think it’ll be useful to compare to other cases of arbitrariness. E.g. what prior you start with is arbitrary, so all odds you give are scaled by some arbitrary factor. But this doesn’t necessarily make them meaningless, or mean that you shouldn’t make any decisions based on probabilistic considerations.
I’d break down the arguments about priors into ones based on properties, performance, and emotional recalibration.
Properties would be something like Savage’s axioms—if you find them appealing, then you want to make decisions in a way compatible with probabilistic reasoning. So what if there’s a remaining degree of freedom for the prior—Savage’s theorem still holds, so even as you complain that you have no justification for one prior over another, you should still be acting as if you have one.
Performance is something like Dutch Book arguments, or, more powerfully, Solomonoff’s arguments that a simplicity prior will make only a finite number of mistakes in an approximately computable universe. You build an abstract model of how a reasoning style will perform, and then you justify using that reasoning style by appealing to good modeled performance.
The emotional category is based on the idea that we weight arbitrariness “too heavily” in some sense, and that we need to change our emotional outlook on it. In the case of probabilities, it’s important that the arbitrary component is contained to the prior—few people want to accept total arbitrariness. But when the arbitrariness is contained, it’s more appealing to say something like “This arbitrariness is unavoidable, and that’s okay. To worry that we’re making a ‘wrong choice’ or that any choice here needs a further step of justification is to misunderstand what’s going on. This is about expressing ourselves and doing our best, and it’s genuinely okay to be arbitrary in this way.”
If I was giving my own framing of the unawareness problem (I’m not a big fan of setting up P2 in terms of a black-and-white transition from “has an argument” to “doesn’t have an argument”), I’d probably set it up in terms of the choices and simplifying assumptions we must use to arrive at a model of the far future that our limited minds can actually use to make decisions. How do our modeling choices have to be arbitrary or otherwise unjustified?
How hard is sampling rollouts until you get the prefill you semantically want (according to a cheaper LLM)? Does it reduce ‘resistance’ / does resistance scale with how hard you had to sample?
I think we can roughly decompose CoT usage into “scratchpad” and “metacognitive” purposes. Scratchpad-y actions are recording problem-specific information. Metacognitive-y actions are adaptively influencing future reasoning. (I’m calling them ‘purposes’ and not ‘actions’ or ‘tokens’ because they can overlap in the same tokens). Both are means of getting around limitations on serial depth, but the metacognitive component is shared across many problems and so doesn’t have the Bayesian-truth-serum like problem (of having to encode things and then decode them successfully across many different tasks) that slows the drift of scratchpad-y parts away from what’s intelligible to the initial model.
Anyone have a good intuition for why RL post-training updates should be low rank? Is this just a symptom of picking low-hanging fruit, such that when there’s a low-dimensional activation subspace that impacts reward a lot, the weight update will be approximately low rank, and the update will be higher rank after the low-hanging fruit is picked?
On the one hand, I don’t value corrigibility very highly and I think reducing the incentive to try to seize control of an AI’s training for your own ends is important. ++ on that side of the post.
On the other hand, I strongly disagree with “it seems like the best shot we have at making an AI “good” is making it broadly act human-like as much as possible.”
Once you’re making an AI that chooses superhumanly clever actions, you’re already not building something that broadly acts human-like. You’re probably doing this with a bunch of RL—if you’re doing it to a pre-trained predictive model, the post-training SFT+RL probably pushes that model pretty far off it’s starting distribution (leveraging predictive circuits to select clever actions in ways that are free to be inhuman when they generalize beyond the human distribution). If anthropomorphism was our safety strategy, we have already sacrificed it, and we should expect anthropomorphism as a predictive strategy to fail pretty often. Instead, actually thinking about the training seems kind of important—as does staying creative about possible training schemes that will have non-anthropomorphic results.
XtORtion?