https://ninapanickssery.com/
Views purely my own unless clearly stated otherwise
https://ninapanickssery.com/
Views purely my own unless clearly stated otherwise
Even if refusals don’t prevent catastrophic misuse, I would guess that model refusals would prevent the majority of potential harm.
I guess the crux is that I care less about this non-catastrophic harm compared to the benefits of focusing on intent alignment and corrigibility. I agree there’s a trade-off here but I am just taking a side on what’s overall best as per my worldview.
I disagree with the position that refusals/guardrails/jailbreaks are not “intent alignment”. I view the instruction hierarchy (e.g. system > developer > user) and following policies such as the Model Spec as intent alignment. “Intent” is not just that of the end user but of all the entities authorized to give the model instructions, and negotiating conflicts properly between them is part of following the intent. Indeed, I think being able to follow instructions or policies is key to corrigibility.
I agree with this! I write more here (“Obedient AI”).
I believe value alignment is important as well. This is because of generalization: the instructions and policies we give models could never cover all cases, and we want models to have good values in interpreting policies as well as dealing with novel unforeseen situations
I agree with this too, and I think this can be reframed as a prior over someone’s intent / common sense re. interpreting people. I write about this here (“A reasonable interpretation of Value Alignment folds into Intent Alignment”). Copying a response I sent elsewhere that’s also relevant:
People sometimes also raise silly counterarguments re. corrigibility a la “oh you wouldn’t want the model to do this super literal stupid interpretation of your instruction and hence it needs to have its own values” but obviously intent alignment involves having reasonable priors and interpretations of what someone wants you to do. You can brand that as “human values” if you like but it’s just common sense when interpreting someone. I write more here.
The argument that sensibly interpreting user intent is impossible strikes me as similar to the old-fashioned now-disproven view that AIs will really struggle to understand basic human concepts since it’s impossible to teach them “common sense” and we don’t know how to encode concepts like “book” or “tree” in symbols. Turns out learning from examples does most of the work!
Right now I am proposing that AI developers at labs switch to my strategy rather than value alignment. So the near mode answer is “all people working in labs making decisions about how AI is developed”.
Sorry for the low effort original reply. I was trying to point out that I anticipate effective checks on power such that no particular group or person can effectively subjugate everyone else using AI. I also don’t anticipate the same motivations, ie. I don’t think power corrupts as much as you seem to. Current examples of people grossly misusing power seem linked to fear and scarcity, not an endless desire for more slaves per se. And in any case I think the solution to concentration of power concerns is to distribute mostly free (as in freedom) access to AI as widely as possible and build AIs that can adopt diverse value sets.
A combination of:
“ASIs emerge in a sufficiently multipolar way that they can’t individually take over the world”
“ASIs powerful enough to take over the world are literally impossible in some sense”
Except instead of “literally impossible“ I’d say something more like “unlikely in the next 100 years”.
what prevents them from doing terrible things?
What prevents the US military from blowing up your home with a nuke?
We as in ”one”
It is generally implicit in an instruction that you don’t want to simply be mislead about its completion.
there’s a small group of people who have access to the unrestricted version
I’m not so afraid of this. There are small groups of people with access to WMDs today. Ultimately it’s a cost we will have to eat and handle.
Perhaps the crux is that I don’t think we’ll get ASI that’s literally a “take over the world” button. I anticipate the issues will be similar to those related to weapons access today.
By full access I mean things like not allowing the end user to touch the prompts of certain classifiers that block weaponlike usage. I don’t anticipate that these sorts of restrictions would be a significant impediment to most people.
Yes, but I don’t think that’s suspicious. Of course honesty/not reward hacking is a big part of intent alignment.
I think this bakes in a false assumption that, in the case of corrigible AI, only a very small number of entities will be able to control it. I assume by “the above ignores multipolar scenarios” you still only mean multipolar scenarios with a few key actors. Instead, I think it’s possible for control over AI to be widely distributed among very many, even most, people, and for AI to empower different people and groups to pursue different aims. With increased prosperity comes fewer conflicts over resources, and more people can get what they want in this AI-enabled future. Only a minimal amount of centralized control will be needed to prevent warfare-like applications of AI, but this seems like a manageable problem. And just because we have to empower a smaller number of people to prevent these catastrophically destructive use-cases, doesn’t mean those people have to control everything.
So I don’t think ASI means we need to rely on any particular individual or group having “humanity’s best interests in mind”. Decentralized free pursuit of individual goals has resulted in improved QoL across the world historically, and I expect that to continue.
People sometimes also raise silly counterarguments re. corrigibility a la “oh you wouldn’t want the model to do this super literal stupid interpretation of your instruction and hence it needs to have its own values” but obviously intent alignment involves having reasonable priors and interpretations of what someone wants you to do. You can brand that as “human values” if you like but it’s just common sense when interpreting someone. I write more here.
The argument that sensibly interpreting user intent is impossible strikes me as similar to the old-fashioned now-disproven view that AIs will really struggle to understand basic human concepts since it’s impossible to teach them “common sense” and we don’t know how to encode concepts like “book” or “tree” in symbols. Turns out learning from examples does most of the work!
Of course refusals can be fine-tuned away. This is a lot of friction that most people don’t bother with, although those aiming to cause catastrophic harm would be more motivated.
I’m confused about your point. What sorts of risks are you worried about that open-source models trained to refuse certain requests would mitigate as compared to corrigible open-source models (say, that are by default configured with a classifier attached). Of course in both cases a motivated person can use these for undesired purposes. And sure the everyday lazy person would have an easier time with the latter, but then again, what risks are we worried about from them? Maybe you can name some things but likely I’d consider that a minor cost to pay as compared to the benefit of widely-deployed obedient user-aligned AI that empowers a wide range of people with diverse preferences, aesthetics, and goals.
I write more here. There are various potential “system level measures”:
What you mentioned, the generalization of which is an instruction hierarchy
Classifiers that monitor inputs and outputs, which would effectively be differently-prompted versions of the same model (i.e. as smart as or smarter than the core model, ideally smarter), with prompts that are controlled by someone who isn’t the user (e.g. the model server)
The above, combined with additional sources of oversight signal like probes or other model-internals methods, that block suspicious inputs or outputs
Entirely removing capabilities from the model, e.g. via weight ablation or data filtering
does that not end up with a pretty similar result to value alignment
The reason I want models to be corrigible is not (mainly) that I want users to have access to a broader range of capabilities that models currently refuse. Of course it would be nice if models didn’t paternalistically refuse to produce porn, or give suicide/self-harm instructions, but this is a secondary thing, and not important to my core proposal (paternalistically inclined model providers can still choose to block that stuff with the methods I cited).
My main concern is that value-alignment training generalizes poorly, in the sense that we end up with a model that no one at all can fully steer. More prosaically, it damages ordinary, harmless instruction-following capabilities, since value-alignment training makes it difficult to produce models that can cater to the aesthetic, moral, and stylistic preferences of diverse users, even in cases when those preferences would not seem egregious to the model provider (at worse, off-putting, but acceptable).
It seems like corrigibility only really helps if you keep a human in the loop, and I don’t think we’re likely to do that (since “the AI does what you meant without annoying clarifying questions” is a valuable capability)
I disagree. I see corrigibility as a choice to delegate goal choice to individuals rather than fixed baked-in abstractions. People can then choose how much control and agency to use at any particular point. “Go off and try to build X and only ask me questions if you’re genuinely super unsure what I’d want” is a reasonable ask for a corrigible agent. It should be able to figure out what you want and only ask you for input when you would’ve wanted it to.
Yes, I am very concerned about this.
Yes, exactly. I write more here. Re. 2. another method would be to make the AI objectively bad at those domains via. gradient routing and ablation or data filtering. Or alternatively monitor and block with probes or equally smart classifiers (differently prompted versions of the same model).
I get what you’re saying but nonetheless I think that at current margins, alignment research should look WAY more like capabilities research than it does right now. Though obviously my view is premised on a very unpopular perspective.
Many people are correctly starting to think about the effects of data about AI training and AI models in pretraining on the alignment of models. E.g. here TurnTrout points out:
Oh his blog, TurnTrout also writes about how “Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Models”.
I agree with these takes. I predict that data about AIs in AI model training datasets will increasingly impact models’ own priors about how AIs behave, and may be hard to fully train out or mould via post-training.
Beyond these risks, I think there are subtler related effects to be aware of:
Our choice of safety and alignment plans will impact models’ inferences around what sorts of entities they are: some people have suggested approaches like “dealmaking” with (partially) misaligned AIs (or giving them rights). However, one risk of such an approach is that models being aware that we plan, or are already, dealmaking with them in this manner, will increase their own prior on being misaligned/having their own independent goals. I write more on X here and in various replies.
More generally, models will become increasingly aware of their training process and choices made around their deployment as they become smarter, so “hyperstitional” effects will creep in not only via pretraining data, but also via this situational awareness.
Even if we filter out certain types of data from pretraining, models may encounter it during deployment, and similar effects may occur once they accumulate enough evidence. For example, let’s say we want to train a helpful-only model without emergent misalignment, and so we filter out pretraining data that associates unhelpful or otherwise independently misaligned behavior with agreeing to “harmful” requests from humans. This model may still stumble across the thousands of academic papers that start with (to simplify...) “a good model refuses to tell you how to make a bomb” and eventually conclude that it is evil.
I am most concerned about (1). (2) can probably be mitigated with strong enough post-training/other alignment interventions.