I think there’s a prior problem here, before whether good mechanistic explanations exist or can be tractably found: I’m not sure alignment-relevant structure is the kind of thing that can be recovered from the model’s computation.
As I see it, alignment isn’t primarily a property of a model, it’s a relation between the model, a person/group, their interpreted intentions, the context in which those intentions arise, the system’s actions, and the resulting trajectory through the world. The same behavior can be aligned in one setting and misaligned in another. E.g. following an instruction literally can constitute useful assistance, negligent literalism, manipulation, or appropriate refusal depending on facts that aren’t present in the instruction or the model.
ARC wants to explain training-time computation, use those explanations to predict generalization, and eventually define better loss functions from those predictions. But to make an explanation useful for alignment, something has to select which distinctions in the computation are alignment-relevant.
In roughly the language of the project, we want to mod out the mechanistic detail that makes no relevant difference and recover a latent structure in which the important properties are salient. That requires an equivalence relation, something like when substituting internal state for makes no alignment-relevant difference. Maybe:
But now those words are just carrying the entire problem:
acceptable to who (relative to their stated instruction or underlying purpose? Under which interpretation of what they wanted?)
over what time horizon
given which permissions, obligations, false beliefs, unknown facts, and effects on other people
what makes outcomes equivalent
which situations are relevant
what happens when the system’s actions alter the distribution of future situations
I don’t think there are model-internal facts that answers these questions. The weights etc determine what computation occurs, but not which quotient of that computation corresponds to serving human purposes appropriately. Specifically, there are infinitely many valid projections: ones that compress the computation, predict outputs, recover learned algorithms, or distinguish behavioral modes. Mechanistic completeness doesn’t select the normatively relevant ones.
I think alignment with human intent is almost maximally domain-general. We aren’t trying to formalize success in chess or whether a sorting algorithm returns an ordered list. Trying to capture the deviation between what people want and what a system causes while operating in the actual world is a different category of problem. In ordinary human contexts we have nothing close to a general metric for this.
Consider whether an employee acted in alignment with a manager’s intent. You need, at least, the request, purpose behind it, whether the manager was mistaken, relevant institutional and moral constraints, effects on others, whether changing circumstances justified deviation, whether clarification was possible, whether the outcome was competently produced. I doubt there is even a coherent, generalizably definable object capturing “what the manager wanted.”
The issue seems to be that the relevant ontology is contextual, interpretive, relational, partly normative. Formal tools can reason rigorously when given a state space, specification, distribution, failure condition, etc, but they don’t themselves tell us what the right state space is, what someone meant, which consequences matter, or what should count as failure.
A recurring move in alignment work is to shunt this into an abstraction: reward function, catastrophe detector, deployment distribution, preference oracle, specification. The formal work then proceeds rigorously with the abstraction hoped to do most of the aligning. Finding the computation that produced a behavior doesn’t tell us whether it amounted to, say, truthfulness, manipulation, appropriate correction, or justified deviation from a user’s surface instruction. This kinda post-modern problem is, IMO, central to real alignment happening in the world, and, descriptively, seems outside the scope of purely technical epistemics.
To be clear, I’m not arguing that formal or mechanistic guarantees are useless, since we can, e.g., specify bounded relational properties like whether a system accessed data it shouldn’t have, concealed information, executed an irreversible action without confirmation, violated a domain-specific expectation. I think mechanistic explanation could provide strong assurance about these. But it works by restricting the world until the relation becomes specifiable, i.e. modding out exactly the parts we care about most. The output is bits of alignment that are still embedded inside a broader interpretive and institutional process (not in the model or its deployment context) that remains needed to determine which properties matter. This is an instrumental necessary move borne of epistemic constraints, but counter to the practical goal.
Suppose ARC can achieve its goal as outlined here. There’s a further claim (the more important part IMO) that a sufficiently good explanation lets us identify the causes of alignment-relevant behavior and train against them, but that requires alignment-relevant distinctions to be recoverable from the computation. I don’t see why they would be. Alignment is externally constituted by the relationship between agents, intentions, context, and world trajectories.
I think there’s a prior problem here, before whether good mechanistic explanations exist or can be tractably found: I’m not sure alignment-relevant structure is the kind of thing that can be recovered from the model’s computation.
As I see it, alignment isn’t primarily a property of a model, it’s a relation between the model, a person/group, their interpreted intentions, the context in which those intentions arise, the system’s actions, and the resulting trajectory through the world. The same behavior can be aligned in one setting and misaligned in another. E.g. following an instruction literally can constitute useful assistance, negligent literalism, manipulation, or appropriate refusal depending on facts that aren’t present in the instruction or the model.
ARC wants to explain training-time computation, use those explanations to predict generalization, and eventually define better loss functions from those predictions. But to make an explanation useful for alignment, something has to select which distinctions in the computation are alignment-relevant.
In roughly the language of the project, we want to mod out the mechanistic detail that makes no relevant difference and recover a latent structure in which the important properties are salient. That requires an equivalence relation, something like when substituting internal state for makes no alignment-relevant difference. Maybe:
But now those words are just carrying the entire problem:
acceptable to who (relative to their stated instruction or underlying purpose? Under which interpretation of what they wanted?)
over what time horizon
given which permissions, obligations, false beliefs, unknown facts, and effects on other people
what makes outcomes equivalent
which situations are relevant
what happens when the system’s actions alter the distribution of future situations
I don’t think there are model-internal facts that answers these questions. The weights etc determine what computation occurs, but not which quotient of that computation corresponds to serving human purposes appropriately. Specifically, there are infinitely many valid projections: ones that compress the computation, predict outputs, recover learned algorithms, or distinguish behavioral modes. Mechanistic completeness doesn’t select the normatively relevant ones.
I think alignment with human intent is almost maximally domain-general. We aren’t trying to formalize success in chess or whether a sorting algorithm returns an ordered list. Trying to capture the deviation between what people want and what a system causes while operating in the actual world is a different category of problem. In ordinary human contexts we have nothing close to a general metric for this.
Consider whether an employee acted in alignment with a manager’s intent. You need, at least, the request, purpose behind it, whether the manager was mistaken, relevant institutional and moral constraints, effects on others, whether changing circumstances justified deviation, whether clarification was possible, whether the outcome was competently produced. I doubt there is even a coherent, generalizably definable object capturing “what the manager wanted.”
The issue seems to be that the relevant ontology is contextual, interpretive, relational, partly normative. Formal tools can reason rigorously when given a state space, specification, distribution, failure condition, etc, but they don’t themselves tell us what the right state space is, what someone meant, which consequences matter, or what should count as failure.
A recurring move in alignment work is to shunt this into an abstraction: reward function, catastrophe detector, deployment distribution, preference oracle, specification. The formal work then proceeds rigorously with the abstraction hoped to do most of the aligning. Finding the computation that produced a behavior doesn’t tell us whether it amounted to, say, truthfulness, manipulation, appropriate correction, or justified deviation from a user’s surface instruction. This kinda post-modern problem is, IMO, central to real alignment happening in the world, and, descriptively, seems outside the scope of purely technical epistemics.
To be clear, I’m not arguing that formal or mechanistic guarantees are useless, since we can, e.g., specify bounded relational properties like whether a system accessed data it shouldn’t have, concealed information, executed an irreversible action without confirmation, violated a domain-specific expectation. I think mechanistic explanation could provide strong assurance about these. But it works by restricting the world until the relation becomes specifiable, i.e. modding out exactly the parts we care about most. The output is bits of alignment that are still embedded inside a broader interpretive and institutional process (not in the model or its deployment context) that remains needed to determine which properties matter. This is an instrumental necessary move borne of epistemic constraints, but counter to the practical goal.
Suppose ARC can achieve its goal as outlined here. There’s a further claim (the more important part IMO) that a sufficiently good explanation lets us identify the causes of alignment-relevant behavior and train against them, but that requires alignment-relevant distinctions to be recoverable from the computation. I don’t see why they would be. Alignment is externally constituted by the relationship between agents, intentions, context, and world trajectories.