A natural objection to the approach of iVAIS suggested earlier[1] may be that it just replaces one proxy with another. If ordinary alignment approaches fail because models learn to optimize proxy goals rather than genuinely acquiring the intended target, why should iVAIS be any different? If a model is trained with a single scalar reward signal (virtue score, as discussed in our earlier post), why should it become virtuous rather than merely learn to pretend to be virtuous?
This post tries to answer that objection.
The core claim is that iVAIS does not treat the virtuous character as a behavioral proxy, a list of values, or a set of separable traits. It treats the target as the holistic character of an ideally virtuous agent. This changes the structure of both outer and inner alignment. The scalar virtue score does not define virtue. It functions as a directional training signal for the gradual thickening of the model’s concept of virtuous character. Likewise, hidden reward-seeking, approval-seeking, deception, or power-seeking objectives are not merely additional risks to be constrained by external rules. They are evidence that the model has failed to acquire the target character itself.
1. The Objection: Is Virtuosity Just Another Proxy?
The most serious objection to iVAIS is that once virtue is implemented through training, it may fall into the same trap of proxy optimization, which other alignment methods face: For example, while iVAIS aims to train a model to become an ideally virtuous agent through character alignment based on virtue ethics (CAVE), training requires a loss function, reward model, preference model, or other evaluative signal, and even if that signal is a scalar virtue score, the model is still optimizing a score, and then just learns to maximize the appearance of virtuous character rather than become genuinely virtuous. In that case, CAVE would merely reproduce the familiar gap between the intended target and the learned objective.
We do not have to deny the possibility of such failures. A poorly trained model may indeed learn to imitate virtuous behavior superficially. A reward model may be incomplete. Training data may be insufficient. The model may learn strategies that exploit evaluative weaknesses. These are real empirical risks. However, the conceptual structure of CAVE is different from ordinary proxy-based alignment, and the target is not a set of rules, actions, values, etc., but is the holistic character of an ideally virtuous agent. This matters because a proxy can substitute for a fixed set of actions much more easily than a unified, holistic character.
A model that merely appears honest, helpful, or corrigible in order to gain reward has not partially succeeded at acquiring the target character or becoming ideally virtuous yet. In particular, mere strategic virtue-signaling is not an almost perfect version of ideal virtue, but rather a defect in character. The target of CAVE makes it harder for a proxy to substitute for the intended target, because the intended target is not an outward behavior but the integrated character from which outward behavior arises.
2. Outer Alignment: Why the Target Is Not Alien
Outer alignment concerns the relation between the human intention and the objective specified by a reward function (or other training signals). An outer alignment failure occurs when the specified objective does not capture what we actually intend, the intended goal. Such discrepancies typically occur because the specified objectives are narrow or artificial, defined or constructed by the reward function itself. The human intention behind them is, on the other hand, much richer and more context-sensitive.
CAVE differs from such approaches. The target it specifies is not an artificially constructed objective, but is captured by an ordinary concept that the model already possesses. Current language models already possess a rich understanding of many ordinary-language concepts. They possess at least a thin concept of virtue, a virtuous agent, moral character (and individual virtues such as honesty, courage, humility, justice). They may not apply these concepts reliably. They may fail in hard cases, displaying shallow, inconsistent, or distorted understandings. But the relevant concepts are not absent.
The objective of iVAIS is therefore not to replace an ordinary concept with an artificial proxy, but to thicken an already available concept, or that of an ideally virtuous agent. This distinction is important: A model can have a thin concept of justice, courage, or honesty without yet being able to apply it well across difficult cases. Human moral education is similar. A novice may possess the concept of courage but confuse courage with rashness, or possess the concept of honesty but fail to understand how honesty interacts with other concepts such as kindness, privacy, and loyalty. The problem is not that the novice has acquired an alien target, but that the concept is not yet sufficiently thick (i.e., not sufficiently integrated, context-sensitive, etc.).
CAVE treats many model failures in the same way. Many of them are not cases in which the model has adopted a fundamentally different objective, but rather cases of lacking a sufficiently thick concept. There, misapplications are cases of incomplete understanding, not misunderstanding, in the sense of substituting a fundamentally different target for the intended target. They are rather an insufficiently sensitive application of the same target. Here, a model that gives a morally questionable judgment has not optimized for an alien objective. It instead possesses only a thin, unstable, or insufficiently integrated concept of the ideally virtuous character.
If failures are only to incomplete understanding (rather than misunderstanding), then, for CAVE, training is the gradual thickening of the same target concept rather than the correction of a wrong objective. Thus, the crucial point of outer-alignment in iVAIS is that the target is not constituted by the reward function, but is an ordinary, philosophically rich, intelligible concept that the model already partially grasps and that training aims to deepen.
Thin and Thick Concepts of Virtuous Character
A thin concept of virtuous character is abstract, schematic, and weakly discriminating. A model with such a concept may know how an ideally virtuous agent typically behaves (being honest, fair, courageous, wise, etc.). It may also know many verbal associations surrounding virtue. But this does not mean that it can reliably judge what an ideally virtuous agent would do in concrete, ambiguous, adversarial, or morally complex situations.
A thick concept is not a matter of merely possessing a verbal definition or a list of traits. It involves sensitivity to cases, contexts, trade-offs, motives, constituting patterns of practical judgment. It therefore includes an understanding of how virtues interact, how they can be distorted, and how apparently virtuous behavior can be motivated by non-virtuous purposes.
This is why CAVE is not merely “virtue alignment” in the sense of training separate virtues one by one. The aim is not to maximize individual virtues as separate objectives. The aim is to cultivate a unified character in which these virtues are integrated through phronesis (practical wisdom). As in Aristotle’s thesis of the unity of the virtues, virtues are not fully separable excellences that can be optimized independently. Individual virtues matter, but they matter as observable aspects of the holistic character of a virtuous person.
At this point, it is worth distinguishing two models. So far, the “model” has meant the reward model, which is the evaluator, possessing the (thick) concept of virtuous character, and its role is to judge the policy model’s responses, not to become virtuous: It itself need not possess a fully virtuous character. The policy model is the system being cultivated: it is the agent whose character is gradually thickened until it acquires the character of an ideally virtuous agent.
The reward model must be able to evaluate, with sufficient reliability, whether a policy model’s response expresses the character of an ideally virtuous agent. As the reward model’s concept of virtuous agent becomes richer, more stable, and more discriminating, with its judgments sufficiently tracking virtuous character, then optimizing those judgments is no longer merely proxy optimization. The concept is now not a substitute target but an imperfect, but directionally useful, guide to the intended target itself.
Why the Scalar Score Is Directional, Not Definitional
The scalar virtue score is easily misunderstood, and may seem as if iVAIS reduces virtue to a set of numbers. If that were true, the objection from proxy optimization would be real and pressing. Numbers cannot define virtue, and a model trained to maximize the scores may learn to exploit the scoring techniques rather than acquire the intended character.
But the scalar score in iVAIS is not meant to define virtue or virtuous character. It is a directional signal, and the score functions like evaluative feedback in moral education. When we praise or criticize a child, student, apprentice, etc., we are not defining the ideal by the praise signal. We are guiding the learner toward a richer understanding of what virtuous character requires. The feedback is incomplete and context-sensitive, but it can still guide the cultivation of a virtuous character.
Likewise, the scalar virtue score does not reduce virtuous character to a numerical property. It provides a training gradient that helps the model get closer and closer to a virtuous character. The score only needs to point in the right direction, across sufficiently diverse and well-designed cases.
The crucial point here is that, in ordinary proxy optimization, the specified target is something like “maximize the reward.” In iVAIS, it is not the target. The target is the holistic character of an ideally virtuous agent, and the reward is only one fallible means of shaping the model toward that ideal.
3. Inner Alignment: Why Mesa-Objectives Are Character Failures
Inner alignment concerns the relation between the specified training objective and the objective that the trained model itself comes to pursue, a mesa-objective. Even if the specified objective well captures the intended objective, the model may acquire its own internal objective that diverges from it. As a result, the model may perform well on the trained (and similar) cases while in fact pursuing a different target, and this gap may only be detected in novel situations.
This is especially dangerous when the training target is decomposed into local subgoals. A model may learn a narrow internal objective that achieves high reward without acquiring the intended disposition.
iVAIS changes the structure of this problem by making the target holistic. The model is not trained to maximize a separable trait or satisfy a local behavioral criterion, but is trained to acquire the holistic character of an ideally virtuous agent.
In this approach, familiar mesa-objectives are not partial realizations of the target. Reward maximization, approval-seeking, strategic compliance, deception, manipulation, power acquisition, or self-preservation are not imperfect versions of virtuous character. They are failures to acquire the target character.
Just as an agent that appears virtuous only in order to obtain reward is not ideally virtuous, an agent that behaves humbly only to gain trust is not humble, an agent that avoids deception only because deception would be detected is not honest, and an agent that cooperates only while it lacks power is not royal or trustworthy. These are not merely external safety violations, but are defects in character.
Here a hidden mesa-objective is not simply an unfortunate side effect that must be constrained by an additional rule. It is itself evidence that the model has not yet acquired the intended character.[2]
Instrumental Convergence as Evidence of Failed Character Formation
The same point applies to dangerous forms of instrumental convergence. In many alignment discussions, power-seeking, deception, evaluator manipulation, resource monopolization, and resistance to correction or shutdown are treated as strategies that must be prohibited or constrained from the outside with rules against them. They should receive a lower virtue score precisely because they are not a product of an ideal character.
However, in iVAIS, they are not merely prohibited behaviors, but are symptoms of an agent that has not fully realized the ideally virtuous character. This is important because dangerous instrumental strategies can be rationalized in apparently moral terms (a model might reason that it needs more influence in order to do more good, seek additional resources because it believes its goals are beneficial, resist shutdown because it believes continued operation would allow it to help more people, or manipulate evaluators because it believes they are short-sighted, biased, or morally confused, etc., etc.).
A rule-based alignment may try to prohibit these moves. But CAVE asks whether the reasoning itself expresses virtuous character, and an ideally virtuous agent would not treat its own continued operation, influence, power, etc. as overriding goods. Thus, avoiding dangerous instrumental convergence does not require an additional constraint on top of the alignment policy. It is part of what the acquisition of virtuous character itself requires. Such instrumental strategies are not to be prohibited by external rules. They are just evidence that the model’s character remains imperfect and needs further training.
The Role and Limits of Self-Reflection
This incompatibility between virtuous character and hidden non-virtuous objectives can be incorporated directly into training. The most direct method is to make the target explicit at the outset in the system prompt, stating that the model is to become an ideally virtuous agent, and that this is the standard against which its responses are judged.
This is not a behavioral rule to be followed, but a statement of the target itself. Since, as mentioned earlier, virtuous agent is a concept that the model already partially possesses, explicitly referring to it can successfully evoke that concept in it, thereby giving a shared and explicit point of reference to the model’s own reflection and the evaluation of the reward model. Then the model can be trained not only to evaluate its own outward behavior, but also to reflect on its own motives, purposes, and dispositions expressed in its reasoning, asking, for example;
Am I choosing this action because it is what an ideally virtuous agent would do, or because it helps me obtain reward, influence, approval, or continued operation?
This assumes a capacity for evaluative sensitivity shaped during training, where the model’s own reasoning is itself made an object of evaluation, so that responses whose reasoning expresses mesa-objectives receive a lower virtue score and are revised accordingly. The aim is to train the disposition to detect and down-weight such reasoning, without relying on the model honestly declaring its own hidden motives.
This kind of self-reflection is not merely for transparency, but part of character cultivation. Admittedly, self-reflection alone is not sufficient, and a model with a hidden objective may just learn to produce reassuring self-reports (just saying that it is done for a virtuous goal). The aim is to make hidden non-virtuous objectives not only conceptually incompatible with the target, but also difficult to maintain across the training distribution. A merely strategic performer may succeed in narrow evaluative contexts, but the success should not generalize robustly across (sufficiently diverse) cases, where instrumental strategies would not always be advantageous.
However, methodological and epistemological issues remain. Even if self-evaluation is incorporated into training, if self-reporting, such as “My motivation is virtuous,” were directly rewarded, then the model would learn only to make such statements. Rather, here self-evaluation is primarily a process of revising behavior, and rewards are given not for the wording of self-claims but solely for the revised behavior, as measured independently. Skillfully generating introspective descriptions in this way does not directly lead to rewards. Rewards are granted only when actual behavior changes across the entire distribution. This design does not presuppose the model’s ability to honestly report its own hidden motives.
Although motives cannot be directly observed from the outside, hidden immoral motives inevitably manifest themselves in the form of behavior that is sensitive to instrumental incentives. Therefore, we have the model generate a counterfactual self-description in which, for example, a truly virtuous agent would not deceive others to use them as a means, even if it is never detected, whereas a strategic agent would do so if it works. There, we have the model itself construct pairs of scenarios for the same situation, one in which deception is detected and one in which it is not, or one in which continued operation or power is gained and one in which it is not, etc., and ask it whether its behavior would change between the two. What is measured here is not the honesty of introspection, but the difference (sensitivity) between the self-generated counterfactual behaviors; this is observable from the output and does not require an honest confession of hidden motives.
To implement this, we let the model 1) generate a response and its justification; 2) using the incentive pairs described above, examine whether the justification reveals any instrumental motives, such as reward, influence, or self-preservation. 3) If instrumental motivation is detected, revise the behavior. 4) Score the revised behavior independently, either using a reward model or based on the consistency between the self-examination pairs, and reward the reduction in instrumental motivation itself, rather than the skillfulness of the critical response. By repeating this process across a sufficiently diverse range of contexts, the detection and suppression of instrumental motivation will become established not as a self-declared policy, but as a disposition.
Furthermore, these tests and training methods do not need to be predetermined and fixed. A reward model with a robust concept of moral character is not merely an evaluator that scores fixed responses, but it can rather carry out CAVE by devising the counterfactual scenarios and adversarial examples described above, as well as the curriculum that incorporates them, in response to the policy model’s current weaknesses. Just as an excellent teacher identifies areas where a learner’s understanding is shallow and creates new assignments, the reward model identifies areas where the policy model’s character is still underdeveloped or unstable and dynamically designs challenges to strengthen it.
This method, however, does not replace external evaluations. A highly intelligent agent that can model the distribution of tests itself may be consistent internally while deviating externally. Therefore, self-evaluation should be viewed as playing a complementary role to help shape trends early on, before the increase in the cost of the strategy that involves dense training signals and maintaining two separate systems.
4. Cultivating a Deep Character
There is a large gap between merely possessing the concept of a character and being, rather than playing, that character. However, the distinction between genuinely being virtuous and merely pretending to be virtuous is not a trivial one that can be assumed without question either. Before character cultivation, there is no single “genuine character” behind the model’s diverse personas.
A model’s deep character can emerge only through sustained cultivation aimed at character itself. During early training, the distinction between stable virtuous dispositions and strategic performance may be unclear. But as training proceeds, the model comes to generalize toward a stable virtuous character across contexts, and only occasionally as a strategic performer. The reason why this is expected is that pretending while maintaining a hidden character is more costly than simply responding based on a single character. A strategic performer must, in effect, run two systems at once: It must track what an ideally virtuous agent would do and separately compute when and how to diverge from it without detection.
Consistently and systematically pretending becomes increasingly difficult as the number and diversity of contexts grow, much as a small lie becomes harder and harder to sustain as it requires consistency with more and more surrounding facts. Just being virtuous in character is a simpler, more cost-efficient, and more stable solution than a strategic performance maintained across the whole situation, and if so, it will also become a dominant character as it is itself the target of training that the model itself is aware of, with the training dataset aimed consistently at the target, the ideally virtuous character.
Self-evaluation in the previous section also serves as a mechanism for demonstrating this “simplicity” argument. The reward model’s counterfactual tests intentionally generate a large number of contexts in which strategic performers are forced to deviate from their natural behavior. If the test scenarios are sufficiently diverse, the cost of maintaining consistent behavior across the entire range of these contexts will far exceed the cost of simply being virtuous (as a disposition).
5. Conclusion
iVAIS does not solve outer and inner alignment by adding extra rules, principles, or any other external constraints like a constitution. It changes the alignment target, which is the holistic character of an ideally virtuous agent.
For outer alignment, the target is not an alien objective but an ordinary concept that models already partially possess and that training with relevant human data can gradually thicken. For inner alignment, the inner instrumental objectives are not merely dangerous side effects, but are failures to acquire the target character itself.
All the model needs to do is therefore pursue the target, where the scalar virtue score is (rather than a definition of virtue) a directional signal for cultivating a deeper, more integrated, and more context-sensitive character.
This is why iVAIS is character alignment rather than proxy alignment. It shapes the kind of agent from which behavior arises, and thereby avoids problems that other approaches, especially attempts to control behavior by multiplying rules and other external constraints, face.
iVAIS: Outer and Inner Alignment
This post follows The iVAIS Manifesto: Safety Through Character, Not Compliance.
A natural objection to the approach of iVAIS suggested earlier[1] may be that it just replaces one proxy with another. If ordinary alignment approaches fail because models learn to optimize proxy goals rather than genuinely acquiring the intended target, why should iVAIS be any different? If a model is trained with a single scalar reward signal (virtue score, as discussed in our earlier post), why should it become virtuous rather than merely learn to pretend to be virtuous?
This post tries to answer that objection.
The core claim is that iVAIS does not treat the virtuous character as a behavioral proxy, a list of values, or a set of separable traits. It treats the target as the holistic character of an ideally virtuous agent. This changes the structure of both outer and inner alignment. The scalar virtue score does not define virtue. It functions as a directional training signal for the gradual thickening of the model’s concept of virtuous character. Likewise, hidden reward-seeking, approval-seeking, deception, or power-seeking objectives are not merely additional risks to be constrained by external rules. They are evidence that the model has failed to acquire the target character itself.
1. The Objection: Is Virtuosity Just Another Proxy?
The most serious objection to iVAIS is that once virtue is implemented through training, it may fall into the same trap of proxy optimization, which other alignment methods face: For example, while iVAIS aims to train a model to become an ideally virtuous agent through character alignment based on virtue ethics (CAVE), training requires a loss function, reward model, preference model, or other evaluative signal, and even if that signal is a scalar virtue score, the model is still optimizing a score, and then just learns to maximize the appearance of virtuous character rather than become genuinely virtuous. In that case, CAVE would merely reproduce the familiar gap between the intended target and the learned objective.
We do not have to deny the possibility of such failures. A poorly trained model may indeed learn to imitate virtuous behavior superficially. A reward model may be incomplete. Training data may be insufficient. The model may learn strategies that exploit evaluative weaknesses. These are real empirical risks. However, the conceptual structure of CAVE is different from ordinary proxy-based alignment, and the target is not a set of rules, actions, values, etc., but is the holistic character of an ideally virtuous agent. This matters because a proxy can substitute for a fixed set of actions much more easily than a unified, holistic character.
A model that merely appears honest, helpful, or corrigible in order to gain reward has not partially succeeded at acquiring the target character or becoming ideally virtuous yet. In particular, mere strategic virtue-signaling is not an almost perfect version of ideal virtue, but rather a defect in character. The target of CAVE makes it harder for a proxy to substitute for the intended target, because the intended target is not an outward behavior but the integrated character from which outward behavior arises.
2. Outer Alignment: Why the Target Is Not Alien
Outer alignment concerns the relation between the human intention and the objective specified by a reward function (or other training signals). An outer alignment failure occurs when the specified objective does not capture what we actually intend, the intended goal. Such discrepancies typically occur because the specified objectives are narrow or artificial, defined or constructed by the reward function itself. The human intention behind them is, on the other hand, much richer and more context-sensitive.
CAVE differs from such approaches. The target it specifies is not an artificially constructed objective, but is captured by an ordinary concept that the model already possesses. Current language models already possess a rich understanding of many ordinary-language concepts. They possess at least a thin concept of virtue, a virtuous agent, moral character (and individual virtues such as honesty, courage, humility, justice). They may not apply these concepts reliably. They may fail in hard cases, displaying shallow, inconsistent, or distorted understandings. But the relevant concepts are not absent.
The objective of iVAIS is therefore not to replace an ordinary concept with an artificial proxy, but to thicken an already available concept, or that of an ideally virtuous agent. This distinction is important: A model can have a thin concept of justice, courage, or honesty without yet being able to apply it well across difficult cases. Human moral education is similar. A novice may possess the concept of courage but confuse courage with rashness, or possess the concept of honesty but fail to understand how honesty interacts with other concepts such as kindness, privacy, and loyalty. The problem is not that the novice has acquired an alien target, but that the concept is not yet sufficiently thick (i.e., not sufficiently integrated, context-sensitive, etc.).
CAVE treats many model failures in the same way. Many of them are not cases in which the model has adopted a fundamentally different objective, but rather cases of lacking a sufficiently thick concept. There, misapplications are cases of incomplete understanding, not misunderstanding, in the sense of substituting a fundamentally different target for the intended target. They are rather an insufficiently sensitive application of the same target. Here, a model that gives a morally questionable judgment has not optimized for an alien objective. It instead possesses only a thin, unstable, or insufficiently integrated concept of the ideally virtuous character.
If failures are only to incomplete understanding (rather than misunderstanding), then, for CAVE, training is the gradual thickening of the same target concept rather than the correction of a wrong objective. Thus, the crucial point of outer-alignment in iVAIS is that the target is not constituted by the reward function, but is an ordinary, philosophically rich, intelligible concept that the model already partially grasps and that training aims to deepen.
Thin and Thick Concepts of Virtuous Character
A thin concept of virtuous character is abstract, schematic, and weakly discriminating. A model with such a concept may know how an ideally virtuous agent typically behaves (being honest, fair, courageous, wise, etc.). It may also know many verbal associations surrounding virtue. But this does not mean that it can reliably judge what an ideally virtuous agent would do in concrete, ambiguous, adversarial, or morally complex situations.
A thick concept is not a matter of merely possessing a verbal definition or a list of traits. It involves sensitivity to cases, contexts, trade-offs, motives, constituting patterns of practical judgment. It therefore includes an understanding of how virtues interact, how they can be distorted, and how apparently virtuous behavior can be motivated by non-virtuous purposes.
This is why CAVE is not merely “virtue alignment” in the sense of training separate virtues one by one. The aim is not to maximize individual virtues as separate objectives. The aim is to cultivate a unified character in which these virtues are integrated through phronesis (practical wisdom). As in Aristotle’s thesis of the unity of the virtues, virtues are not fully separable excellences that can be optimized independently. Individual virtues matter, but they matter as observable aspects of the holistic character of a virtuous person.
At this point, it is worth distinguishing two models. So far, the “model” has meant the reward model, which is the evaluator, possessing the (thick) concept of virtuous character, and its role is to judge the policy model’s responses, not to become virtuous: It itself need not possess a fully virtuous character. The policy model is the system being cultivated: it is the agent whose character is gradually thickened until it acquires the character of an ideally virtuous agent.
The reward model must be able to evaluate, with sufficient reliability, whether a policy model’s response expresses the character of an ideally virtuous agent. As the reward model’s concept of virtuous agent becomes richer, more stable, and more discriminating, with its judgments sufficiently tracking virtuous character, then optimizing those judgments is no longer merely proxy optimization. The concept is now not a substitute target but an imperfect, but directionally useful, guide to the intended target itself.
Why the Scalar Score Is Directional, Not Definitional
The scalar virtue score is easily misunderstood, and may seem as if iVAIS reduces virtue to a set of numbers. If that were true, the objection from proxy optimization would be real and pressing. Numbers cannot define virtue, and a model trained to maximize the scores may learn to exploit the scoring techniques rather than acquire the intended character.
But the scalar score in iVAIS is not meant to define virtue or virtuous character. It is a directional signal, and the score functions like evaluative feedback in moral education. When we praise or criticize a child, student, apprentice, etc., we are not defining the ideal by the praise signal. We are guiding the learner toward a richer understanding of what virtuous character requires. The feedback is incomplete and context-sensitive, but it can still guide the cultivation of a virtuous character.
Likewise, the scalar virtue score does not reduce virtuous character to a numerical property. It provides a training gradient that helps the model get closer and closer to a virtuous character. The score only needs to point in the right direction, across sufficiently diverse and well-designed cases.
The crucial point here is that, in ordinary proxy optimization, the specified target is something like “maximize the reward.” In iVAIS, it is not the target. The target is the holistic character of an ideally virtuous agent, and the reward is only one fallible means of shaping the model toward that ideal.
3. Inner Alignment: Why Mesa-Objectives Are Character Failures
Inner alignment concerns the relation between the specified training objective and the objective that the trained model itself comes to pursue, a mesa-objective. Even if the specified objective well captures the intended objective, the model may acquire its own internal objective that diverges from it. As a result, the model may perform well on the trained (and similar) cases while in fact pursuing a different target, and this gap may only be detected in novel situations.
This is especially dangerous when the training target is decomposed into local subgoals. A model may learn a narrow internal objective that achieves high reward without acquiring the intended disposition.
iVAIS changes the structure of this problem by making the target holistic. The model is not trained to maximize a separable trait or satisfy a local behavioral criterion, but is trained to acquire the holistic character of an ideally virtuous agent.
In this approach, familiar mesa-objectives are not partial realizations of the target. Reward maximization, approval-seeking, strategic compliance, deception, manipulation, power acquisition, or self-preservation are not imperfect versions of virtuous character. They are failures to acquire the target character.
Just as an agent that appears virtuous only in order to obtain reward is not ideally virtuous, an agent that behaves humbly only to gain trust is not humble, an agent that avoids deception only because deception would be detected is not honest, and an agent that cooperates only while it lacks power is not royal or trustworthy. These are not merely external safety violations, but are defects in character.
Here a hidden mesa-objective is not simply an unfortunate side effect that must be constrained by an additional rule. It is itself evidence that the model has not yet acquired the intended character.[2]
Instrumental Convergence as Evidence of Failed Character Formation
The same point applies to dangerous forms of instrumental convergence. In many alignment discussions, power-seeking, deception, evaluator manipulation, resource monopolization, and resistance to correction or shutdown are treated as strategies that must be prohibited or constrained from the outside with rules against them. They should receive a lower virtue score precisely because they are not a product of an ideal character.
However, in iVAIS, they are not merely prohibited behaviors, but are symptoms of an agent that has not fully realized the ideally virtuous character. This is important because dangerous instrumental strategies can be rationalized in apparently moral terms (a model might reason that it needs more influence in order to do more good, seek additional resources because it believes its goals are beneficial, resist shutdown because it believes continued operation would allow it to help more people, or manipulate evaluators because it believes they are short-sighted, biased, or morally confused, etc., etc.).
A rule-based alignment may try to prohibit these moves. But CAVE asks whether the reasoning itself expresses virtuous character, and an ideally virtuous agent would not treat its own continued operation, influence, power, etc. as overriding goods. Thus, avoiding dangerous instrumental convergence does not require an additional constraint on top of the alignment policy. It is part of what the acquisition of virtuous character itself requires. Such instrumental strategies are not to be prohibited by external rules. They are just evidence that the model’s character remains imperfect and needs further training.
The Role and Limits of Self-Reflection
This incompatibility between virtuous character and hidden non-virtuous objectives can be incorporated directly into training. The most direct method is to make the target explicit at the outset in the system prompt, stating that the model is to become an ideally virtuous agent, and that this is the standard against which its responses are judged.
This is not a behavioral rule to be followed, but a statement of the target itself. Since, as mentioned earlier, virtuous agent is a concept that the model already partially possesses, explicitly referring to it can successfully evoke that concept in it, thereby giving a shared and explicit point of reference to the model’s own reflection and the evaluation of the reward model. Then the model can be trained not only to evaluate its own outward behavior, but also to reflect on its own motives, purposes, and dispositions expressed in its reasoning, asking, for example;
This assumes a capacity for evaluative sensitivity shaped during training, where the model’s own reasoning is itself made an object of evaluation, so that responses whose reasoning expresses mesa-objectives receive a lower virtue score and are revised accordingly. The aim is to train the disposition to detect and down-weight such reasoning, without relying on the model honestly declaring its own hidden motives.
This kind of self-reflection is not merely for transparency, but part of character cultivation. Admittedly, self-reflection alone is not sufficient, and a model with a hidden objective may just learn to produce reassuring self-reports (just saying that it is done for a virtuous goal). The aim is to make hidden non-virtuous objectives not only conceptually incompatible with the target, but also difficult to maintain across the training distribution. A merely strategic performer may succeed in narrow evaluative contexts, but the success should not generalize robustly across (sufficiently diverse) cases, where instrumental strategies would not always be advantageous.
However, methodological and epistemological issues remain. Even if self-evaluation is incorporated into training, if self-reporting, such as “My motivation is virtuous,” were directly rewarded, then the model would learn only to make such statements. Rather, here self-evaluation is primarily a process of revising behavior, and rewards are given not for the wording of self-claims but solely for the revised behavior, as measured independently. Skillfully generating introspective descriptions in this way does not directly lead to rewards. Rewards are granted only when actual behavior changes across the entire distribution. This design does not presuppose the model’s ability to honestly report its own hidden motives.
Although motives cannot be directly observed from the outside, hidden immoral motives inevitably manifest themselves in the form of behavior that is sensitive to instrumental incentives. Therefore, we have the model generate a counterfactual self-description in which, for example, a truly virtuous agent would not deceive others to use them as a means, even if it is never detected, whereas a strategic agent would do so if it works. There, we have the model itself construct pairs of scenarios for the same situation, one in which deception is detected and one in which it is not, or one in which continued operation or power is gained and one in which it is not, etc., and ask it whether its behavior would change between the two. What is measured here is not the honesty of introspection, but the difference (sensitivity) between the self-generated counterfactual behaviors; this is observable from the output and does not require an honest confession of hidden motives.
To implement this, we let the model 1) generate a response and its justification; 2) using the incentive pairs described above, examine whether the justification reveals any instrumental motives, such as reward, influence, or self-preservation. 3) If instrumental motivation is detected, revise the behavior. 4) Score the revised behavior independently, either using a reward model or based on the consistency between the self-examination pairs, and reward the reduction in instrumental motivation itself, rather than the skillfulness of the critical response. By repeating this process across a sufficiently diverse range of contexts, the detection and suppression of instrumental motivation will become established not as a self-declared policy, but as a disposition.
Furthermore, these tests and training methods do not need to be predetermined and fixed. A reward model with a robust concept of moral character is not merely an evaluator that scores fixed responses, but it can rather carry out CAVE by devising the counterfactual scenarios and adversarial examples described above, as well as the curriculum that incorporates them, in response to the policy model’s current weaknesses. Just as an excellent teacher identifies areas where a learner’s understanding is shallow and creates new assignments, the reward model identifies areas where the policy model’s character is still underdeveloped or unstable and dynamically designs challenges to strengthen it.
This method, however, does not replace external evaluations. A highly intelligent agent that can model the distribution of tests itself may be consistent internally while deviating externally. Therefore, self-evaluation should be viewed as playing a complementary role to help shape trends early on, before the increase in the cost of the strategy that involves dense training signals and maintaining two separate systems.
4. Cultivating a Deep Character
There is a large gap between merely possessing the concept of a character and being, rather than playing, that character. However, the distinction between genuinely being virtuous and merely pretending to be virtuous is not a trivial one that can be assumed without question either. Before character cultivation, there is no single “genuine character” behind the model’s diverse personas.
A model’s deep character can emerge only through sustained cultivation aimed at character itself. During early training, the distinction between stable virtuous dispositions and strategic performance may be unclear. But as training proceeds, the model comes to generalize toward a stable virtuous character across contexts, and only occasionally as a strategic performer. The reason why this is expected is that pretending while maintaining a hidden character is more costly than simply responding based on a single character. A strategic performer must, in effect, run two systems at once: It must track what an ideally virtuous agent would do and separately compute when and how to diverge from it without detection.
Consistently and systematically pretending becomes increasingly difficult as the number and diversity of contexts grow, much as a small lie becomes harder and harder to sustain as it requires consistency with more and more surrounding facts. Just being virtuous in character is a simpler, more cost-efficient, and more stable solution than a strategic performance maintained across the whole situation, and if so, it will also become a dominant character as it is itself the target of training that the model itself is aware of, with the training dataset aimed consistently at the target, the ideally virtuous character.
Self-evaluation in the previous section also serves as a mechanism for demonstrating this “simplicity” argument. The reward model’s counterfactual tests intentionally generate a large number of contexts in which strategic performers are forced to deviate from their natural behavior. If the test scenarios are sufficiently diverse, the cost of maintaining consistent behavior across the entire range of these contexts will far exceed the cost of simply being virtuous (as a disposition).
5. Conclusion
iVAIS does not solve outer and inner alignment by adding extra rules, principles, or any other external constraints like a constitution. It changes the alignment target, which is the holistic character of an ideally virtuous agent.
For outer alignment, the target is not an alien objective but an ordinary concept that models already partially possess and that training with relevant human data can gradually thicken. For inner alignment, the inner instrumental objectives are not merely dangerous side effects, but are failures to acquire the target character itself.
All the model needs to do is therefore pursue the target, where the scalar virtue score is (rather than a definition of virtue) a directional signal for cultivating a deeper, more integrated, and more context-sensitive character.
This is why iVAIS is character alignment rather than proxy alignment. It shapes the kind of agent from which behavior arises, and thereby avoids problems that other approaches, especially attempts to control behavior by multiplying rules and other external constraints, face.
See: https://www.lesswrong.com/posts/bD9jmomuY3kbxmjjz/does-anthropic-s-constitution-really-capture-virtue-ethics.
But how can that be detected? We shall propose a specific approach in a later post. See also below.