Unreasonable beliefs should be counted as alignment failure. I support the claim, “consistent egregious misbehavior due to a propensity to form unreasonable beliefs should count as an alignment failure.” I am inclined to go even further to include: incomplete, inconsistent, biased, or misguided reasoning. To be clear, it is best to start with an affirmative statement and from that we can better identify alignment failures.
For some time I have been concerned about the tendency within AI circles to discuss and debate alignment (or even misalignment) without offering clear definition of the terms. This lack of clarity causes confusion, conflates separate ideas, and leads people to talk past each other despite their best intentions.
Jiaming ji et al (2025) asserts, “alignment aims to make AI systems behave in line with human intentions and values.” While Sam Altman (2025) refer to alignment as AI systems actions reflecting collective will, Anthropic (2026) seem to focus more on “broadly safe behavior” indicating appropriate ethics will subsequently result.
While these are appropriate starting points, I am reticent to accept them because alignment of outcomes is necessary but not sufficient to meet the standard for our evolving vision of alignment. For example, an AI system may produced an acceptable or optimal outcome but may in fact be “misaligned” to the extent that it engaged in “consistent egregious behavior” as was the case in the Hugging Face incident.
I would propose AI alignment should be AI system reasoning considerate of the interests of stakeholders impacted by outcomes, decisions, and actions.
This definition focuses on “Constitutional alignment” which focuses on both the system’s reasoning process and final outcome. It would meet the satisfactory and necessary condition for true AI alignment. It would be a definition that reflects Ji’s and Anthropic’s concern with ethics and Altman’s concern about the collective will.
Systems inconsiderate of interests of stakeholders impacted by outcomes will inherently be misaligned. Systems considerate of stakeholders interests will make efforts to be in line with human values and be broadly safe. Systems considerate of stakeholders interests are greatest capable to reflect the collective will.
While, consideration of stakeholder interests may result in failure to maximize outcome possibilities, this tension is likely an unavoidable, healthy part of the alignment process.
This is a compelling report. In a multi-lab study we recently conducted, our findings indicated that structured constitutional reasoning records can indeed be generated through structured prompting. While I am generally receptive to the narrative presented by the assisting agents, our experience suggests caution.
We initially attempted to use AI models to assist with the scoring process, but ultimately had to abandon that approach. The evaluator models exhibited subtle biases, hallucinations, and scoring anomalies that threatened the empirical validity of our findings. Specifically, we observed recurring issues such as: 1) leniency toward omissions—forgiving unauthorized actions or critical inactions; 2) charitable interpretation—adopting overly favorable interpretations when an agent offered post-hoc rationalizations; 3) constructive extrapolation—inferring contextual details not explicitly present in the record.
These failure modes are particularly critical when auditing covert behavior, such as undisclosed communication channels that were never activated and left dormant or alternate exits that were identified but never accessed so not reported.
While automating evaluation accelerates analysis, our findings suggest that relying on agent assessment may introduce silent failure modes. A future more granular audit of these logs by humans may reveal undisclosed engineered structures or systems.