Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
06:00 · August 15, 2026 · arXiv cs.AI RSS

Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority labels is often high. However, rationale-level analysis reveals systematic divergence in the moral grounds expressed by human annotators and models. In particular, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority. Our results show that agreement should not be treated as equivalent to alignment. Label-based evaluation can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments.
Summary
Agreement with human judgments serves as a standard proxy for assessing whether large language models are aligned with human values, yet matching final labels on moral tasks does not establish that models and humans rely on the same underlying principles. Two parties can endorse the same verdict while invoking distinct considerations, such as differing interpretations of context, emphasis on particular duties, or weighting of excuses and consequences. The authors examine this gap through a 500-item benchmark derived from the ETHICS dataset, covering five domains of moral reasoning: commonsense morality, deontology, justice, utilitarianism, and virtue ethics. Both human annotators and multiple frontier and open models supplied not only binary or scalar judgments but also structured rationales for each item.
Label agreement between models and human majority votes proved high across many items. Rationale-level comparison, however, exposed consistent divergences: models redistributed attention among categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even on cases where the final label coincided with the human majority. These shifts appeared across model families and were not limited to edge cases. The pattern indicates that surface-level performance on existing benchmarks can mask differences in the moral considerations models foreground or downplay.
The findings carry direct implications for evaluation practice. When models justify identical conclusions through different moral grounds, downstream behavior may diverge in novel situations, and explanations offered to users may rest on priorities that human stakeholders would not recognize or endorse. The work therefore advocates supplementing label-based metrics with rationale-aware analysis that inspects the principles and contextual assumptions expressed in model outputs, thereby providing a more transparent basis for assessing ethical alignment.
Why it matters
Directly addresses ethical AI evaluation and transparency, aligning with Dutch government priorities on responsible AI and EU regulatory context; actionable for Dutch researchers and SMEs developing or auditing LLMs.






