Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture
06:00 · August 18, 2026 · arXiv cs.AI RSS

Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored. We posit that this imbalance stems from the field's reliance on descriptive ethics frameworks, such as Moral Foundations Theory and Kohlberg's stages of moral development, which emphasize value representation over normative application. We review existing benchmarks and evaluation methods, and show that they cluster heavily around the value problem, while discussion regarding normative ethics remains underrepresented. We identify three crucial gaps: (i) the absence of high-quality ground-truth data for moral norms and their applications, (ii) insufficient evaluation of intermediate reasoning processes, and (iii) limited attention to the identification of morally relevant features in context. Subsequently, we propose a research agenda that includes the development of standardized formal representations for normative theories, the construction of expert-annotated datasets capturing norm application, and evaluation protocols that explicitly distinguish between values-level and norms-level competence. Our goal is to encourage a more systematic study of normative reasoning in LLMs.
Summary
Current evaluations of moral reasoning in large language models concentrate on whether model outputs match patterns of human moral values, an issue the authors term the moral value problem. These assessments typically draw on descriptive ethics instruments such as Moral Foundations Theory questionnaires, the Moral Machine dilemma set, and large-scale value surveys. While such methods can reveal whether models reproduce observed human preferences, they leave aside the moral norm problem: the capacity to identify morally relevant features in a given context and to apply structured principles that determine how values should guide judgments in that setting.
The paper traces the imbalance to the field’s reliance on descriptive frameworks that prioritize value representation over normative application. As a result, existing benchmarks rarely supply high-quality ground-truth data on norm application, seldom examine intermediate reasoning steps against expert standards, and pay limited attention to how models detect context-specific moral salience. A model may therefore align with aggregate human attitudes while still failing to construct coherent arguments within established ethical theories or to recognize when particular principles are relevant.
To address these shortcomings, the authors outline a research agenda centered on three elements: the creation of standardized formal representations of normative theories, the construction of expert-annotated datasets that capture norm application across varied scenarios, and the design of evaluation protocols that separately measure values-level alignment and norms-level competence. The proposed direction aims to support more systematic assessment of normative reasoning without assuming that value matching alone suffices for moral competence.
Why it matters
This article is highly relevant for Dutch AI researchers and practitioners focused on ethical AI, aligning strongly with the Netherlands' and EU's emphasis on transparent and trustworthy AI systems. It provides a critical framework for advancing LLM evaluation beyond simple value alignment toward robust normative reasoning.










