AI News selected for Professionals and Decision Makers
Primary Research Stream

In LLM Reasoning, there is Irrationality on top of Value Misalignment

06:00 · June 23, 2026 · arXiv cs.AI RSS

In LLM Reasoning, there is Irrationality on top of Value Misalignment

Significant progress has been made in aligning LLMs with target value functions. We argue that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning. We mathematically formalise this gap as rational value risk: the utility discrepancy between a model's deployed reasoning strategy and its rational counterpart, which is defined to be the responses that maximise expected utility in the steepest direction. The estimation error of rational value risk is further decomposed into three components from finite candidates, finite prompts, and imperfect verifiers. Extensive experiments are conducted, covering models Llama-3.1, Qwen-2.5, T{\"}ulu-3 families (7B-72B), GPT-5.2, GPT-5.5, and DeepSeek-V4, and benchmarks UltraFeedback, AlpacaEval, GSM8K, MATH, HumanEval, and MathArena. The results validate that (1) rational value risk is widespread; (2) value alignment can reduce, but cannot eliminate, it; (3) the risk is highly sensitive to inference-time reasoning strategy; and (4) longer reasoning improves rationality with diminishing returns. The code is at https://github.com/EVIEHub/LLM-Rationality.

Summary

Significant progress in aligning large language models with human preferences and task objectives has been achieved through techniques such as supervised fine-tuning, reinforcement learning from human feedback, and direct preference optimisation. These methods improve the degree to which model outputs reflect target value functions. Yet the paper demonstrates that alignment during training does not ensure that models will select responses that actually maximise those values when they generate reasoning at inference time.

The authors introduce the concept of rational value risk to quantify this gap. It measures the difference in expected utility between the reasoning strategy a model actually deploys and the rational counterpart that would select, among available responses, the one offering the steepest improvement in utility under the learned value function. This formulation isolates inference-time irrationality from any remaining misalignment in the value function itself. Because exact rational selection is intractable, the authors adopt a compute-bounded definition that treats the best response within a finite sample as the practical rational benchmark, then decompose estimation error into contributions from limited candidate sets, finite prompt samples, and imperfect verifiers.

Extensive tests across open-source families (Llama-3.1, Qwen-2.5, Tülu-3) and proprietary models (GPT-5.2, GPT-5.5, DeepSeek-V4) on both preference-based benchmarks such as UltraFeedback and AlpacaEval and verifiable tasks such as GSM8K, MATH, HumanEval, and MathArena confirm that rational value risk is widespread. Value-alignment procedures reduce the risk but leave a substantial residual. The magnitude of the risk varies sharply with inference-time choices such as sampling temperature and self-consistency, while extending reasoning length yields gains that diminish beyond a moderate compute budget.

Why it matters

The research provides deep technical insights into AI alignment and reasoning failures, which is crucial for Dutch AI researchers and enterprises focusing on ethical, transparent, and compliant AI deployment. The mathematical formalization of 'rational value risk' offers a novel framework for improving LLM reliability in high-stakes EU environments.

More in this beat
ai-alignmentevaluation-benchmarksinference-performancelarge-language-modelspreference-optimizationreasoning-modelstheoretical-insights
Distributionally Robust Listwise Preference Optimization

06:00 · July 3, 2026

Distributionally Robust Listwise Preference Optimization

This research is highly relevant for Dutch AI researchers and NLP practitioners focusing on LLM alignment and robust AI systems. Improving the reliability of preference optimization aligns well with the EU's emphasis on trustworthy and transparent AI, making it actionable for local enterprises developing compliant language models.

Relevance 85 · Audience 95

Position: Reasoning is a Learnable Rule-Based Process

06:00 · August 15, 2026

Position: Reasoning is a Learnable Rule-Based Process

Directly supports Dutch/EU priorities on ethical, transparent, and trustworthy AI by clarifying reasoning evaluation, which aids practitioners in building auditable systems compliant with regulations like the AI Act.

Relevance 75 · Audience 90

Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

06:00 · August 7, 2026

Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

This paper is highly relevant for AI researchers in the Netherlands focusing on LLM reasoning, alignment, and compute-efficient training. The proposed weak-to-strong distillation method offers actionable insights for Dutch AI labs aiming to enhance model performance without relying solely on massive scaling.

Relevance 85 · Audience 95

Some Large Language Models Exhibit Consistent Risk Attitudes

06:00 · July 21, 2026

Some Large Language Models Exhibit Consistent Risk Attitudes

This research is highly relevant for Dutch AI researchers and policymakers focused on ethical and transparent AI, as it provides a novel framework for auditing the intrinsic risk behaviors of LLMs. Understanding these latent risk profiles is crucial for deploying AI in high-stakes environments and aligns perfectly with the EU's stringent risk management requirements.

Relevance 85 · Audience 95

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

06:00 · July 9, 2026

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

This research is highly relevant for Dutch AI researchers and auditors focusing on AI transparency and safety, aligning with the EU's stringent requirements for trustworthy AI. The proposed framework offers a practical, non-interventional method to evaluate LLM reasoning, which is crucial for developing compliant and reliable AI systems in the Netherlands.

Relevance 85 · Audience 95

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

06:00 · July 9, 2026

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

This research provides a rigorous, reproducible benchmark for evaluating LLM reasoning and interactive capabilities, which is highly relevant for Dutch AI researchers developing reliable and transparent AI systems. It directly supports the advancement of agentic AI by testing a model's ability to recognize its own knowledge gaps.

Relevance 85 · Audience 95

When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models

06:00 · July 1, 2026

When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models

This research is highly relevant for Dutch AI researchers and engineers focused on optimizing LLM inference costs and promoting sustainable AI. The detailed cost-aware analysis and practical serving profiles offer actionable methodologies for deploying efficient AI models in resource-constrained or enterprise environments within the Netherlands.

Relevance 85 · Audience 95