AI News selected for Professionals and Decision Makers
Primary Research Stream

Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors

06:00 · August 17, 2026 · arXiv cs.AI RSS

Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors

High-confidence errors in large language models are often treated as evidence of fragile internal inference. We study a different possibility: stable miscalibration, where a confident wrong answer remains locally stable under small perturbations. We combine two diagnostics: a label-aware output-level audit score that ranks domains by confidence variation and overconfident mistakes under a forced-answer baseline, and an internal sensitivity probe that measures hidden-state movement. On a multi-domain binary factual audit set, this audit score tracks where abstention-aware self-critique reduces decision loss, although direct labeled baselines rank the same gain more strongly. Internally, self-critical prompting consistently reduces hidden-state sensitivity across layers in three open-weight models. This supports prompt-induced local stabilization rather than a purely output-level abstention pattern, but it does not imply calibration: audit-defined overconfident errors are not clearly more locally sensitive than confidently correct answers, so some high-confidence errors may be stable and miscalibrated rather than simply fragile.

Summary

High-confidence errors in large language models are commonly viewed as signs of fragile internal inference, where small input changes readily flip an incorrect output. This paper examines an alternative: stable miscalibration, in which a confidently wrong answer persists under modest perturbations because the model remains near the same internal state. The distinction matters for uncertainty estimation, because fragility-based probes would then fail to separate overconfident mistakes from confidently correct answers.

To test the idea, the author combines two diagnostics on a frozen set of 532 binary factual items spanning eleven domains. An output-level audit score ranks domains according to confidence variation across three policies—a forced-answer baseline, a cautious abstention policy, and a self-critical abstention policy—plus the mass of overconfident mistakes produced by the forced-answer baseline. An internal sensitivity probe tracks layer-wise movement of hidden states under the same policies. Experiments cover three open-weight models, including Llama-3.1 and DeepSeek-R1.

The audit score identifies domains where abstention-aware self-critique reduces policy-aware decision loss, although direct labeled baselines capture the same improvement more strongly. Internally, self-critical prompting lowers hidden-state sensitivity across layers, indicating prompt-induced local stabilization rather than a purely output-level abstention effect. However, items that the audit flags as overconfidently wrong do not exhibit greater local sensitivity than confidently correct items. This pattern suggests that some high-confidence errors are locally stable yet miscalibrated, rather than merely fragile.

The findings imply that abstention policies and uncertainty estimates should account for stable miscalibration in addition to perturbation sensitivity when deciding when a model should refrain from answering.

Why it matters

This research is highly relevant for Dutch AI researchers and practitioners focused on trustworthy and transparent AI, a key priority in the Netherlands and the EU. Understanding stable miscalibration provides actionable insights for improving LLM reliability, auditing high-confidence hallucinations, and developing better abstention-aware decision policies for enterprise deployment.

More in this beat
confidence-calibrationhigh-confidence errorsllm-benchmarksrisk-and-limitationsstable miscalibrationuncertainty estimation
ASI-Bench: At the Dawn of Artificial Superintelligence

06:00 · August 19, 2026

ASI-Bench: At the Dawn of Artificial Superintelligence

Offers a novel, high-depth evaluation framework that Dutch AI researchers and advanced labs can directly apply to measure progress toward autonomous scientific agents, aligning with the Netherlands' strengths in ethical AI and SME-driven innovation.

Relevance 62 · Audience 88

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

06:00 · August 18, 2026

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

This article is highly relevant for Dutch AI researchers and practitioners focused on ethical AI, aligning strongly with the Netherlands' and EU's emphasis on transparent and trustworthy AI systems. It provides a critical framework for advancing LLM evaluation beyond simple value alignment toward robust normative reasoning.

Relevance 85 · Audience 95

AI Evaluation Should Work With Humans

06:00 · August 17, 2026

AI Evaluation Should Work With Humans

This paper aligns strongly with the Dutch and EU focus on ethical, human-centric AI and human oversight. It provides researchers with a conceptual foundation to develop new evaluation frameworks that prioritize human-AI collaboration over autonomous replacement, which is highly actionable for Dutch AI policy and enterprise deployment.

Relevance 85 · Audience 90

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

06:00 · August 15, 2026

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

This research is highly relevant for Dutch AI researchers and institutions focused on ethical AI deployment. It provides a concrete framework to evaluate and mitigate research misconduct risks when integrating LLMs into scientific workflows, aligning perfectly with the EU's emphasis on trustworthy AI.

Relevance 85 · Audience 95

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

06:00 · August 13, 2026

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

This research is highly relevant for Dutch AI researchers and MLOps practitioners developing autonomous agents, as it provides a rigorous, open-source framework for testing agent reliability and safety. Its focus on risk assessment and operational side-effects aligns strongly with the Netherlands' strategic emphasis on transparent, ethical, and secure AI deployments.

Relevance 85 · Audience 95

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

06:00 · August 4, 2026

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

This research is highly relevant for Dutch AI researchers developing autonomous LLM agents, providing a rigorous framework for evaluating continuous learning in realistic deployment scenarios. Understanding how model capabilities gate self-evolution is crucial for building robust and reliable AI systems.

Relevance 85 · Audience 95

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

06:00 · August 4, 2026

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

Directly actionable for Dutch AI researchers and advanced practitioners developing or continuously evaluating LLMs: reduces compute overhead while preserving ranking fidelity, aligns with EU emphasis on efficient and transparent AI, and requires no historical logs.

Relevance 72 · Audience 88

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

06:00 · August 3, 2026

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

Directly supports ethical, transparent AI development emphasized in Dutch/EU policy and the AI Act by validating safety measurements for LLM agents; Dutch practitioners can apply the released harness and findings to avoid over-reliance on unvalidated benchmarks in regulated deployments.

Relevance 82 · Audience 88