AI News selected for Professionals and Decision Makers
Primary Research Stream

Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs

06:00 · July 24, 2026 · arXiv cs.AI RSS

Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs

When a language model gives different answers on repeated runs, does that variation reveal what it does not know? Self-consistency turns the variation into a per-question uncertainty estimate via majority voting. But does the same variation reveal cross-question structure -- related questions flipping together, the way a diverse ensemble does? We compare two regimes on the same questions: one model run $100$ times at $\tau=1$ versus an ensemble of $24$ LLMs run once each at $\tau=0$. A Marchenko--Pastur random-matrix test separates signal from sampling noise on both sides. Within any single model, at most one dimension rises above noise across five families and three benchmarks (MMLU, HellaSwag, GSM8K). Across the ensemble, four eigenvalues clear the noise edge, while a matched-difficulty Bernoulli null produces at most one in $500$ Monte Carlo draws. Self-consistency gives accurate per-question uncertainty but no detectable cross-question structure; only a diverse ensemble surfaces what a model does not know.

Summary

Temperature-based stochastic sampling within a single language model produces at most one dimension of structured epistemic uncertainty that exceeds sampling noise, according to a Marchenko–Pastur analysis of the run-by-question correctness matrix. When the same questions are presented to an ensemble of 24 diverse models, four eigenvalues rise above the noise threshold on MMLU, while matched-difficulty Bernoulli null matrices yield at most one such eigenvalue across 500 Monte Carlo trials. The gap persists across five model families and three benchmarks, including HellaSwag and GSM8K chain-of-thought.

Self-consistency methods convert repeated samples from one model into reliable per-question success probabilities through majority voting. Yet the same samples show near-zero correlations between questions, even within the same subject category, indicating that temperature-induced variation supplies independent noise rather than coordinated epistemic structure. In contrast, the across-model regime reveals correlated error patterns that survive the random-matrix null test, suggesting that model diversity captures shared limitations that single-model sampling does not.

The experimental design isolates this distinction by restricting analysis to borderline questions whose pass rates lie away from zero and one, then comparing the resulting correlation spectra against Tracy–Widom fluctuations around the Marchenko–Pastur edge. Within-model results remain consistent when temperature is varied between 0.5 and 1.5 and when the number of runs is reduced, confirming that the dimensionality ceiling is not an artifact of a single sampling regime. These findings imply that uncertainty estimates derived solely from intra-model stochasticity remain epistemically shallow relative to the richer signal obtained from heterogeneous model ensembles.

Why it matters

Directly actionable for Dutch AI teams building reliable LLM systems: ensembles outperform deeper sampling at far lower cost for selective prediction and uncertainty quantification, aligning with EU emphasis on transparent, trustworthy AI under the AI Act.

More in this beat
chain-of-thoughtconfidence-calibrationevaluation-benchmarkslarge-language-modelsmmlutheoretical-insights
Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

06:00 · July 11, 2026

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

This survey provides a rigorous, structured framework for evaluating medical LLMs, which is highly valuable for Dutch AI researchers and healthcare institutions developing transparent and safe clinical AI. Its focus on mitigating hallucinations and ensuring reliable reasoning aligns well with the EU AI Act and the Netherlands' emphasis on ethical AI deployment.

Relevance 85 · Audience 95

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

06:00 · August 4, 2026

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

Directly actionable for Dutch AI researchers and advanced practitioners developing or continuously evaluating LLMs: reduces compute overhead while preserving ranking fidelity, aligns with EU emphasis on efficient and transparent AI, and requires no historical logs.

Relevance 72 · Audience 88

Interpreting Latent CoT Reasoning as Dynamical Systems

06:00 · July 14, 2026

Interpreting Latent CoT Reasoning as Dynamical Systems

The article is highly relevant for AI researchers in the Netherlands focusing on LLM interpretability and trustworthy AI. Understanding the internal dynamics of latent reasoning aligns strongly with EU and Dutch priorities for transparent and explainable AI systems.

Relevance 85 · Audience 95

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

06:00 · July 1, 2026

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

This research provides a rigorous framework for evaluating the reasoning and reliability of LLMs in dynamic, multi-turn environments. For Dutch AI researchers and developers, understanding and benchmarking these epistemic updates is crucial for building trustworthy, transparent AI systems that align with EU standards.

Relevance 85 · Audience 95

PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

06:00 · June 23, 2026

PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

This research is highly relevant for Dutch AI researchers and advanced practitioners focusing on LLM reliability and multi-agent systems. The introduction of a dynamic, bias-reducing routing protocol aligns with the Netherlands' strategic emphasis on transparent, ethical, and robust AI development, offering actionable methodologies with open-source code.

Relevance 85 · Audience 95

In LLM Reasoning, there is Irrationality on top of Value Misalignment

06:00 · June 23, 2026

In LLM Reasoning, there is Irrationality on top of Value Misalignment

The research provides deep technical insights into AI alignment and reasoning failures, which is crucial for Dutch AI researchers and enterprises focusing on ethical, transparent, and compliant AI deployment. The mathematical formalization of 'rational value risk' offers a novel framework for improving LLM reliability in high-stakes EU environments.

Relevance 85 · Audience 95

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

06:00 · August 15, 2026

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

This research is highly relevant for Dutch AI researchers and institutions focused on ethical AI deployment. It provides a concrete framework to evaluate and mitigate research misconduct risks when integrating LLMs into scientific workflows, aligning perfectly with the EU's emphasis on trustworthy AI.

Relevance 85 · Audience 95

Position: Reasoning is a Learnable Rule-Based Process

06:00 · August 15, 2026

Position: Reasoning is a Learnable Rule-Based Process

Directly supports Dutch/EU priorities on ethical, transparent, and trustworthy AI by clarifying reasoning evaluation, which aids practitioners in building auditable systems compliant with regulations like the AI Act.

Relevance 75 · Audience 90

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

06:00 · August 13, 2026

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

SAPO provides a highly actionable, structured approach to prompt engineering that Dutch AI researchers and enterprise teams can use to build more reliable and interpretable LLM applications. Its focus on modular, non-destructive prompt updates aligns with the EU's demand for robust, transparent, and controllable AI systems.

Relevance 85 · Audience 95