AI News selected for Professionals and Decision Makers
Primary Research Stream

RoPoLL: Robust Panel of LLM Judges

06:00 · July 1, 2026 · arXiv cs.AI RSS

RoPoLL: Robust Panel of LLM Judges

The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its statistical behavior remains poorly understood. We formalize the LLM Jury under the Huber contamination model and show that PoLL incurs unbounded bias under any positive contamination, regardless of jury size, whenever a single judge fails in a biased, LLM-typical way (mode collapse, sycophancy, safety refusal). Framing jury consensus as classical robust mean estimation, we propose RoPoLL (Robust Panel of LLM-as-Judge), which preserves the PoLL panel but replaces the aggregation function with a robust mean estimator, instantiated with the geometric median (GM): tuning-free, with the optimal finite-sample breakdown point 1/2. A finite-sample error bound and a matching information-theoretic minimax lower bound agree on the parametric rate sigma*sqrt(d/N) and differ on the breakdown floor by a factor of sqrt(d), a statistical-computational gap that polynomial-time RoPoLL pays relative to the intractable Tukey halfspace median. Across 13 open-weight judges (4B-675B), three reward-model benchmarks, and four corruption regimes at rates up to 50%, RoPoLL dominates PoLL on every biased corruption type: by about 19% on cross-dimensional attacks at matched compute, and by orders of magnitude on heavy-tailed Byzantine adversaries. A 3-judge RoPoLL committee at 38B beats Mistral-Large-3 (675B) by 1.31x on HelpSteer-2 under 30% bimodal-random corruption, an 18x parameter advantage at better accuracy; a Noisy-GT control confirms the premium is paid against biased contamination, not benign imprecision.

Summary

RoPoLL addresses a core weakness in panel-based LLM evaluation by treating the jury as a statistical estimation problem under the Huber contamination model. Standard PoLL approaches aggregate scores from multiple judges via arithmetic mean, which works when errors are light-tailed and symmetric but produces unbounded bias as soon as any judge exhibits the biased, LLM-typical failures common in practice: mode collapse, sycophancy, cross-attribute confusion, or heavy-tailed parser hallucinations. These failures map directly onto the contamination distribution in the Huber framework, and a direct calculation shows that the resulting bias grows linearly with the corruption shift regardless of jury size.

The method keeps the same diverse panel of smaller open-weight judges but replaces the mean with the geometric median as the aggregation step. This estimator is tuning-free, operates on the full score vector to preserve cross-attribute structure, and attains the optimal finite-sample breakdown point of one half. A finite-sample error bound of order sigma times square root of d over N is derived, together with a matching information-theoretic minimax lower bound that reveals a statistical-computational gap of order square root of d between the geometric median and the intractable Tukey halfspace median.

Extensive experiments across thirteen open-weight judges ranging from 4 B to 675 B parameters, three reward-model benchmarks, and four corruption regimes up to fifty percent contamination show that RoPoLL consistently outperforms arithmetic-mean aggregation on every biased attack. Gains reach roughly nineteen percent on cross-dimensional corruption at matched compute and orders of magnitude on heavy-tailed Byzantine adversaries. A three-judge RoPoLL committee with thirty-eight billion total parameters exceeds a 675-billion-parameter single judge on HelpSteer-2 under thirty percent bimodal-random corruption, confirming that the robustness premium arises specifically against structured contamination rather than benign noise. The authors also release the full judge-output corpus to support further calibration and aggregation research.

Why it matters

Directly actionable for Dutch research teams and SMEs building LLM evaluation pipelines; aligns with EU emphasis on reliable and transparent AI; high technical depth and novelty for advanced readers.

More in this beat
evaluation-benchmarkslarge-language-modelsllm-as-judgellm-judgesmode-collapseropoll
Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

06:00 · August 17, 2026

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

This research is highly relevant for Dutch AI researchers and enterprises focused on developing trustworthy and transparent AI systems. By reducing false-pass rates in automated agent evaluation, it provides a robust methodology that aligns with the EU's stringent requirements for AI reliability and safety.

Relevance 85 · Audience 95

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

06:00 · August 15, 2026

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

This research is highly relevant for Dutch AI researchers and institutions focused on ethical AI deployment. It provides a concrete framework to evaluate and mitigate research misconduct risks when integrating LLMs into scientific workflows, aligning perfectly with the EU's emphasis on trustworthy AI.

Relevance 85 · Audience 95

Position: Reasoning is a Learnable Rule-Based Process

06:00 · August 15, 2026

Position: Reasoning is a Learnable Rule-Based Process

Directly supports Dutch/EU priorities on ethical, transparent, and trustworthy AI by clarifying reasoning evaluation, which aids practitioners in building auditable systems compliant with regulations like the AI Act.

Relevance 75 · Audience 90

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

06:00 · August 13, 2026

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

SAPO provides a highly actionable, structured approach to prompt engineering that Dutch AI researchers and enterprise teams can use to build more reliable and interpretable LLM applications. Its focus on modular, non-destructive prompt updates aligns with the EU's demand for robust, transparent, and controllable AI systems.

Relevance 85 · Audience 95

Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

06:00 · August 7, 2026

Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

This paper is highly relevant for AI researchers in the Netherlands focusing on LLM reasoning, alignment, and compute-efficient training. The proposed weak-to-strong distillation method offers actionable insights for Dutch AI labs aiming to enhance model performance without relying solely on massive scaling.

Relevance 85 · Audience 95

TriQua: Reconciling Granularity and Context in Factuality Evaluation

06:00 · August 7, 2026

TriQua: Reconciling Granularity and Context in Factuality Evaluation

This research is highly relevant for Dutch AI researchers and practitioners focused on trustworthy AI and LLM deployment. Improving factuality evaluation directly supports the Netherlands and EU strategic emphasis on transparent, reliable, and ethical AI systems.

Relevance 85 · Audience 95