AI News selected for Professionals and Decision Makers
Primary Research Stream

LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation

06:00 · July 29, 2026 · arXiv cs.AI RSS

LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation

Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective. In this paper, our question is: \emph{Do LLMs know what other LLMs don't? And can we leverage this dynamic for evaluation?} We present \textbf{LivingArena}, an automated, contamination-resistant evaluation framework. In this framework, models take turns proposing questions, aiming to pose items that opponents cannot answer correctly. Questioners are encouraged to actively identify and exploit opponents' knowledge boundaries, receiving rewards when the answerer fails, while the answerer is rewarded otherwise. To ensure questions contain objectively verifiable answers, a judge panel of strong models validates them, penalizing questioners if validation fails. Evaluating ten frontier LLMs, LivingArena yields a stable Elo leaderboard. Our behavioral analyses show that models identify and exploit their peers' cognitive boundaries: self-play and tournament logs indicate that they localize and double down on opponents' weak dimensions. Beyond static knowledge recall, peer probing measures factual rigor and the higher-order ability to probe an opponent's weaknesses, correlating only weakly with human preference and offering a scalable, low-cost approach to continuous evaluation.

Summary

LivingArena addresses the core limitations of static benchmarks such as MMLU and MMLU-Pro, which suffer from data contamination through leakage into training sets, rapid saturation that collapses distinctions among leading models, and high expert authoring costs. Human-preference arenas avoid fixed test sets but rely on subjective ratings that do not isolate factual accuracy or reasoning failures. The framework instead lets frontier models probe one another directly by generating questions intended to expose verifiable gaps in an opponent’s knowledge or reasoning.

In each round, one model acts as questioner and must supply both a question and a correct reference answer; a panel of strong models then validates the item through majority consensus, rejecting any question that fails verification and penalizing the questioner. Validated questions proceed to the answerer, with scoring awarded only when the opponent fails. This zero-sum dynamic incentivizes models to localize and repeatedly target an opponent’s weaker dimensions, such as logic or abstract reasoning, rather than sampling uniformly across topics. The resulting tournament produces an Elo ranking from objective, binary outcomes instead of stylistic comparisons.

When ten models from OpenAI, Anthropic, Google, and DeepSeek competed in 360 matches, the system generated a stable leaderboard while revealing distinct strategies: some models aggressively test difficult items at the risk of self-harm penalties, while others remain more calibrated. Performance diverged most clearly in quantitative reasoning and code, with consistent weakness in logic. The approach correlates only weakly with human preference scores, indicating that peer probing captures a distinct dimension of factual rigor and opponent modeling. A full ten-model evaluation completes in roughly 100 minutes of wall-clock time and requires no static question bank or human annotation.

Why it matters

This research provides Dutch AI researchers and developers with a novel, open-source framework for dynamically evaluating LLMs, addressing critical challenges like benchmark saturation and data contamination. Its rigorous, automated testing methodology aligns well with the EU's growing emphasis on robust AI evaluation and compliance.

More in this beat
anthropicevaluation-benchmarksfrontier-modelslarge-language-modelsLivingArenallm-benchmarksmmlu-proopenai
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

06:00 · August 15, 2026

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

This research is highly relevant for Dutch AI researchers and institutions focused on ethical AI deployment. It provides a concrete framework to evaluate and mitigate research misconduct risks when integrating LLMs into scientific workflows, aligning perfectly with the EU's emphasis on trustworthy AI.

Relevance 85 · Audience 95

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

06:00 · August 4, 2026

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

Directly actionable for Dutch AI researchers and advanced practitioners developing or continuously evaluating LLMs: reduces compute overhead while preserving ranking fidelity, aligns with EU emphasis on efficient and transparent AI, and requires no historical logs.

Relevance 72 · Audience 88

Chinese military researchers tap US AI models to train defense systems

14:34 · July 31, 2026

Chinese military researchers tap US AI models to train defense systems

Directly addresses military AI applications, dual-use model distillation, and NATO-relevant export control challenges, providing actionable insights for defense technologists and strategists on adversary capabilities and technology transfer risks.

Relevance 85 · Audience 90

Secure Code Warrior Research Reveals AI-Generated Code Introduces an Average of 15 Vulnerabilities Per Codebase

15:18 · July 21, 2026

Secure Code Warrior Research Reveals AI-Generated Code Introduces an Average of 15 Vulnerabilities Per Codebase

This research provides crucial empirical data on the security risks of AI-assisted development, directly supporting the Dutch AI market's focus on secure, ethical, and transparent AI deployment. It offers actionable insights for Dutch researchers and CISOs to benchmark LLMs and implement necessary guardrails in enterprise software development.

Relevance 85 · Audience 90

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

06:00 · July 14, 2026

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

This research is highly relevant for Dutch AI researchers and enterprises developing LLMs, as it offers a mathematically rigorous method to drastically reduce the computational cost and time required for model evaluation. This aligns with the European and Dutch focus on sustainable, resource-efficient AI development (Green AI).

Relevance 85 · Audience 95

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

06:00 · July 9, 2026

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

This research provides a rigorous, reproducible benchmark for evaluating LLM reasoning and interactive capabilities, which is highly relevant for Dutch AI researchers developing reliable and transparent AI systems. It directly supports the advancement of agentic AI by testing a model's ability to recognize its own knowledge gaps.

Relevance 85 · Audience 95

APeB: Benchmarking Personalization Ability of Large Language Model Agents

06:00 · July 7, 2026

APeB: Benchmarking Personalization Ability of Large Language Model Agents

This research provides a valuable benchmark and methodology for Dutch AI researchers and enterprises, particularly in e-commerce and customer service, developing personalized LLM agents. Improving intent discovery from user histories directly impacts the effectiveness of AI-driven consumer applications prevalent in the Netherlands.

Relevance 75 · Audience 90

Towards Evaluation of Implicit Software World Models in Coding LLMs

06:00 · June 29, 2026

Towards Evaluation of Implicit Software World Models in Coding LLMs

It provides AI researchers with a new framework for evaluating coding LLMs beyond standard metrics. For the Dutch AI ecosystem, which emphasizes efficient and robust AI engineering, improving how models predict execution resources is crucial for developing sustainable and optimized software.

Relevance 75 · Audience 90