AI News selected for Professionals and Decision Makers
Primary Research Stream

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

06:00 · July 14, 2026 · arXiv cs.AI RSS

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions. We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced by wrapper choice, and the Parseability Sensitivity Index (PSI), the corresponding range in answer parseability. Across 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by over 30x across models and is largely explained by compliance failures. A fixed-effects regression shows that parseability remains a strong predictor of accuracy even after controlling for task, model, and wrapper. We argue that reporting accuracy without wrapper variance and compliance is statistically fragile, and we give practical recommendations for both benchmarking and structured-output deployments.

Summary

Prompt wrappers—templates that add formatting instructions such as JSON schemas, key-value delimiters, or scratchpad sections around otherwise identical task content—can alter LLM benchmark scores by amounts large enough to reverse model rankings. The study quantifies this effect through two range-based metrics. The Format Sensitivity Index measures the spread in accuracy across five wrapper families, while the Parseability Sensitivity Index tracks the corresponding spread in the fraction of outputs that an automated extractor can map to a canonical answer. Both are evaluated under a token-controlled protocol that pads prompts to a fixed character budget and logs residual token variation, across 140,000 generations on seven question-answering tasks and four instruct models ranging from 7 B to 72 B parameters.

Results reveal extreme model-dependent differences. Mean FSI varies by more than thirtyfold, with Qwen-2.5-72B remaining nearly invariant (mean FSI 0.024) while Phi-4 exhibits a mean FSI of 0.763; normalized sensitivity accentuates the contrast further. A weighted least-squares regression with fixed effects for model, task, and wrapper shows that parseability remains a strong predictor of accuracy (coefficient 0.819) even after these controls, indicating that most wrapper-induced score changes stem from outright compliance failures rather than subtle reasoning effects. Bootstrap confidence intervals and a matched low-token-spread subset corroborate the robustness of these patterns.

The authors conclude that single-point accuracy figures reported without wrapper variance or compliance statistics are statistically fragile. They recommend that benchmarking suites track distributions over wrapper families and that structured-output deployments adopt explicit parseability monitoring or constrained decoding to reduce the observed sensitivity.

Why it matters

Directly actionable for Dutch AI researchers and labs evaluating or deploying LLMs: wrapper variance can flip leaderboard results and affect real structured-output reliability, aligning with EU emphasis on transparent, reproducible AI.

More in this beat
evaluation-benchmarksexperimental-benchmarksFormat Sensitivity Indexllm-benchmarksphi-4prompt wrappersqwen
ASI-Bench: At the Dawn of Artificial Superintelligence

06:00 · August 19, 2026

ASI-Bench: At the Dawn of Artificial Superintelligence

Offers a novel, high-depth evaluation framework that Dutch AI researchers and advanced labs can directly apply to measure progress toward autonomous scientific agents, aligning with the Netherlands' strengths in ethical AI and SME-driven innovation.

Relevance 62 · Audience 88

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

06:00 · July 9, 2026

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

This research provides a rigorous, reproducible benchmark for evaluating LLM reasoning and interactive capabilities, which is highly relevant for Dutch AI researchers developing reliable and transparent AI systems. It directly supports the advancement of agentic AI by testing a model's ability to recognize its own knowledge gaps.

Relevance 85 · Audience 95

Measuring Intelligence Beyond Human Scale

06:00 · July 9, 2026

Measuring Intelligence Beyond Human Scale

This research is highly relevant for Dutch AI researchers and auditors developing robust evaluation frameworks for advanced AI systems, aligning with the EU AI Act's focus on rigorous model benchmarking. It offers a scalable solution to the saturation of current human-authored benchmarks.

Relevance 85 · Audience 95

Towards Evaluation of Implicit Software World Models in Coding LLMs

06:00 · June 29, 2026

Towards Evaluation of Implicit Software World Models in Coding LLMs

It provides AI researchers with a new framework for evaluating coding LLMs beyond standard metrics. For the Dutch AI ecosystem, which emphasizes efficient and robust AI engineering, improving how models predict execution resources is crucial for developing sustainable and optimized software.

Relevance 75 · Audience 90

Is it agentic enough? Benchmarking open models on your own tooling

02:00 · June 18, 2026

Is it agentic enough? Benchmarking open models on your own tooling

It provides ML Engineers with actionable insights and a new open-source tool to benchmark and optimize their own libraries for agentic use. Understanding the trade-offs in token consumption and latency across different model sizes is crucial for building cost-effective and reliable AI systems.

Relevance 85 · Audience 95

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

06:00 · August 20, 2026

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

This research is highly relevant for Dutch AI researchers and the strong local fintech and banking sector exploring customer-facing LLM agents. It provides a rigorous, reproducible framework to test agent compliance and security against fraud, aligning with strict EU financial and AI regulations.

Relevance 85 · Audience 95

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

06:00 · August 18, 2026

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

This article is highly relevant for Dutch AI researchers and practitioners focused on ethical AI, aligning strongly with the Netherlands' and EU's emphasis on transparent and trustworthy AI systems. It provides a critical framework for advancing LLM evaluation beyond simple value alignment toward robust normative reasoning.

Relevance 85 · Audience 95