AI News selected for Professionals and Decision Makers
Primary Research Stream

PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

06:00 · July 2, 2026 · arXiv cs.AI RSS

PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex. We introduce PHREEQC-MCQ-200, a benchmark for evaluating tool-augmented agents on deterministic aqueous-geochemistry simulations. The benchmark contains 200 multiple-choice questions derived from 21 validated PHREEQC scenarios, requiring agents to construct simulator inputs, execute PHREEQC, inspect structured outputs, and commit to final answers. Across multiple frontier and mid-tier model families, simulator access substantially improves aggregate accuracy, confirming that grounded execution is necessary for many scientific-computation tasks. However, the gains are not monotonic: tool-augmented agents also lose items they answered correctly without tools, revealing regressions that average accuracy alone hides. We further show that output-access protocol matters. A table-of-contents interface can reduce token cost while preserving or improving accuracy for stronger models, but it degrades performance for mid-tier models that cannot reliably navigate structured simulator outputs. PHREEQC-MCQ-200 therefore frames scientific tool use as an end-to-end diagnostic problem rather than a simple tool-calling capability. We argue that evaluations of scientific agents should report not only accuracy, but also item-level retention, output-access sensitivity, trajectory failures, and where the computation chain breaks.

Summary

The article presents PHREEQC-MCQ-200, a benchmark of 200 multiple-choice questions drawn from 21 validated scenarios in the open-source PHREEQC simulator for aqueous geochemistry. Agents must generate valid input files, run the deterministic simulator, parse its structured output, and select the correct answer, providing a controlled setting in which tool use can be measured against reproducible ground truth rather than open-ended generation.

Experiments across frontier and mid-tier models show that simulator access raises aggregate accuracy by 15 to 41.5 percentage points compared with text-only baselines. The improvement confirms that many scientific-computation tasks require grounded execution rather than scaled chain-of-thought reasoning alone. At the same time, the gains are not monotonic: agents lose between 10 and 32 items they had previously solved without tools, yielding retention rates between 56 and 86 percent depending on the model.

The benchmark further isolates the effect of output-access protocols. A table-of-contents interface reduces input tokens by 11 to 57 percent while preserving or slightly improving accuracy for stronger models; the same interface degrades performance for mid-tier models that struggle to navigate the structured simulator output. These trade-offs cluster by capability tier across vendors rather than by model family.

The authors therefore treat scientific tool use as an end-to-end diagnostic problem. They recommend that evaluations of tool-augmented agents report item-level retention, output-access sensitivity, and the location of chain failures in addition to overall accuracy.

Why it matters

This research is highly relevant for Dutch AI researchers and enterprises developing agentic workflows for scientific discovery (AI4Science). Its focus on the reliability, transparency, and diagnostic evaluation of tool-augmented LLMs aligns well with the EU's push for trustworthy AI systems.

More in this beat
evaluation-benchmarksexperimental-benchmarksllm-agentsllm-benchmarksPHREEQCPHREEQC-MCQ-200scientific-discoverytool-use
ASI-Bench: At the Dawn of Artificial Superintelligence

06:00 · August 19, 2026

ASI-Bench: At the Dawn of Artificial Superintelligence

Offers a novel, high-depth evaluation framework that Dutch AI researchers and advanced labs can directly apply to measure progress toward autonomous scientific agents, aligning with the Netherlands' strengths in ethical AI and SME-driven innovation.

Relevance 62 · Audience 88

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

06:00 · August 4, 2026

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

This research is highly relevant for Dutch AI researchers developing autonomous LLM agents, providing a rigorous framework for evaluating continuous learning in realistic deployment scenarios. Understanding how model capabilities gate self-evolution is crucial for building robust and reliable AI systems.

Relevance 85 · Audience 95

MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents

06:00 · July 7, 2026

MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents

This research is highly relevant for Dutch AI researchers and health-tech enterprises focusing on clinical decision support systems. The proposed benchmark and agent framework align with the Netherlands' strong emphasis on robust, validated, and ethical AI applications in healthcare.

Relevance 85 · Audience 95

Is it agentic enough? Benchmarking open models on your own tooling

02:00 · June 18, 2026

Is it agentic enough? Benchmarking open models on your own tooling

It provides ML Engineers with actionable insights and a new open-source tool to benchmark and optimize their own libraries for agentic use. Understanding the trade-offs in token consumption and latency across different model sizes is crucial for building cost-effective and reliable AI systems.

Relevance 85 · Audience 95

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

06:00 · August 20, 2026

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

This research is highly relevant for Dutch AI researchers and the strong local fintech and banking sector exploring customer-facing LLM agents. It provides a rigorous, reproducible framework to test agent compliance and security against fraud, aligning with strict EU financial and AI regulations.

Relevance 85 · Audience 95

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

06:00 · August 15, 2026

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

This research is highly relevant for Dutch AI researchers and institutions focused on ethical AI deployment. It provides a concrete framework to evaluate and mitigate research misconduct risks when integrating LLMs into scientific workflows, aligning perfectly with the EU's emphasis on trustworthy AI.

Relevance 85 · Audience 95