AI News selected for Professionals and Decision Makers
Primary Research Stream

MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents

06:00 · July 7, 2026 · arXiv cs.AI RSS

MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents

Current benchmarks for evaluating large language models (LLMs) in medical calculation are largely based on simplified settings, where each patient case corresponds to a single calculator and the required tool is explicitly specified in the query. However, real clinical scenarios often require multiple calculators for joint evaluation, nested-scale calculation, and fuzzy queries that do not directly specify the target calculator. To this end, we propose a new medical calculation benchmark, MedCalc-Pro, which covers three progressively challenging task settings: single-calculator, multi-calculator, and nested-calculator calculation settings. MedCalc-Pro contains 2,268 real-world clinical cases, covering 77 medical calculators across 14 clinical departments. Meanwhile, to address the limited performance of existing frameworks and methods in complex clinical scenarios, we further propose a more generalizable agent framework that supports multi-tool selection and nested-tool calling, while suppressing parameter error propagation through structured validation and evidence review. We conduct systematic comparisons across open-source, closed-source, and medical-specialized LLMs, and the results show that our framework achieves the best performance across all three task settings. This work provides a new benchmark and method for evaluating and applying LLMs in challenging medical calculation scenarios.

Summary

Existing benchmarks for medical calculation with large language models typically assume simplified conditions in which each clinical case maps to a single, explicitly named calculator. Real patient records, however, frequently require several calculators to be applied together, sometimes with one calculator depending on the output of another, and clinicians rarely state the required tools in advance. MedCalc-Pro addresses these gaps by providing a benchmark that spans three graduated difficulty levels: single-calculator tasks, multi-calculator tasks that demand joint evaluation, and nested-calculator tasks that involve explicit tool dependencies.

The dataset comprises 2,268 real-world cases drawn from 14 clinical departments and covering 77 distinct medical calculators. Queries are formulated in a goal-driven style that withholds tool names, forcing models to infer the appropriate calculators from clinical intent. This design tests query understanding, tool selection, parameter extraction, and multi-step execution under conditions closer to actual practice than prior collections.

To improve performance on these harder settings, the authors introduce a generalizable agent framework organized into four stages: query rewriting to clarify clinical intent, retrieval and reranking to surface candidate calculators, tool selection that permits multiple or nested calls, and dependency-aware execution supported by structured validation and evidence review. The validation steps are intended to limit the propagation of parameter errors across sequential tool invocations.

Systematic experiments compare the framework against representative open-source, closed-source, and medically fine-tuned models. Results indicate that the proposed approach attains the strongest outcomes across all three task settings while exhibiting greater robustness when queries become less explicit or when calculator chains grow more complex.

Why it matters

This research is highly relevant for Dutch AI researchers and health-tech enterprises focusing on clinical decision support systems. The proposed benchmark and agent framework align with the Netherlands' strong emphasis on robust, validated, and ethical AI applications in healthcare.

More in this beat
ai-agentsevaluation-benchmarksllm-agentsllm-benchmarksMedCalc-Promedical-aitool-use
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

06:00 · August 4, 2026

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

This research is highly relevant for Dutch AI researchers developing autonomous LLM agents, providing a rigorous framework for evaluating continuous learning in realistic deployment scenarios. Understanding how model capabilities gate self-evolution is crucial for building robust and reliable AI systems.

Relevance 85 · Audience 95

AI Tool Discovery at Scale: All You Need is DNS

06:00 · July 22, 2026

AI Tool Discovery at Scale: All You Need is DNS

This research is highly relevant for Dutch AI infrastructure developers and researchers building multi-agent systems. Its decentralized governance model aligns well with European data sovereignty and transparent AI goals, offering a scalable alternative to centralized tool registries.

Relevance 85 · Audience 95

SAAG: Structured Agent Assessment and Grounding

06:00 · July 22, 2026

SAAG: Structured Agent Assessment and Grounding

This research provides a rigorous framework for diagnosing and mitigating hallucinations in AI agents, directly supporting the Dutch and EU focus on transparent and trustworthy AI. It offers researchers new methodologies to evaluate agentic systems beyond simple binary exact-match metrics.

Relevance 85 · Audience 95

Is it agentic enough? Benchmarking open models on your own tooling

02:00 · June 18, 2026

Is it agentic enough? Benchmarking open models on your own tooling

It provides ML Engineers with actionable insights and a new open-source tool to benchmark and optimize their own libraries for agentic use. Understanding the trade-offs in token consumption and latency across different model sizes is crucial for building cost-effective and reliable AI systems.

Relevance 85 · Audience 95

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

06:00 · August 13, 2026

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

This research is highly relevant for Dutch AI researchers and MLOps practitioners developing autonomous agents, as it provides a rigorous, open-source framework for testing agent reliability and safety. Its focus on risk assessment and operational side-effects aligns strongly with the Netherlands' strategic emphasis on transparent, ethical, and secure AI deployments.

Relevance 85 · Audience 95

Calibrated Selective Fact-Checking via Evidence Chain Evaluation

06:00 · July 22, 2026

Calibrated Selective Fact-Checking via Evidence Chain Evaluation

This research is highly relevant for Dutch AI researchers and practitioners focusing on trustworthy and ethical AI, a key priority in the Netherlands and the EU. The abstention mechanism directly addresses LLM hallucination and reliability issues, offering actionable methodologies for building compliant, high-stakes verification pipelines under EU AI regulations.

Relevance 85 · Audience 95

Deterministic Replay for AI Agent Systems

06:00 · July 21, 2026

Deterministic Replay for AI Agent Systems

Directly actionable for Dutch AI researchers and advanced practitioners working on agent systems, offering high technical depth, reproducibility resources, and alignment with EU emphasis on transparent, reliable AI.

Relevance 85 · Audience 90

Cura 1T: Specialized Model for Agentic Healthcare

06:00 · July 20, 2026

Cura 1T: Specialized Model for Agentic Healthcare

This research is highly relevant for Dutch AI researchers and healthcare institutions developing specialized clinical models. The data-centric, self-evolving training methodology offers a transparent and rigorous approach to building reliable healthcare AI, aligning with EU regulatory standards for clinical deployment.

Relevance 85 · Audience 95