AI News selected for Professionals and Decision Makers
Primary Research Stream

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

06:00 · July 2, 2026 · arXiv cs.AI RSS

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics. To address this generalization gap, we formalize OpenAgent (Tool-Use Agent in Open-World), a problem setting characterized by distributional shifts across query, action, observation, and domain dimensions. To systematically diagnose its impact, we construct a controlled sandbox environment where we define fine-grained environmental shifts across a four-tier hierarchy, Perception, Interaction, Reasoning, and Internalization, and conduct a comprehensive series of experiments. Our analysis yields a series of key insights, demonstrating that agents trained via both Supervised Fine-Tuning(SFT) and Reinforcement Learning suffer from varying degrees of performance degradation when confronting open environmental shifts. Building on these insights, we propose Perturbation-Augmented Fine-Tuning, a disturbance-based intervention strategy for SFT that lays the foundation for enhancing agent robustness and utility in realistic environments. Our code will be released at: https://github. com/LAMDA-NeSy/OpenAgent.

Summary

Large Language Model agents achieve strong results on static tool-use benchmarks yet encounter sharp performance drops once deployed in non-stationary environments where user queries, available tools, interaction outcomes, and task domains can all change. The paper formalizes this setting as OpenAgent, a problem defined by distributional shifts along four axes: query intent, action space, observation dynamics, and domain. To isolate these effects from the noise of live APIs, the authors built a controlled sandbox that injects perturbations at four diagnostic levels—Perception, Interaction, Reasoning, and Internalization—while preserving a clean closed-set baseline.

Experiments with both Supervised Fine-Tuning and Reinforcement Learning agents show consistent degradation under these shifts. SFT models tend to overfit training trajectories and anchor too rigidly to surface-level symbols, while RL agents, though better at semantic grounding, suffer from boundary blindness induced by reward structures that emphasize goal completion over robustness. These failure modes compound along multi-step trajectories because an early misstep alters subsequent observations and policy decisions.

To address the observed fragility, the authors introduce Perturbation-Augmented Fine-Tuning, a data-centric intervention that injects controlled observation anomalies and symbolic noise into SFT trajectories. The resulting models exhibit improved resilience to open-world shifts without sacrificing performance on the original static tasks. The work provides both a reproducible evaluation framework and concrete evidence that current post-training paradigms remain insufficient for realistic tool-use deployment.

Why it matters

This research is highly relevant for Dutch AI researchers and developers building autonomous agents, as it addresses critical robustness and generalization challenges in real-world tool use. Improving agent reliability aligns with the EU's focus on trustworthy AI, making the proposed fine-tuning strategies actionable for enterprise AI deployments in the Netherlands.

More in this beat
ai-agentsdeployment-readinessevaluation-benchmarksOpenAgentpeft-and-fine-tuningreinforcement-learningsupervised-fine-tuningtool-use
SAAG: Structured Agent Assessment and Grounding

06:00 · July 22, 2026

SAAG: Structured Agent Assessment and Grounding

This research provides a rigorous framework for diagnosing and mitigating hallucinations in AI agents, directly supporting the Dutch and EU focus on transparent and trustworthy AI. It offers researchers new methodologies to evaluate agentic systems beyond simple binary exact-match metrics.

Relevance 85 · Audience 95

AI Tool Discovery at Scale: All You Need is DNS

06:00 · July 22, 2026

AI Tool Discovery at Scale: All You Need is DNS

This research is highly relevant for Dutch AI infrastructure developers and researchers building multi-agent systems. Its decentralized governance model aligns well with European data sovereignty and transparent AI goals, offering a scalable alternative to centralized tool registries.

Relevance 85 · Audience 95

Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents

06:00 · July 8, 2026

Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents

This synthesis is highly relevant for Dutch AI researchers and developers building autonomous agents, as it provides a structured understanding of current LLM limitations. Its focus on safety, security, and measurement validity aligns strongly with the Netherlands' and EU's regulatory emphasis on robust, transparent, and ethical AI systems.

Relevance 85 · Audience 95

MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents

06:00 · July 7, 2026

MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents

This research is highly relevant for Dutch AI researchers and health-tech enterprises focusing on clinical decision support systems. The proposed benchmark and agent framework align with the Netherlands' strong emphasis on robust, validated, and ethical AI applications in healthcare.

Relevance 85 · Audience 95

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

06:00 · June 29, 2026

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

This research is highly relevant for AI researchers and advanced practitioners in the Netherlands developing autonomous LLM agents. The proposed training paradigm offers actionable methodologies to overcome the reactive limitations of current agents, aligning with the Dutch focus on advanced, capable, and reliable AI systems.

Relevance 85 · Audience 95

MosaicLeaks: Can your research agent keep a secret?

20:13 · June 18, 2026

MosaicLeaks: Can your research agent keep a secret?

Directly addresses production challenges for ML Engineers building agents: privacy leakage via queries, balancing accuracy vs. data exposure, and sample-efficient RL training. Strong quantitative benchmarks and actionable training recipe. EU GDPR relevance for Dutch enterprises handling sensitive data.

Relevance 78 · Audience 85

Position: Behavioral Systems Require Behavioral Tests

06:00 · August 20, 2026

Position: Behavioral Systems Require Behavioral Tests

The article is highly relevant for Dutch AI researchers and practitioners focused on ethical and transparent AI. By proposing behavioral tests to evaluate AI alignment, safety, and decision-making processes, it provides a crucial methodological framework that supports compliance with EU regulations like the AI Act and advances responsible AI deployment.

Relevance 85 · Audience 95