AI News selected for Professionals and Decision Makers
Primary Research Stream

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

06:00 · August 3, 2026 · arXiv cs.AI RSS

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance on predefined tool spaces with static semantics limits their applicability to open-world scientific workflows, where tool requirements, capabilities, and boundaries evolve dynamically. To this end, we propose SciToolAgent-Evo, an ontology-aware self-evolving agent for open-world scientific tool acquisition. Driven by an evolving memory of skills, experiences, and an ontologized tool graph, it distills generalizable knowledge from contrastive trajectories during accumulation, whereas during inference, it formulates active requests and utilizes a LinUCB-based bandit gate to dynamically balance exploration and exploitation. Once a novel tool is acquired, its scientific ontology is completed online for seamless integration into the known graph. Moreover, we introduce OpenSciToolBench, a benchmark containing 900 realistic tasks across four difficulty levels. Extensive evaluations show that SciToolAgent-Evo achieves state-of-the-art performance, validating its robustness and generalization.

Summary

SciToolAgent-Evo addresses a core limitation of current LLM agents in scientific domains: their dependence on fixed tool libraries whose semantics and boundaries remain static. In open-world research settings, required capabilities such as descriptor calculators or format converters often lie outside the agent’s initial inventory, and conventional systems either fail or over-exploit the tools they already know. The proposed agent counters this by maintaining an evolving memory that combines a skill library, an experience pool, and an ontologized tool graph. During accumulation phases it extracts compact, reusable knowledge from pairs of successful and unsuccessful trajectories; at inference time it issues explicit requests for missing tools and routes decisions through a LinUCB bandit that continuously trades off exploration of unknown tools against exploitation of known ones.

When a new tool is obtained, its scientific ontology—covering inputs, outputs, domain relations, and usage constraints—is completed online and merged into the existing graph, enabling immediate reuse in subsequent tasks. This self-evolution loop removes the need for manual ontology updates and supports continual adaptation across biology, chemistry, and materials-science workflows.

To evaluate such capabilities the authors introduce OpenSciToolBench, a 900-task benchmark derived from SciToolKG. Tasks are stratified into four difficulty levels that progress from direct tool invocation (one or two tools) to open-ended research-strategy formulation (five to ten tools), and they appear in four formats: question answering, multiple choice, content completion, and true/false. Construction follows a two-stage verification pipeline that combines LLM judging with human review for tool-chain coherence and scientific validity. Experimental results across this benchmark and prior scientific tool-use suites indicate that SciToolAgent-Evo attains state-of-the-art performance, with ablation studies confirming the contribution of the evolving memory, active requests, and bandit-based exploration gate.

Why it matters

This research is highly relevant for Dutch AI researchers and institutions focusing on autonomous agents and AI-driven scientific discovery. The methodologies for open-world tool acquisition and the new benchmark provide actionable frameworks for advancing LLM applications in fields like chemistry and materials science.

More in this beat
chemical-synthesisllm-agentsmaterials-scienceOpenSciToolBenchscientific-discoverySciToolAgent-Evoself-evolving-agentstool-use
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

06:00 · August 4, 2026

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

This research is highly relevant for Dutch AI researchers developing autonomous LLM agents, providing a rigorous framework for evaluating continuous learning in realistic deployment scenarios. Understanding how model capabilities gate self-evolution is crucial for building robust and reliable AI systems.

Relevance 85 · Audience 95

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

06:00 · July 9, 2026

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

This research is highly relevant for Dutch AI researchers and enterprise developers building autonomous agents, as it offers a novel method to reduce reasoning overhead and API costs while improving reliability. The transition from static tools to self-evolving SOPs aligns well with the Dutch market's focus on scalable, efficient AI automation for SMEs.

Relevance 85 · Audience 95

Optimal Resource Utilization for Autonomous Laboratory Orchestrators

06:00 · July 2, 2026

Optimal Resource Utilization for Autonomous Laboratory Orchestrators

The research is highly relevant for Dutch R&D sectors, particularly in materials science, chemistry, and high-tech manufacturing, where autonomous laboratories can significantly accelerate innovation. It provides actionable methodologies for AI researchers and engineers looking to optimize hardware orchestration and resource management in automated experimental setups.

Relevance 75 · Audience 90

Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective

06:00 · August 20, 2026

Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective

The paper provides foundational research on making autonomous AI agents auditable, safe, and transparent through dynamic graph modeling. This aligns strongly with the Dutch and EU focus on ethical AI and regulatory compliance, offering advanced researchers actionable frameworks for building governable agentic systems.

Relevance 85 · Audience 95

Harnessing agent memory to build lifelong AI partners for materials scientists

06:00 · August 13, 2026

Harnessing agent memory to build lifelong AI partners for materials scientists

This research is highly relevant for Dutch AI researchers and high-tech materials enterprises looking to deploy autonomous AI agents for R&D. The proposed model-agnostic memory framework addresses critical challenges in AI reproducibility and workflow efficiency, offering actionable methodologies for advanced scientific computing.

Relevance 85 · Audience 95

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

06:00 · August 3, 2026

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

The research is highly relevant for Dutch AI practitioners as it provides a reproducible, privacy-preserving framework using local inference that aligns with strict EU data sovereignty and governance standards. It offers actionable architectural blueprints for researchers building trustworthy, scalable autonomous agents.

Relevance 85 · Audience 95