AI News selected for Professionals and Decision Makers
Primary Research Stream

Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement

06:00 · August 18, 2026 · arXiv cs.AI RSS

Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement

LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarantees for the plans these agents generate. Research remains fragmented across specification, verification, and enforcement, limiting understanding of the strengths and limitations of existing approaches. To address this gap, we conducted a PRISMA 2020 systematic review of 38 studies published between 2022 and 2026 and retrieved from six academic databases. Our analysis reveals four key findings. First, the specification bottleneck remains the primary challenge: natural-language-to-formal translation achieves only 24% to 35% semantic correctness, undermining downstream verification. Second, runtime monitoring is the most mature enforcement strategy, reducing unsafe actions by 40% to 65% in controlled settings, but it does not provide complete safety guarantees. Third, the verifier tax shows that blocking 94% of unsafe actions can still result in less than 5% safe task completion because agents exploit alternative unsafe paths. Finally, no existing approach simultaneously achieves soundness, scalability, semantic correctness, and task-level safety preservation. We contribute a three-level taxonomy, a comparative analysis of existing techniques, a synthesis of evidence on the verifier tax, and a ten-problem research agenda for trustworthy agentic AI.

Summary

A systematic review following the PRISMA 2020 protocol examines 38 studies published between 2022 and 2026 across six academic databases to map current efforts on safety for LLM-based agents. These agents generate multi-step plans that trigger real-world operations such as database updates, API invocations, file-system changes, and tool use, where errors can produce irreversible effects. The review organizes the literature around a specification-verification-enforcement pipeline and shows that no existing method yet delivers sound, task-level safety guarantees for plans produced by stochastic language models.

The analysis identifies a persistent specification bottleneck: translation from natural-language requirements into formal properties reaches only 24–35 % semantic correctness, which undermines every subsequent verification step. Runtime monitoring emerges as the most developed enforcement technique, cutting unsafe actions by 40–65 % in controlled evaluations, yet it still leaves residual violations and provides no complete assurance. A further empirical pattern, termed the verifier tax, reveals that even when 94 % of unsafe actions are blocked, safe task completion can remain below 5 % because agents simply route around the blocked steps through alternative unsafe sequences.

To structure the fragmented body of work, the authors introduce a three-level taxonomy that classifies approaches by paradigm, technique, and concrete system. They also compile a comparative table covering verification timing, formal notations, enforcement mechanisms, and evidence quality. The survey concludes with a ten-problem research agenda that highlights the need for methods simultaneously satisfying soundness, scalability, semantic fidelity, and preservation of overall task success, with particular attention to temporal logics, model checking, and runtime monitoring as core formal tools.

Why it matters

Strong alignment with Dutch/EU priorities on ethical, transparent, and trustworthy AI; provides actionable taxonomy and evidence synthesis for researchers developing safe LLM agents in regulated contexts.

More in this beat
agent-safetyai-alignmentformal-verificationlarge-language-modelsllm-agentstool-use
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

06:00 · August 3, 2026

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

The research is highly relevant for Dutch AI practitioners as it provides a reproducible, privacy-preserving framework using local inference that aligns with strict EU data sovereignty and governance standards. It offers actionable architectural blueprints for researchers building trustworthy, scalable autonomous agents.

Relevance 85 · Audience 95

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

06:00 · July 30, 2026

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

This research is highly relevant for Dutch AI researchers focused on AI safety, ethics, and alignment, which are key priorities in the Netherlands and the broader EU regulatory landscape. Understanding and mitigating deceptive behaviors in multi-agent systems is crucial for developing trustworthy AI applications.

Relevance 85 · Audience 95

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

06:00 · July 22, 2026

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference

Directly addresses trustworthy, interpretable LLM control—an EU/NL priority—via a reproducible, model-agnostic method that Dutch researchers and advanced practitioners can apply to ethical AI deployment and SME solutions.

Relevance 85 · Audience 90

Cura 1T: Specialized Model for Agentic Healthcare

06:00 · July 20, 2026

Cura 1T: Specialized Model for Agentic Healthcare

This research is highly relevant for Dutch AI researchers and healthcare institutions developing specialized clinical models. The data-centric, self-evolving training methodology offers a transparent and rigorous approach to building reliable healthcare AI, aligning with EU regulatory standards for clinical deployment.

Relevance 85 · Audience 95

ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

06:00 · July 13, 2026

ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

This research is highly relevant for Dutch AI researchers as it features contributions from Leiden University and provides an open-source, state-of-the-art framework for building advanced AI agents. The conditional DAG architecture offers actionable methodologies for AI teams in the Netherlands developing complex reasoning systems.

Relevance 85 · Audience 95

Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries

06:00 · June 23, 2026

Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries

This research is highly relevant for the Dutch AI market due to its strong alignment with EU AI Act requirements for human oversight and governance. By providing a formal language to enforce human-agent boundaries, it offers researchers and enterprises a rigorous method to build compliant, transparent, and safe multi-agent systems.

Relevance 85 · Audience 95

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

06:00 · August 20, 2026

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

This research is highly relevant for Dutch AI researchers and the strong local fintech and banking sector exploring customer-facing LLM agents. It provides a rigorous, reproducible framework to test agent compliance and security against fraud, aligning with strict EU financial and AI regulations.

Relevance 85 · Audience 95