AI News selected for Professionals and Decision Makers
Primary Research Stream

DEMM-Bench: A Cross-Regime Benchmark for Agent-Runtime Governance-Evidence Sufficiency

06:00 · June 23, 2026 · arXiv cs.AI RSS

DEMM-Bench: A Cross-Regime Benchmark for Agent-Runtime Governance-Evidence Sufficiency

Agent-runtime systems emit traces, ledgers, provenance graphs, policy logs, delegation tokens, cache events, and tool-firewall records, but those containers do not necessarily answer governance questions about a specific decision. DEMM-Bench is a cross-regime benchmark for agent-runtime governance-evidence sufficiency, grounded in the Decision Evidence Maturity Model (DEMM): it measures whether records across eight evidence regimes are sufficient to reconstruct decision-level properties rather than merely present. The benchmark normalizes the regimes through adapters, asks property questions over actor, authority, action, policy, decision basis, resource touch, lifecycle context, and verification strength, and applies eight deterministic degradation conditions. Across 64 manuscript cases, trace-present and schema-present baselines overclaim on 75% of cases, ledger-present overclaims on 50%, and the redacted property-level candidate scorer has zero overclaim with 56.25% mean Property Sufficiency Accuracy. The deposited package provides the 64-case dataset, construction-oracle labels, baselines, and adapters, supporting reproducible evaluation of decision-evidence maturity across heterogeneous agent-runtime evidence substrates.

Summary

Agent-runtime systems generate diverse outputs such as traces, ledgers, provenance graphs, policy logs, delegation tokens, cache events, and tool-firewall records. These outputs often fail to resolve concrete governance questions about individual decisions, even when the records themselves are present. DEMM-Bench addresses this gap by providing a benchmark that tests whether evidence across eight distinct regimes is sufficient to reconstruct decision-level properties rather than simply confirming that some record exists.

The benchmark rests on the Decision Evidence Maturity Model and normalizes the eight regimes through purpose-built adapters. It poses targeted questions about actor identity, authority, action taken, applicable policy, decision basis, resource access, lifecycle context, and verification strength. Eight deterministic degradation conditions are then applied to each of the 64 manuscript cases, allowing controlled measurement of how evidence quality affects the ability to answer those questions.

Baseline methods that rely on trace presence or schema presence overclaim sufficiency in 75 percent of cases, while ledger presence alone overclaims in half of the cases. In contrast, a redacted property-level candidate scorer records zero overclaims and achieves a mean Property Sufficiency Accuracy of 56.25 percent. The accompanying release supplies the full 64-case dataset, construction-oracle labels, baselines, and adapters, supporting reproducible assessment of evidence maturity across heterogeneous agent-runtime substrates.

Why it matters

Directly supports EU AI Act transparency and accountability requirements; Dutch researchers and advanced practitioners can apply the benchmark to evaluate agent systems for regulatory compliance and ethical AI deployment.

More in this beat
agent-safetyai-agentsdata-security-governanceDEMM-Benchevaluation-benchmarksexperimental-benchmarksllm-agents
FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

06:00 · August 20, 2026

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

This research is highly relevant for Dutch AI researchers and the strong local fintech and banking sector exploring customer-facing LLM agents. It provides a rigorous, reproducible framework to test agent compliance and security against fraud, aligning with strict EU financial and AI regulations.

Relevance 85 · Audience 95

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

06:00 · June 23, 2026

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

This research directly supports the Dutch and EU strategic focus on safe, ethical, and reliable AI deployment. For researchers and advanced practitioners in the Netherlands, it provides actionable methodologies to build autonomous agents that comply with stringent safety constraints in dynamic environments.

Relevance 85 · Audience 90

Is it agentic enough? Benchmarking open models on your own tooling

02:00 · June 18, 2026

Is it agentic enough? Benchmarking open models on your own tooling

It provides ML Engineers with actionable insights and a new open-source tool to benchmark and optimize their own libraries for agentic use. Understanding the trade-offs in token consumption and latency across different model sizes is crucial for building cost-effective and reliable AI systems.

Relevance 85 · Audience 95

Measuring Cross-Task Behavioral Consistency in Language Model Agents

06:00 · August 17, 2026

Measuring Cross-Task Behavioral Consistency in Language Model Agents

The article provides a novel, quantifiable method for assessing the reliability and behavioral consistency of AI agents, which is crucial for compliance with EU AI regulations and the Dutch focus on transparent AI. Researchers can directly apply the open-source BCM framework to evaluate and improve the predictability of enterprise AI deployments.

Relevance 85 · Audience 95

SAAG: Structured Agent Assessment and Grounding

06:00 · July 22, 2026

SAAG: Structured Agent Assessment and Grounding

This research provides a rigorous framework for diagnosing and mitigating hallucinations in AI agents, directly supporting the Dutch and EU focus on transparent and trustworthy AI. It offers researchers new methodologies to evaluate agentic systems beyond simple binary exact-match metrics.

Relevance 85 · Audience 95

AI Tool Discovery at Scale: All You Need is DNS

06:00 · July 22, 2026

AI Tool Discovery at Scale: All You Need is DNS

This research is highly relevant for Dutch AI infrastructure developers and researchers building multi-agent systems. Its decentralized governance model aligns well with European data sovereignty and transparent AI goals, offering a scalable alternative to centralized tool registries.

Relevance 85 · Audience 95