DEMM-Bench: A Cross-Regime Benchmark for Agent-Runtime Governance-Evidence Sufficiency
06:00 · June 23, 2026 · arXiv cs.AI RSS

Agent-runtime systems emit traces, ledgers, provenance graphs, policy logs, delegation tokens, cache events, and tool-firewall records, but those containers do not necessarily answer governance questions about a specific decision. DEMM-Bench is a cross-regime benchmark for agent-runtime governance-evidence sufficiency, grounded in the Decision Evidence Maturity Model (DEMM): it measures whether records across eight evidence regimes are sufficient to reconstruct decision-level properties rather than merely present. The benchmark normalizes the regimes through adapters, asks property questions over actor, authority, action, policy, decision basis, resource touch, lifecycle context, and verification strength, and applies eight deterministic degradation conditions. Across 64 manuscript cases, trace-present and schema-present baselines overclaim on 75% of cases, ledger-present overclaims on 50%, and the redacted property-level candidate scorer has zero overclaim with 56.25% mean Property Sufficiency Accuracy. The deposited package provides the 64-case dataset, construction-oracle labels, baselines, and adapters, supporting reproducible evaluation of decision-evidence maturity across heterogeneous agent-runtime evidence substrates.
Summary
Agent-runtime systems generate diverse outputs such as traces, ledgers, provenance graphs, policy logs, delegation tokens, cache events, and tool-firewall records. These outputs often fail to resolve concrete governance questions about individual decisions, even when the records themselves are present. DEMM-Bench addresses this gap by providing a benchmark that tests whether evidence across eight distinct regimes is sufficient to reconstruct decision-level properties rather than simply confirming that some record exists.
The benchmark rests on the Decision Evidence Maturity Model and normalizes the eight regimes through purpose-built adapters. It poses targeted questions about actor identity, authority, action taken, applicable policy, decision basis, resource access, lifecycle context, and verification strength. Eight deterministic degradation conditions are then applied to each of the 64 manuscript cases, allowing controlled measurement of how evidence quality affects the ability to answer those questions.
Baseline methods that rely on trace presence or schema presence overclaim sufficiency in 75 percent of cases, while ledger presence alone overclaims in half of the cases. In contrast, a redacted property-level candidate scorer records zero overclaims and achieves a mean Property Sufficiency Accuracy of 56.25 percent. The accompanying release supplies the full 64-case dataset, construction-oracle labels, baselines, and adapters, supporting reproducible assessment of evidence maturity across heterogeneous agent-runtime substrates.
Why it matters
Directly supports EU AI Act transparency and accountability requirements; Dutch researchers and advanced practitioners can apply the benchmark to evaluate agent systems for regulatory compliance and ethical AI deployment.







