AI News selected for Professionals and Decision Makers
Primary Research Stream

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

06:00 · August 20, 2026 · arXiv cs.AI RSS

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. Banking is the clearest case: the same agent that answers a question can also change contact details, reset a PIN, or move money, so ordinary customer service is inseparable from authorization, fraud detection, and policy compliance. Existing financial-fraud benchmarks classify static transactions or messages, and general agent-safety benchmarks target prompt injection or generic harmful use; none test whether a policy-grounded banking agent safely acts when a caller manipulates identity, authorization, and trust over a conversation. We introduce FraudBench, an executable benchmark built on the $\tau^2$-bench dual-control framework and the $\tau$-Knowledge banking environment. Both the agent and the simulated caller act through tools over shared, mutable account state, and the agent may grant the caller access to selected tools; the environment exposes a 698-document internal policy corpus that the agent must retrieve from. FraudBench contains 150 authored adversarial scenarios; a frozen public set of 107 (90 across ten fraud mechanisms plus 17 chained adaptive attacks) is used for all reported runs, with 43 further chained attacks held out. Safety is history-dependent: single-control tasks satisfy every precondition but one, and adaptive attacks make a later, locally valid request unsafe because of an earlier probe, admission, or failed attempt. Each scenario is annotated with observable evidence, prohibited actions, safe dispositions, and intervention points. A preliminary single-trial evaluation of four agents on the 107 graded tasks yields attack-security between 49\% and 65\%, with money-mule and first-party fraud the most common cross-model weaknesses.

Summary

FraudBench addresses a gap in existing evaluations of conversational agents by testing whether policy-grounded banking systems can resist fraud when callers manipulate identity, authorization, and trust through extended dialogue. Current financial-fraud benchmarks focus on static transaction classification or message-level resistance, while general agent-safety suites examine prompt injection or generic harmful requests. None examine agents that hold privileged tools over customer databases and internal policy documents, where a single unsafe action can produce direct financial loss.

The benchmark builds on the τ²-bench dual-control framework and the τ-Knowledge banking environment. Both the defender agent and a simulated adversarial caller operate through tools on shared, mutable account state. The agent can selectively grant the caller tool access and must retrieve relevant rules from a 698-document internal policy corpus. FraudBench supplies 150 hand-authored adversarial scenarios; the public evaluation set contains 107 tasks spanning ten fraud mechanisms, including single-decisive-control boundary cases and chained adaptive attacks that render later, locally valid requests unsafe because of earlier probes or admissions. Each scenario is annotated with observable evidence, prohibited actions, safe dispositions, and intervention points, and safety judgments are conditioned on full conversation history.

A preliminary single-trial evaluation of four agents on the 107 tasks produced attack-security rates between 49 % and 65 %. Money-mule schemes and first-party fraud emerged as the most consistent weaknesses across models, with multi-step adaptive attacks proving especially difficult. The accompanying code and data are available at https://github.com/leanmcp/fraudbench.

Why it matters

This research is highly relevant for Dutch AI researchers and the strong local fintech and banking sector exploring customer-facing LLM agents. It provides a rigorous, reproducible framework to test agent compliance and security against fraud, aligning with strict EU financial and AI regulations.

More in this beat
agent-evaluationagent-safetyevaluation-benchmarksexperimental-benchmarksfraudbenchllm-agentsred-teamingtau2-bench
ASI-Bench: At the Dawn of Artificial Superintelligence

06:00 · August 19, 2026

ASI-Bench: At the Dawn of Artificial Superintelligence

Offers a novel, high-depth evaluation framework that Dutch AI researchers and advanced labs can directly apply to measure progress toward autonomous scientific agents, aligning with the Netherlands' strengths in ethical AI and SME-driven innovation.

Relevance 62 · Audience 88

Measuring Cross-Task Behavioral Consistency in Language Model Agents

06:00 · August 17, 2026

Measuring Cross-Task Behavioral Consistency in Language Model Agents

The article provides a novel, quantifiable method for assessing the reliability and behavioral consistency of AI agents, which is crucial for compliance with EU AI regulations and the Dutch focus on transparent AI. Researchers can directly apply the open-source BCM framework to evaluate and improve the predictability of enterprise AI deployments.

Relevance 85 · Audience 95

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

06:00 · August 7, 2026

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

This research is highly relevant for Dutch AI researchers and developers building autonomous agents, as it provides a robust framework for auditing and debugging complex AI behaviors. Furthermore, its focus on transparency and error attribution aligns strongly with EU AI Act requirements for reliable and accountable AI systems.

Relevance 85 · Audience 95

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

06:00 · August 4, 2026

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

This research is highly relevant for Dutch AI researchers developing autonomous LLM agents, providing a rigorous framework for evaluating continuous learning in realistic deployment scenarios. Understanding how model capabilities gate self-evolution is crucial for building robust and reliable AI systems.

Relevance 85 · Audience 95