AI News selected for Professionals and Decision Makers
Primary Research Stream

Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records

06:00 · July 1, 2026 · arXiv cs.AI RSS

Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records

To ensure safe on-road behavior, pre-deployment testing and failure discovery of Autonomous Driving Systems (ADS) is crucial. Present day simulation based testing methods focus largely on mathematical models for efficient search of optimal scenarios, assuming a fixed scenario representation. On the other hand, real-world testing involves substantial manual effort to design scenario templates for testing. These templates represent distinct failure scenarios consisting of pre-deployment vehicle movements, map types, etc. Historical failure records for ADS are a reliable source of real-world failure conditions, which can be used for scenario generation. In this work, we propose a scenario generation pipeline using categorical and contextual information available from historical records in natural language format. Our approach consists of modular LLM based synthetic scenario generation, compatible with the testing constraints of a given system. We successfully apply our method to generate a diverse set of scenarios for testing autonomous navigation on Metadrive simulator using the NHTSA ADS crash records. Our approach results in accurate and diverse scenario generation with a combination of 4 road types, 3 non ego vehicle movement types, including on road anomalies in the form of working zones. Generated scenarios align with the provided testing conditions, and reveals interesting failures of the system within a limited testing budget of 20 scenarios. Code is available at https://github.com/anjaliParashar/crash2scenario.

Summary

The article introduces a pipeline that converts historical crash records into simulation-ready test scenarios for autonomous driving systems. Rather than relying on fixed mathematical templates or exhaustive manual design, the method extracts both categorical attributes—such as road type and pre-crash vehicle movements—and contextual details from natural-language narratives in existing incident reports. Large language models then translate these elements into synthetic scenarios that respect the operational constraints of a given ADS under test.

The approach is modular, allowing the generated scenarios to be adapted to different simulators and testing budgets. In the reported validation, the pipeline processed NHTSA ADS crash records to produce scenarios for the MetaDrive simulator running an IDM-based navigation policy. The resulting suite combined four road types, three categories of non-ego vehicle motion, and road anomalies such as work zones, while remaining aligned with the original crash conditions.

Within a budget of only twenty scenarios, the generated tests exposed several system failures that standard combinatorial templates would have missed. The authors release the full pipeline as open-source code, enabling reproducible application to other ADS platforms and additional crash datasets.

Why it matters

This research is highly relevant for Dutch AI researchers and automotive tech companies focusing on smart mobility and AI safety. It provides an actionable, open-source methodology using LLMs to improve the safety and regulatory compliance of autonomous systems, aligning well with EU AI safety standards.

More in this beat
autonomous-drivingevaluation-benchmarkslarge-language-modelsMetaDriveNHTSAnovel-methodologiesScenario Generation
Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

06:00 · August 7, 2026

Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

This paper is highly relevant for AI researchers in the Netherlands focusing on LLM reasoning, alignment, and compute-efficient training. The proposed weak-to-strong distillation method offers actionable insights for Dutch AI labs aiming to enhance model performance without relying solely on massive scaling.

Relevance 85 · Audience 95

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

06:00 · July 14, 2026

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

This research is highly relevant for Dutch AI researchers and enterprises developing LLMs, as it offers a mathematically rigorous method to drastically reduce the computational cost and time required for model evaluation. This aligns with the European and Dutch focus on sustainable, resource-efficient AI development (Green AI).

Relevance 85 · Audience 95

L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

06:00 · July 13, 2026

L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

This research is highly relevant for Dutch AI researchers and LegalTech developers building multi-agent systems for high-stakes, regulatory, or compliance domains. It provides actionable insights into preventing hallucination and over-deliberation, aligning with the Netherlands' strong focus on transparent, ethical, and reliable AI.

Relevance 85 · Audience 95

Synthetic Consumer Insight Generation with Large Language Models

06:00 · July 8, 2026

Synthetic Consumer Insight Generation with Large Language Models

This article is highly relevant for researchers and advanced readers in the Dutch AI market as it addresses the growing need for synthetic data generation, which is crucial for navigating strict EU GDPR privacy regulations. The methodological insights into prompt engineering and model evaluation provide valuable frameworks for Dutch AI practitioners in marketing and consumer analytics.

Relevance 85 · Audience 95

Autonomous discovery of traffic laws with AI traffic scientists

06:00 · July 3, 2026

Autonomous discovery of traffic laws with AI traffic scientists

This research is highly relevant for Dutch AI researchers and urban planners, given the Netherlands' strong focus on smart city infrastructure and advanced traffic management. The introduction of an agentic AI for autonomous scientific discovery offers actionable methodologies for institutions like TU Delft or Rijkswaterstaat to optimize urban mobility.

Relevance 85 · Audience 95

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

06:00 · July 1, 2026

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

This research provides a rigorous framework for evaluating the reasoning and reliability of LLMs in dynamic, multi-turn environments. For Dutch AI researchers and developers, understanding and benchmarking these epistemic updates is crucial for building trustworthy, transparent AI systems that align with EU standards.

Relevance 85 · Audience 95

Investigating Multi-Agent Deliberation in Law

06:00 · July 1, 2026

Investigating Multi-Agent Deliberation in Law

This research is highly relevant for Dutch AI researchers and legal tech practitioners, as it introduces novel multi-agent frameworks for legal reasoning. Given the Netherlands' strong emphasis on ethical AI and transparent legal applications, these law-inspired deliberation models offer actionable methodologies for developing robust AI systems in regulated domains.

Relevance 85 · Audience 95

NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning

06:00 · June 29, 2026

NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning

This research is highly relevant to the Dutch AI market's strong emphasis on ethical, transparent, and socially responsible AI. The benchmark provides Dutch researchers and enterprises with actionable tools to evaluate and improve the social compliance of embodied AI agents, aligning with EU regulatory frameworks for safe AI deployment.

Relevance 85 · Audience 95

An LLM-Explainable DRL Framework for Passenger-Directed Autonomous Driving

06:00 · June 23, 2026

An LLM-Explainable DRL Framework for Passenger-Directed Autonomous Driving

This research aligns with the Dutch AI market's focus on ethical, transparent AI and smart mobility. It provides researchers with a novel approach to Explainable AI (XAI) that could help autonomous systems comply with strict EU transparency regulations.

Relevance 85 · Audience 95