AI News selected for Professionals and Decision Makers
Primary Research Stream

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research

06:00 · July 7, 2026 · arXiv cs.AI RSS

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research

AI tools are accelerating scientific publication while the systems that review it struggle to keep up, and independent verification of published research has become both harder and more important. As manual replication is slow and expensive, a growing line of work uses coding agents to automate parts of the process. Existing efforts are largely packaged as benchmarks with companion agents that only run inside the benchmark's own pipeline, and no general-purpose replication tool exists. We present VERITAS, a domain-agnostic replication framework built around CLI coding agents. Given a paper, a code repository, or both, VERITAS extracts the paper's claims, runs the methodology while resolving issues as they arise, and judges each claim against the evidence from experiment runs. The pipeline returns an importance-weighted Replication Score, a severity-rated log of every fix applied, and the patched codebase. We evaluate VERITAS on CORE-Bench and ReplicationBench, 65 papers spanning computer science, social science, medicine, and astrophysics. Against two strong Claude Code baselines on the same model and host environment, VERITAS achieves state-of-the-art performance and leads on every metric on both benchmarks.

Summary

VERITAS addresses the widening gap between rapid AI-driven scientific publication and the capacity of traditional peer review to verify results. The framework provides a general-purpose pipeline that accepts a paper, an associated code repository, or both, then uses CLI coding agents to extract empirical claims, implement or adapt the described methodology, resolve runtime issues, and produce a structured replication verdict.

The system operates in six phases. An analyze step converts the input into typed claims annotated with importance levels. When no repository is supplied, a dedicated codegen phase produces an initial implementation from the paper alone. Subsequent plan, replicate, assess-fixes, and verify stages draft an execution procedure, run experiments while logging every intervention, rate the severity of those interventions, and compare observed outcomes against the extracted claims. The final output is an importance-weighted Replication Score together with a severity-rated fix log and the patched codebase.

VERITAS supports three input modes—full (paper plus repository), paper-only, and repo-only—and keeps the claim set hidden from the replication agent to reduce leakage. It has been evaluated on CORE-Bench and ReplicationBench, covering 65 papers across computer science, medicine, social science, and astrophysics. On identical model and host environments, it surpasses two strong Claude Code baselines on every reported metric, including per-capsule pass rate, per-task match rate, and adapted measures of trajectory faithfulness and unauthorized source access.

Why it matters

Directly addresses reproducibility and verification needs for Dutch AI researchers and labs; offers actionable tooling that can be adopted in NL/EU academic and industrial settings focused on trustworthy AI.

More in this beat
ai-agentsclaude-codeevaluation-benchmarksreproducibility-assetsscientific-discoveryVERITAS
FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

06:00 · July 8, 2026

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

This research is highly relevant to the Dutch AI market's focus on transparent and ethical AI. By making LLM-generated scientific hypotheses auditable and inspectable, it aligns with EU regulatory priorities and offers Dutch researchers a robust tool for accountable AI-driven scientific discovery.

Relevance 85 · Audience 95

Autonomous discovery of traffic laws with AI traffic scientists

06:00 · July 3, 2026

Autonomous discovery of traffic laws with AI traffic scientists

This research is highly relevant for Dutch AI researchers and urban planners, given the Netherlands' strong focus on smart city infrastructure and advanced traffic management. The introduction of an agentic AI for autonomous scientific discovery offers actionable methodologies for institutions like TU Delft or Rijkswaterstaat to optimize urban mobility.

Relevance 85 · Audience 95

Position: Behavioral Systems Require Behavioral Tests

06:00 · August 20, 2026

Position: Behavioral Systems Require Behavioral Tests

The article is highly relevant for Dutch AI researchers and practitioners focused on ethical and transparent AI. By proposing behavioral tests to evaluate AI alignment, safety, and decision-making processes, it provides a crucial methodological framework that supports compliance with EU regulations like the AI Act and advances responsible AI deployment.

Relevance 85 · Audience 95

ASI-Bench: At the Dawn of Artificial Superintelligence

06:00 · August 19, 2026

ASI-Bench: At the Dawn of Artificial Superintelligence

Offers a novel, high-depth evaluation framework that Dutch AI researchers and advanced labs can directly apply to measure progress toward autonomous scientific agents, aligning with the Netherlands' strengths in ethical AI and SME-driven innovation.

Relevance 62 · Audience 88

Measuring Cross-Task Behavioral Consistency in Language Model Agents

06:00 · August 17, 2026

Measuring Cross-Task Behavioral Consistency in Language Model Agents

The article provides a novel, quantifiable method for assessing the reliability and behavioral consistency of AI agents, which is crucial for compliance with EU AI regulations and the Dutch focus on transparent AI. Researchers can directly apply the open-source BCM framework to evaluate and improve the predictability of enterprise AI deployments.

Relevance 85 · Audience 95

MobileMem: Learning from a Year of Mobile Experiences

06:00 · August 17, 2026

MobileMem: Learning from a Year of Mobile Experiences

This research is highly relevant for Dutch AI researchers and developers focusing on edge AI and personal assistants. Its emphasis on on-device, local-first memory processing aligns perfectly with the EU's strict GDPR privacy standards, offering a practical framework for building compliant, personalized AI systems.

Relevance 85 · Audience 95

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

06:00 · August 15, 2026

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

This research is highly relevant for Dutch AI researchers and institutions focused on ethical AI deployment. It provides a concrete framework to evaluate and mitigate research misconduct risks when integrating LLMs into scientific workflows, aligning perfectly with the EU's emphasis on trustworthy AI.

Relevance 85 · Audience 95

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

06:00 · August 13, 2026

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

This research is highly relevant for Dutch AI researchers and MLOps practitioners developing autonomous agents, as it provides a rigorous, open-source framework for testing agent reliability and safety. Its focus on risk assessment and operational side-effects aligns strongly with the Netherlands' strategic emphasis on transparent, ethical, and secure AI deployments.

Relevance 85 · Audience 95

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

06:00 · August 3, 2026

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

This research is highly relevant for Dutch AI researchers and developers building autonomous agents, as it provides a structured methodology for diagnosing and repairing complex AI systems. It aligns well with the EU's focus on AI robustness, transparency, and safety by offering a standardized way to trace and mitigate agent failures.

Relevance 85 · Audience 95