AI News selected for Professionals and Decision Makers
Primary Research Stream

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

06:00 · July 8, 2026 · arXiv cs.AI RSS

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect. We introduce FirstResearch, a first-principles research-question formation framework for scientific LLM agents whose core artifact is a structured Research Question Certificate. The certificate records primitive definitions, assumptions, a mechanism model, a tension or contradiction, a falsifiable hypothesis, a minimal decisive test, and a failure update rule, making the proposed question inspectable before downstream execution. On ten LLM-agent research topics, FirstResearch outperforms controlled prompt-level baselines inspired by AI co-scientist, Agent Laboratory, and AI Scientist-v2 under a primary DeepSeek-blind-judge protocol. A Gemini-2.5-Flash independent-judge rescore of the same 40 baseline packages preserves the system-level ranking, with FirstResearch scoring 4.86/5 versus 4.38/5 for the strongest baseline and Pearson agreement of 0.865 on average score. A one-repeat ablation checkpoint further suggests that the certificate-centered core is the strongest component: certificate-only scoring reaches 4.90/5 under DeepSeek and 4.88/5 under Gemini, while removing certificates drops below 1/5 under both judges. These results are preliminary and use LLM judges rather than human domain experts, but they support a narrow scientific-discovery claim: explicit derivation constraints are a promising mechanism for making LLM-generated scientific questions more auditable. Code, prompts, saved outputs, and reproduction scripts are available at https://github.com/louiswang524/FirstResearch.

Summary

FirstResearch addresses a persistent gap in LLM-driven scientific discovery: while current agents can generate plausible research questions, those outputs often conceal the underlying assumptions, mechanisms, and falsifying conditions that human scientists need to inspect. The framework enforces a structured derivation process that begins with explicit primitive definitions and first-principles assumptions, then builds a mechanism model, surfaces tensions or contradictions, and produces a Research Question Certificate for each candidate. This certificate records a falsifiable hypothesis, a minimal decisive test, expected observations, and a failure-update rule, allowing the question to be audited before any downstream experiment design or execution begins.

A novelty-aware gate further refines weak but formally valid certificates by requiring mechanism-boundary signals such as thresholds, interactions, or failure regimes. On a benchmark of ten LLM-agent research topics, the system was compared against controlled prompt-level baselines modeled on AI co-scientist, Agent Laboratory, and AI Scientist-v2. Under a primary DeepSeek judge, FirstResearch produced higher average rubric scores; an independent Gemini-2.5-Flash rescore preserved the ranking, with FirstResearch at 4.86/5 versus 4.38/5 for the strongest baseline and strong inter-judge agreement. An ablation isolating the certificate component confirmed its central contribution, while removal of the certificate caused scores to collapse.

The authors emphasize that the results remain preliminary because they rely on LLM judges rather than domain-expert human review, and they position FirstResearch as a narrow, complementary layer rather than a full autonomous research pipeline. All code, prompts, and reproduction scripts are released publicly to support further validation of explicit derivation constraints as a route to more auditable LLM-generated scientific questions.

Why it matters

This research is highly relevant to the Dutch AI market's focus on transparent and ethical AI. By making LLM-generated scientific hypotheses auditable and inspectable, it aligns with EU regulatory priorities and offers Dutch researchers a robust tool for accountable AI-driven scientific discovery.

More in this beat
ai-agentsevaluation-benchmarksFirstResearchnovel-methodologiesResearch Impactscientific-discoverytrustworthy-ai-practices
Autonomous discovery of traffic laws with AI traffic scientists

06:00 · July 3, 2026

Autonomous discovery of traffic laws with AI traffic scientists

This research is highly relevant for Dutch AI researchers and urban planners, given the Netherlands' strong focus on smart city infrastructure and advanced traffic management. The introduction of an agentic AI for autonomous scientific discovery offers actionable methodologies for institutions like TU Delft or Rijkswaterstaat to optimize urban mobility.

Relevance 85 · Audience 95

Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents

06:00 · July 13, 2026

Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents

This research is highly relevant to the Dutch AI market's strong emphasis on transparent, ethical, and auditable AI systems. It provides researchers with a concrete methodology to build explainable AI scientists, aligning with EU regulatory standards for AI traceability and accountability.

Relevance 85 · Audience 95

ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation

06:00 · July 8, 2026

ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation

This research is highly relevant for the Dutch AI market, particularly for its strong high-tech manufacturing and engineering sectors (e.g., ASML, VDL, Philips). Researchers and advanced practitioners can leverage these text-to-CAD advancements to automate and optimize complex industrial design workflows in the Netherlands.

Relevance 85 · Audience 95

Auto-FL-Research: Agentic Search for Federated Learning Algorithms

06:00 · July 3, 2026

Auto-FL-Research: Agentic Search for Federated Learning Algorithms

Federated Learning is crucial for the Dutch AI market due to strict EU data privacy regulations (GDPR), especially in collaborative sectors like healthcare. This research provides advanced practitioners with an automated, agent-driven approach to optimize FL pipelines, directly supporting scalable and privacy-preserving AI development in the Netherlands.

Relevance 85 · Audience 95

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

06:00 · July 2, 2026

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

This research directly supports the Dutch and EU focus on ethical, transparent, and trustworthy AI by providing a rigorous method to audit LLM reasoning. It offers researchers and advanced practitioners a novel framework to mitigate hallucinations and ensure compliance with emerging AI regulations.

Relevance 85 · Audience 95

Self-Evolving Agents with Anytime-Valid Certificates

06:00 · July 2, 2026

Self-Evolving Agents with Anytime-Valid Certificates

This research is highly relevant for Dutch AI researchers and practitioners because it addresses the critical need for auditable and safe autonomous agents, aligning perfectly with the EU AI Act's emphasis on transparency and risk management. The introduction of anytime-valid certificates provides a mathematically grounded approach to deploying self-evolving AI in enterprise environments.

Relevance 85 · Audience 95

Optimal Resource Utilization for Autonomous Laboratory Orchestrators

06:00 · July 2, 2026

Optimal Resource Utilization for Autonomous Laboratory Orchestrators

The research is highly relevant for Dutch R&D sectors, particularly in materials science, chemistry, and high-tech manufacturing, where autonomous laboratories can significantly accelerate innovation. It provides actionable methodologies for AI researchers and engineers looking to optimize hardware orchestration and resource management in automated experimental setups.

Relevance 75 · Audience 90

Beyond expert users: agents should help users construct preferences, not just elicit them

06:00 · July 1, 2026

Beyond expert users: agents should help users construct preferences, not just elicit them

This research is highly relevant for AI researchers and developers focusing on user-centric and transparent AI, a key priority in the Dutch AI market. By providing a formal framework and benchmark for improving how agents assist non-expert users, it offers actionable insights for enhancing conversational AI and recommender systems in enterprise applications.

Relevance 85 · Audience 95

Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering

06:00 · July 1, 2026

Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering

This research is highly relevant for Dutch AI researchers and practitioners as it offers a concrete methodology to reduce compute costs and improve the efficiency of AI development through transfer learning in multi-agent systems. Its focus on resource efficiency aligns well with the Dutch AI market's emphasis on sustainable and scalable AI solutions for enterprises and SMEs.

Relevance 85 · Audience 95