AI News selected for Professionals and Decision Makers
Primary Research Stream

Agentic Neural Architecture Search

06:00 · July 11, 2026 · arXiv cs.AI RSS

Agentic Neural Architecture Search

Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces that require substantial domain expertise and must be rebuilt for every new task. Large language models (LLMs) can generate architectures in an open-ended space, but how to optimally divide the labor between LLM-driven design and NAS-driven search remains unexplored. We propose a mechanism that bridges these two paradigms: an LLM produces a high-quality seed architecture, then decomposes it into a "slotted architecture", a scaffold with named, interchangeable module slots that automatically defines a bounded, task-specific search space for conventional NAS to explore, without manual engineering. We instantiate this mechanism in AgentNAS, a modular three-phase pipeline in which each component's contribution can be measured independently. On 17 tasks spanning classification, dense regression, segmentation, and multi-label tagging across diverse modalities (NAS-Bench-360 and Unseen NAS), AgentNAS establishes a new state of the art on 11 tasks, outperforming published baselines including task-specific expert designs. Ablation studies show that the two search mechanisms are broadly complementary: the LLM-generated seed already surpasses published baselines on the majority of tasks, and NAS delivers additional gains in most cases through combinatorial recombination across slots, a mode of search that independent LLM samples cannot replicate. These patterns hold across three LLMs of different capability levels, confirming that the division of labor is robust. Our code is available at https://github.com/alroimfebruary/AgentNAS.

Summary

Neural architecture search has long been constrained by the need for manually designed search spaces that embed domain expertise and must be rebuilt for each new task. AgentNAS addresses this limitation through a hybrid pipeline that lets a large language model first propose a seed architecture and then decompose it into a slotted scaffold. The scaffold consists of named, interchangeable module slots whose alternatives are automatically enumerated, thereby creating a bounded yet task-specific search space that conventional NAS algorithms can explore without further human engineering.

The resulting three-phase AgentNAS system separates the contributions of the LLM and the NAS stage so each can be measured independently. In the first phase the LLM iteratively proposes, implements, and evaluates candidate networks until performance saturates. The second phase converts the best seed into the slotted form, exposing combinatorial degrees of freedom at the module level while preserving the LLM’s macro-level decisions on depth, width progression, and backbone type. A standard NAS procedure then recombines the slot alternatives in the third phase.

Evaluated across the 17 tasks of NAS-Bench-360 and Unseen NAS—covering image classification, dense regression, segmentation, and multi-label tagging in multiple modalities—AgentNAS reaches state-of-the-art accuracy on 11 tasks and surpasses both published baselines and task-specific expert designs. Ablation experiments conducted with three LLMs of varying capability show that the LLM-generated seed already exceeds most baselines on the majority of tasks, while the subsequent NAS recombination step supplies additional gains that independent LLM sampling at matched compute budgets cannot replicate. The observed complementarity between the two mechanisms remains consistent across model scales.

Why it matters

High technical depth, novelty in bridging open-ended LLM generation with combinatorial NAS, full reproducibility via public code, and direct applicability for Dutch researchers advancing AutoML and agentic systems.

More in this beat
AgentNASexperimental-benchmarkslarge-language-modelsmodel-architectureNAS-Bench-360novel-methodologiesUnseen NAS
Agentic evolution of physically constrained foundation models

06:00 · June 25, 2026

Agentic evolution of physically constrained foundation models

This research is highly relevant for Dutch AI researchers and infrastructure engineers focusing on efficient, sustainable AI deployment. By drastically reducing the hardware requirements for massive foundation models, it enables local, cost-effective deployment for SMEs and aligns with European goals for green AI and data sovereignty.

Relevance 85 · Audience 95

Latent Goal Prediction from Language for Model-Based Planning

06:00 · June 23, 2026

Latent Goal Prediction from Language for Model-Based Planning

This research is highly relevant for AI researchers and practitioners in the Netherlands, particularly those focused on robotics, autonomous systems, and logistics. The LAGO framework offers actionable methodologies for improving long-horizon planning and text-guided control, aligning well with the Dutch high-tech sector's focus on advanced automation.

Relevance 85 · Audience 95

Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

06:00 · August 7, 2026

Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

This paper is highly relevant for AI researchers in the Netherlands focusing on LLM reasoning, alignment, and compute-efficient training. The proposed weak-to-strong distillation method offers actionable insights for Dutch AI labs aiming to enhance model performance without relying solely on massive scaling.

Relevance 85 · Audience 95

MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers

06:00 · July 22, 2026

MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers

This research is highly relevant for Dutch AI and Operations Research practitioners, particularly in the logistics, manufacturing, and supply chain sectors where MILP solvers are foundational. The focus on generating explicit, interpretable ('white-box') solver logic aligns perfectly with the Netherlands' and EU's strategic emphasis on transparent and trustworthy AI.

Relevance 85 · Audience 95

Some Large Language Models Exhibit Consistent Risk Attitudes

06:00 · July 21, 2026

Some Large Language Models Exhibit Consistent Risk Attitudes

This research is highly relevant for Dutch AI researchers and policymakers focused on ethical and transparent AI, as it provides a novel framework for auditing the intrinsic risk behaviors of LLMs. Understanding these latent risk profiles is crucial for deploying AI in high-stakes environments and aligns perfectly with the EU's stringent risk management requirements.

Relevance 85 · Audience 95

How Far Can Root Cause Analysis Go on Real-World Telemetry Data?

06:00 · July 16, 2026

How Far Can Root Cause Analysis Go on Real-World Telemetry Data?

This research is highly relevant for AI researchers and AIOps practitioners in the Netherlands managing complex cloud-native environments. It provides actionable insights into improving LLM-based multi-agent systems for automated diagnostics, a critical area for Dutch tech enterprises and infrastructure providers.

Relevance 85 · Audience 95

Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems

06:00 · July 14, 2026

Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems

This research provides a rigorous mathematical foundation for building robust, distributed multi-agent systems, directly addressing the reliability and traceability requirements crucial for enterprise AI deployment. Its focus on verifiable semantic rollbacks and transparent belief lineages aligns strongly with the EU's regulatory emphasis on AI safety and oversight, making it highly valuable for Dutch AI researchers and infrastructure developers.

Relevance 85 · Audience 95

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

06:00 · July 14, 2026

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

This research is highly relevant for Dutch AI researchers and enterprises developing LLMs, as it offers a mathematically rigorous method to drastically reduce the computational cost and time required for model evaluation. This aligns with the European and Dutch focus on sustainable, resource-efficient AI development (Green AI).

Relevance 85 · Audience 95

Interpreting Latent CoT Reasoning as Dynamical Systems

06:00 · July 14, 2026

Interpreting Latent CoT Reasoning as Dynamical Systems

The article is highly relevant for AI researchers in the Netherlands focusing on LLM interpretability and trustworthy AI. Understanding the internal dynamics of latent reasoning aligns strongly with EU and Dutch priorities for transparent and explainable AI systems.

Relevance 85 · Audience 95