AI News selected for Professionals and Decision Makers
Primary Research Stream

When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory

06:00 · August 24, 2026 · arXiv cs.AI RSS

When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory

Agentic memory under a fixed budget involves two stages: retention and retrieval. Existing retrieval-centered paradigms implicitly assume necessary evidence survives eviction, but we challenge this by isolating a pre-retrieval failure mode: structurally indirect prerequisite eviction, in which upstream blocks weakly aligned with the query are discarded under budget pressure. We provide an operational definition of this failure, a reproducible deterministic benchmark, and per-seed trace diagnostics. Finally, we evaluate Dependency-aware Semantic Garbage Collection (DSGC), a one-hop graph-aware rule. In our main suite, DSGC improves full-chain retention from 0.03 to 0.90 under a lexical encoder and from 0.23 to 1.00 under a sentence encoder. Robustness checks then identify the budget and scaling regimes where the one-hop rule holds or degrades. Our released pipeline and failure postmortem support mechanistic analysis of retention before retrieval as a distinct failure boundary.

Summary

The article examines a distinct failure mode in agentic memory systems that operate under a fixed token budget. Retention and retrieval are treated as separate stages: retention decides which context blocks survive eviction, while retrieval ranks what remains. The work isolates structurally indirect prerequisite eviction, a pre-retrieval failure in which a necessary upstream block is discarded because its surface similarity to the current query is lower than that of a downstream block that depends on it. Once evicted, the block cannot be recovered by any subsequent retrieval step.

To make the failure reproducible, the author supplies an operational definition that distinguishes it from retrieval errors, simple recency-based forgetting, and reasoning mistakes. A deterministic benchmark provides fixed seeds, explicit prerequisite edges, and chain-level targets whose joint retention is required to answer each query. Trace diagnostics record, for each failure, the evicted block, the competing block that displaced it, and the similarity margin under the chosen encoder.

The proposed mitigation is Dependency-aware Semantic Garbage Collection (DSGC), a one-hop retention rule inspired by tracing garbage collection. Query-similar blocks serve as roots; liveness then propagates along declared prerequisite edges to protect the immediate structural predecessors that would otherwise be evicted. In the main experimental suite, DSGC raised full-chain retention from 0.03 to 0.90 with a lexical encoder and from 0.23 to 1.00 with a sentence encoder. Additional robustness checks map the budget and scaling regimes in which the one-hop rule remains effective or begins to degrade.

The released pipeline and per-seed postmortem traces allow direct inspection of retention decisions before retrieval occurs, establishing a controlled base case for studying structural reachability separately from graph induction or downstream ranking.

Why it matters

This research is highly relevant for AI researchers and advanced practitioners developing long-horizon LLM agents. It provides a novel, reproducible methodology to solve memory retention failures, directly actionable for Dutch AI labs and enterprises building advanced agentic systems.

More in this beat
Function-Level Execution Feedback for Code Preference Optimization

06:00 · August 26, 2026

Function-Level Execution Feedback for Code Preference Optimization

This research provides a highly actionable and novel methodology for aligning code generation models, which is directly applicable to Dutch AI researchers and software-heavy enterprises. The open-source nature and rigorous mathematical foundation make it an excellent resource for advanced AI practitioners in the Netherlands looking to improve LLM coding capabilities.

Relevance 85 · Audience 95

Automata from Agent Traces: Failure and Next-Step Prediction

06:00 · August 26, 2026

Automata from Agent Traces: Failure and Next-Step Prediction

This research is highly relevant for Dutch AI researchers and practitioners focusing on AI safety and compliance with the EU AI Act. The proposed FSM-based monitoring offers a transparent, model-agnostic tool for auditing LLM agents and ensuring reliable deployment in enterprise environments.

Relevance 85 · Audience 95

LLM Agents Perform Controlled Experiments Using Simulation Models

06:00 · August 26, 2026

LLM Agents Perform Controlled Experiments Using Simulation Models

This research is highly relevant for Dutch AI researchers and industrial R&D teams, particularly in the strong local chemical, pharmaceutical, and high-tech manufacturing sectors. It provides a novel, actionable framework for grounding LLM reasoning in scientific simulations, addressing the critical need for reliable and evidence-based AI decision support in enterprise environments.

Relevance 85 · Audience 95

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

06:00 · August 26, 2026

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

This article presents a breakthrough in autonomous AI-driven scientific discovery using multi-agent systems. It is highly relevant for Dutch AI researchers focusing on AI for Science, multi-agent collaboration, and transparent AI methodologies, offering open-source tools and reproducible mathematical findings.

Relevance 85 · Audience 95

Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware

06:00 · August 26, 2026

Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware

This research is highly relevant for Dutch AI infrastructure researchers and HPC operators looking to optimize the serving of emerging diffusion LLMs. The findings on CPU bottlenecks and step-level parallelism provide actionable design principles for building efficient, scalable, and cost-effective AI inference systems in the Netherlands.

Relevance 85 · Audience 95

How much of a measured AI preference is the model, and how much is the instrument?

06:00 · August 26, 2026

How much of a measured AI preference is the model, and how much is the instrument?

The Netherlands strongly emphasizes ethical, transparent, and safe AI development. For Dutch researchers focusing on AI alignment and ethics, this paper provides critical methodological insights into the unreliability of current techniques used to measure AI 'preferences' or welfare.

Relevance 75 · Audience 90

AI Agents Push Humans Out of the Loop

06:00 · August 26, 2026

AI Agents Push Humans Out of the Loop

Directly addresses ethical AI deployment and human oversight mandated by the EU AI Act, relevant to Dutch enterprises and regulators prioritizing transparent, human-centered AI. Offers actionable design and organizational recommendations for Dutch AI practitioners building or deploying agents.

Relevance 68 · Audience 82