Memory in the Loop: In-Process Retrieval as ExtendedWorking Memory for Language Agents
06:00 · July 8, 2026 · arXiv cs.AI RSS

Language agents run a loop - observe, reason, act - but the memory they reason over sits outside it: a store queried at most once per turn. We study the regime where memory moves inside the loop, read and written on every step. The obstacle has always been latency: networked stores answer in tens to hundreds of milliseconds, and in-loop retrieval can inflate end-to-end latency by up to 83x when retrieval is expensive. Prior work manages that cost rather than questioning it: serving-layer scheduling hides it, "memory-first" designs ration retrieval to once per turn. We argue latency is a property of where the store lives, not the in-loop pattern: an in-process store answers in ~100us, three orders of magnitude below the network regime, and at that speed the per-step tax collapses. By the extended-mind thesis's parity principle, a store fast enough to be constantly and directly available becomes extended working memory, not a tool the agent merely consults. The premise is causal: holding a fixed per-turn memory-latency budget and varying only the store's answer speed, redundant actions rise monotonically with latency - 0.0 of 12 at in-process speed, 7.2 of 12 at a 110ms cloud round trip (gpt-5-nano, gpt-5-mini; exact permutation p=0.0079). We demonstrate the regime end-to-end: across four GPT-5-class models under a bounded window, recall improves from 0/5 to 3.6-4.8/5 with in-loop memory, store ops at p50 80-165us - though an instructed restate-every-reply baseline also solves it perfectly, at a token cost that grows with the working set. The store never lost a fact in any run (244 of 244 writes kept); every miss traces to the agent's read policy, not the store. Our measurements also relocate the bottleneck: the dominant per-step cost is embedding (~200-400ms over the network); pairing the in-process store with a small local embedder returns the complete operation to a measured ~40us.
Summary
Language agents operate in an observe-reason-act cycle, yet conventional designs keep external memory outside that loop and query it at most once per turn. Networked vector stores impose latencies of 50–200 ms, which quickly renders repeated retrievals prohibitive and forces designers either to hide the cost through scheduling or to restrict access to the start of each turn. The paper demonstrates that relocating the store to an in-process implementation changes the economics: answer times fall to roughly 100 µs, three orders of magnitude below the network regime, so that the cumulative per-turn overhead collapses to about 1.7 ms even when memory is consulted on every reasoning step.
At this speed the external store satisfies the latency criteria of the extended-mind thesis and functions as genuine extended working memory rather than a consulted tool. Controlled experiments that hold the per-turn memory budget fixed while varying only store latency confirm the causal link: redundant actions rise monotonically from zero at in-process speeds to 7.2 out of 12 at a 110 ms round-trip, with the difference statistically reliable. End-to-end trials across four GPT-5-class models under bounded context windows show recall improving from zero to between 3.6 and 4.8 out of five facts, while the store itself never dropped a write across 244 operations; every observed miss originated in the agent’s read policy.
The measurements also relocate the remaining bottleneck. Embedding over the network now dominates at 200–400 ms per step. Pairing the in-process store with a compact local embedder reduces the complete operation to approximately 40 µs, restoring the regime in which memory participates directly in every reasoning cycle without token-cost inflation or external scheduling.
Why it matters
This research is highly relevant for Dutch AI researchers and engineers developing autonomous language agents, offering a practical architectural shift to drastically reduce latency and improve agent reasoning. It provides deep technical insights into optimizing memory loops, which is crucial for building efficient, scalable AI software in the Netherlands.





