AI News selected for Professionals and Decision Makers
Primary Research Stream

RIZZ: Routing Interactions to Near Zero-Interference Zones for Continual Adaptation of Black-Box Agents

06:00 · June 23, 2026 · arXiv cs.AI RSS

RIZZ: Routing Interactions to Near Zero-Interference Zones for Continual Adaptation of Black-Box Agents

Large language models are increasingly deployed as long-lived agents that must adapt across users, tasks, domains, modalities, and feedback regimes without access to model weights. Existing black-box adaptation methods typically optimize a single prompt, maintain an undifferentiated memory, or rely on repeated rollout-heavy search. However, these designs struggle when streams of input are nonstationary, feedback is sparse, and failures from one task family can contaminate behavior on another. We introduce RIZZ (Routing Interactions to Near Zero-interference Zones), a continual adaptation framework for compound language-model systems that learns entirely through verifier-gated memory, routing, and prompt compilation. RIZZ organizes input streams into dynamically spawned memory branches. At inference time, either while online or offline, a context-aware router selects or creates a branch that retrieves branch-local, global, graph-structured, and working-memory context, which is compiled into a bounded prompt together with retrieved task evidence. After the model acts, task verifiers score the output, and only verified interactions can update memory, promote reusable rules, demote harmful rules, or create anti-patterns. This yields a black-box agent that improves through persistent natural-language feedback while explicitly controlling interference. RIZZ targets the regime where adaptation must occur online under context budgets. Finally, we demonstrate the effectiveness of our framework against state-of-the-art baselines on competitive benchmarks.

Summary

RIZZ addresses the challenge of adapting black-box large language models that operate as persistent agents across shifting users, tasks, and domains. Because model weights remain inaccessible, adaptation must occur through external mechanisms rather than parameter updates. Standard approaches that rely on a single optimized prompt or a shared undifferentiated memory often fail when input streams are nonstationary and feedback arrives only after actions are taken, allowing experience from one task family to degrade performance on others.

The framework organizes incoming queries into dynamically created memory branches, each functioning as a localized store of verified examples, distilled procedural rules, and reward statistics. A context-aware router, operating at inference time, either selects an existing branch or spawns a new one by combining output-shape signals, embedding similarity, and label information. Relevant branch-local, global, graph-structured, and working-memory elements are then assembled with task evidence into a bounded prompt that respects context limits.

After the model produces a response, deterministic task verifiers evaluate the outcome. Only interactions that meet verification thresholds are permitted to update memory: successful traces reinforce reusable rules, while failures are retained as anti-patterns or demoted. This gated process enables the system to discover its own task organization without requiring ground-truth labels or a predefined taxonomy, while preventing unchecked growth in token usage.

Evaluations across streaming, sequential, conversational, and tool-use benchmarks show that the routed, verifier-controlled design improves accuracy over frozen baselines at modest additional cost and with substantially lower overhead than heavier memory-intensive alternatives.

Why it matters

This research is highly relevant for Dutch AI researchers and developers building compound LLM systems, as it offers a novel method for continuous agent adaptation using black-box models like commercial APIs. Its focus on controlled, verifier-gated memory aligns well with the EU's push for reliable, transparent, and predictable AI behavior.

More in this beat
agent-memorycontext-managementcontinual-learninglarge-language-modelsllm-agentsRIZZtool-useworking-memory
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

06:00 · August 3, 2026

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

The research is highly relevant for Dutch AI practitioners as it provides a reproducible, privacy-preserving framework using local inference that aligns with strict EU data sovereignty and governance standards. It offers actionable architectural blueprints for researchers building trustworthy, scalable autonomous agents.

Relevance 85 · Audience 95

Accurate and Efficient Long-Term Memory for LLM Agents

06:00 · July 21, 2026

Accurate and Efficient Long-Term Memory for LLM Agents

Provides novel, reproducible graph-based memory methods directly applicable to reliable LLM agent development; aligns with Dutch/EU emphasis on ethical, transparent AI and supports SME adoption of robust agent systems.

Relevance 75 · Audience 90

Cura 1T: Specialized Model for Agentic Healthcare

06:00 · July 20, 2026

Cura 1T: Specialized Model for Agentic Healthcare

This research is highly relevant for Dutch AI researchers and healthcare institutions developing specialized clinical models. The data-centric, self-evolving training methodology offers a transparent and rigorous approach to building reliable healthcare AI, aligning with EU regulatory standards for clinical deployment.

Relevance 85 · Audience 95

CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions

06:00 · July 13, 2026

CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions

This research is highly relevant for Dutch AI researchers and engineers building enterprise LLM systems, as it offers a concrete methodology to improve AI reliability and predictability. This aligns strongly with the Netherlands' and EU's regulatory focus on transparent, trustworthy, and controllable AI systems without requiring massive computational resources for model scaling.

Relevance 85 · Audience 95

StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems

06:00 · July 8, 2026

StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems

This research is highly relevant for Dutch AI practitioners developing multi-agent systems, as it directly addresses the need for transparent and auditable AI memory architectures. By preserving data conflicts rather than overwriting them, StateFuse aligns strongly with EU and Dutch priorities for ethical, explainable, and safe AI deployments.

Relevance 85 · Audience 95

Object-Centric Environment Modeling for Agentic Tasks

06:00 · July 7, 2026

Object-Centric Environment Modeling for Agentic Tasks

This research is highly relevant for Dutch AI researchers and developers working on autonomous LLM agents. It provides a structured, programmatic approach to agent memory and environment modeling, which can be directly applied by technical teams in the Netherlands to build more robust and reliable AI systems.

Relevance 75 · Audience 90

AGI Maze as a Benchmark Framework for World-Modeling Agents

06:00 · July 2, 2026

AGI Maze as a Benchmark Framework for World-Modeling Agents

This research is highly relevant for Dutch AI researchers and developers focusing on autonomous agents and LLM reasoning capabilities. It provides a novel benchmarking tool to test and improve the robustness and world-modeling skills of AI systems, aligning with the Netherlands' strong academic focus on advanced, reliable AI.

Relevance 75 · Audience 90

Darwin Mobile Agent: A Roadmap for Self-Evolution

06:00 · June 23, 2026

Darwin Mobile Agent: A Roadmap for Self-Evolution

This research provides a novel, open-source infrastructure for developing autonomous, self-evolving GUI agents, which is highly actionable for Dutch AI researchers and developers working on reinforcement learning and automation. The focus on removing human priors aligns with advanced AI development goals within the Netherlands' strong technical ecosystem.

Relevance 75 · Audience 95