Agent-Native Immune System: Architecture, Taxonomy, and Engineering
06:00 · June 29, 2026 · arXiv cs.AI RSS

The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has fundamentally expanded the AI threat landscape. Current defense mechanisms, such as perimeter security and training-time alignment, remain external to the agent's active reasoning loop. Consequently, they fall short: a fully aligned agent remains highly vulnerable to runtime hijacking via memory poisoning, tool-chain manipulation, or multi-agent protocol attacks. To address this critical gap, we introduce the Agent-Native Immune System (ANIS), the first biologically inspired, endogenous defense architecture embedded directly within the agent's cognitive loop. Our framework presents four primary contributions. First, we design a six-layer Immune Tower (L0-L5), distinctly incorporating Barrier Immunity (L1) as a non-cognitive, physical-and-logical isolation layer. Second, we establish a unified taxonomy of Agent Viruses and Agent Vaccines, formalizing the critical distinction between superficial non-parametric defenses and robust parametric vaccines. Third, we conceptualize the Harness Triad--Meta, Self, and Auto--a self-monitoring, meta-cognitive automation backbone that drives Continual Immune Learning (CIL), enabling vaccines to dynamically adapt to novel threats. Finally, we establish a rigorous theoretical demarcation between model alignment and agent immunity: while alignment provides a static "constitutional" value foundation during training, ANIS serves as the dynamic "law enforcement" mechanism during runtime. We conclude by framing open challenges for the field, including immune protocol standardization, novel evaluation metrics such as the Autoimmunity Rate (false-positive intervention rate), and the co-evolutionary dynamics between pathogens and vaccines within collective intelligence ecosystems.
Summary
The shift from static chat models to autonomous agents equipped with persistent memory, tool-use protocols, and multi-agent coordination has enlarged the attack surface beyond what perimeter controls or training-time alignment can address. A model that has internalized human values during pre-training can still be subverted at runtime through memory poisoning, adversarial tool metadata, or protocol-level manipulation among collaborating agents. Existing safeguards operate outside the agent’s active reasoning loop and therefore cannot respond to threats that appear only after deployment.
The Agent-Native Immune System (ANIS) places defense inside that loop. It draws on biological immunity to define a six-layer Immune Tower (L0–L5) whose lowest non-cognitive tier, Barrier Immunity (L1), enforces physical and logical isolation before any reasoning occurs. Higher layers integrate innate detection, adaptive response, and memory of past encounters. A companion taxonomy distinguishes Agent Viruses—malicious inputs or protocol exploits—from Agent Vaccines, separating superficial non-parametric measures such as prompt filters from parametric interventions that modify internal representations, including steering vectors and defensive LoRA adapters.
Central to sustained protection is the Harness Triad of meta-level search, self-monitoring, and automatic synthesis. Together these mechanisms drive Continual Immune Learning, allowing vaccines to be generated, tested, and refined without external supervision. The framework explicitly separates this runtime “law enforcement” function from the static “constitutional” role of model alignment: alignment supplies enduring value constraints at training time, while ANIS supplies ongoing detection and correction once the agent is operating.
The authors close by identifying open engineering questions, among them the standardization of immune protocols across agent platforms, the definition of an Autoimmunity Rate to quantify false-positive interventions, and the co-evolutionary pressures that arise when multiple agents exchange both pathogens and defensive updates within shared ecosystems.
Why it matters
This research aligns perfectly with the Dutch AI market's strategic focus on secure, ethical, and transparent AI. It provides advanced researchers with a novel, dynamic runtime defense framework necessary for deploying safe autonomous agents within strict EU regulatory environments.






