AI News selected for Professionals and Decision Makers
Primary Research Stream

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

06:00 · August 3, 2026 · arXiv cs.AI RSS

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.

Summary

As AI companions are deployed for repeated social interaction, users often expect stable roles and shared history to persist across sessions. Yet locally coherent replies do not guarantee that a model will retain its assigned persona or accurately recall prior commitments. The paper examines two observable long-horizon failures: persona collapse, in which a model loses its specified role, boundaries, values or style, and behavioral drift, the gradual or recurrent erosion of those properties.

To measure these failures separately, the authors introduce ANCHOR, a controlled synthetic audit framework. The study generated 2,008 conversations that combined 27 authored personas, nine interaction schedules, three memory conditions and four evaluated models. The Identity Probe pairs a sealed 102-item questionnaire with turn-level judgments of role fidelity. The Trajectory Probe scores 110 calibrated counterfactual questions drawn from 35 conversation banks, testing whether models can distinguish actual updates and user-state changes from plausible alternatives.

Results show that no model or configuration reliably preserved both dimensions. Trajectory accuracy averaged 44.4 percent, while recall of changes in user state remained near four-option chance levels. Questionnaire retention varied by model and persona facet, often diverged from observed turn-level behavior, and proved sensitive to evaluator choice. No tested context length or memory setting consistently mitigated the observed failures.

The authors conclude that current systems do not yet support reliable long-horizon companion continuity. They argue that audits should report persona enactment, trajectory recall, evaluator provenance and deployment context as distinct metrics rather than collapsing them into a single stability score.

Why it matters

Provides novel, technically deep methodology for auditing long-term AI companion continuity with strong reproducibility elements; directly supports ethical and transparent AI priorities in Dutch/EU research and SME deployment contexts.

More in this beat
agent-memoryai-alignmentai-companionsANCHORbehavioral-driftpersona-agentspersona collapse
Toward Personal Intelligence Through Cooperative Observation

06:00 · August 19, 2026

Toward Personal Intelligence Through Cooperative Observation

Strong alignment with Dutch/EU priorities on ethical, transparent, and privacy-preserving AI; offers actionable concepts for researchers building user-owned personal agents compliant with GDPR and trustworthy AI guidelines.

Relevance 78 · Audience 85

Personalization, Personas, and Forecasting in Value Alignment

06:00 · July 29, 2026

Personalization, Personas, and Forecasting in Value Alignment

The article provides critical insights into LLM cultural alignment and bias mitigation, which is highly relevant for Dutch AI researchers and enterprises striving to comply with EU ethical AI standards. Understanding how prompt framing impacts value elicitation is essential for developing transparent, localized, and culturally aware AI systems in the Netherlands.

Relevance 85 · Audience 95

Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective

06:00 · August 20, 2026

Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective

The paper provides foundational research on making autonomous AI agents auditable, safe, and transparent through dynamic graph modeling. This aligns strongly with the Dutch and EU focus on ethical AI and regulatory compliance, offering advanced researchers actionable frameworks for building governable agentic systems.

Relevance 85 · Audience 95

How Much Memory Does Your Agent Actually Need?

20:09 · August 18, 2026

How Much Memory Does Your Agent Actually Need?

This article provides highly actionable, production-focused insights for ML Engineers building AI agents. It addresses critical MLOps challenges like balancing inference cost with model accuracy through prompt caching and dynamic context retrieval, which is highly applicable for Dutch tech teams optimizing LLM deployments.

Relevance 85 · Audience 95

Position: AI Lock-In Is in Progress, and We Must Be Prepared

06:00 · August 18, 2026

Position: AI Lock-In Is in Progress, and We Must Be Prepared

The article aligns strongly with the Dutch and EU focus on responsible, ethical AI and human oversight. Its proposed frameworks for mitigating systemic AI dependency offer actionable insights for Dutch policymakers, AI safety researchers, and enterprise leaders navigating AI adoption.

Relevance 85 · Audience 90

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

06:00 · August 18, 2026

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

This article is highly relevant for Dutch AI researchers and practitioners focused on ethical AI, aligning strongly with the Netherlands' and EU's emphasis on transparent and trustworthy AI systems. It provides a critical framework for advancing LLM evaluation beyond simple value alignment toward robust normative reasoning.

Relevance 85 · Audience 95

Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents

06:00 · August 17, 2026

Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents

Agentao's focus on runtime governance, auditability, and permission-mediated execution aligns strongly with the transparency and human-oversight requirements of the EU AI Act. Dutch AI researchers and engineers can leverage this open-source architecture to build compliant, secure, and inspectable local-first AI agents.

Relevance 85 · Audience 90

AI Evaluation Should Work With Humans

06:00 · August 17, 2026

AI Evaluation Should Work With Humans

This paper aligns strongly with the Dutch and EU focus on ethical, human-centric AI and human oversight. It provides researchers with a conceptual foundation to develop new evaluation frameworks that prioritize human-AI collaboration over autonomous replacement, which is highly actionable for Dutch AI policy and enterprise deployment.

Relevance 85 · Audience 90

MobileMem: Learning from a Year of Mobile Experiences

06:00 · August 17, 2026

MobileMem: Learning from a Year of Mobile Experiences

This research is highly relevant for Dutch AI researchers and developers focusing on edge AI and personal assistants. Its emphasis on on-device, local-first memory processing aligns perfectly with the EU's strict GDPR privacy standards, offering a practical framework for building compliant, personalized AI systems.

Relevance 85 · Audience 95

MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents

06:00 · August 15, 2026

MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents

This paper provides advanced AI researchers with a rigorous framework for solving long-term memory and skill evolution in LLM agents. Its structured approach to memory consolidation and feedback aligns with the Dutch AI ecosystem's drive toward robust, transparent, and highly capable autonomous systems.

Relevance 85 · Audience 95