Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
06:00 · August 3, 2026 · arXiv cs.AI RSS

As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.
Summary
As AI companions are deployed for repeated social interaction, users often expect stable roles and shared history to persist across sessions. Yet locally coherent replies do not guarantee that a model will retain its assigned persona or accurately recall prior commitments. The paper examines two observable long-horizon failures: persona collapse, in which a model loses its specified role, boundaries, values or style, and behavioral drift, the gradual or recurrent erosion of those properties.
To measure these failures separately, the authors introduce ANCHOR, a controlled synthetic audit framework. The study generated 2,008 conversations that combined 27 authored personas, nine interaction schedules, three memory conditions and four evaluated models. The Identity Probe pairs a sealed 102-item questionnaire with turn-level judgments of role fidelity. The Trajectory Probe scores 110 calibrated counterfactual questions drawn from 35 conversation banks, testing whether models can distinguish actual updates and user-state changes from plausible alternatives.
Results show that no model or configuration reliably preserved both dimensions. Trajectory accuracy averaged 44.4 percent, while recall of changes in user state remained near four-option chance levels. Questionnaire retention varied by model and persona facet, often diverged from observed turn-level behavior, and proved sensitive to evaluator choice. No tested context length or memory setting consistently mitigated the observed failures.
The authors conclude that current systems do not yet support reliable long-horizon companion continuity. They argue that audits should report persona enactment, trajectory recall, evaluator provenance and deployment context as distinct metrics rather than collapsing them into a single stability score.
Why it matters
Provides novel, technically deep methodology for auditing long-term AI companion continuity with strong reproducibility elements; directly supports ethical and transparent AI priorities in Dutch/EU research and SME deployment contexts.










