FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents
06:00 · August 6, 2026 · arXiv cs.AI RSS

Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored. We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model. On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions. Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.
Summary
FinPerMA addresses the challenge of evaluating whether LLM agents can sustain and update individualized user models across extended interactions in financial advising. Existing benchmarks tend to focus on factual recall or loosely generated dialogues, leaving event-driven shifts in preferences underexplored. The new benchmark constructs frozen longitudinal investor trajectories anchored in verifiable 2020–2026 macro, industry, and personal events drawn from behavioral-finance principles, then measures how well memory systems adapt recommendations after material changes.
Its generation pipeline relies on a deterministic three-layer Impact Model. Layer 1 produces an auditable ImpactConstraint from theory-informed rules, Layer 2 generates a narrated candidate reaction within that constraint, and Layer 3 applies automated validation and retry. Personas are sampled from calibrated demographic, financial, and psychometric distributions and mapped to Behavioral Investor Types, while timelines enforce minimum spacing and category balance. A dedicated Post-Shock checkpoint isolates whether an agent has incorporated a consequential event into its persistent model rather than relying on stale information.
Evaluation covers 2,994 questions across 276 personas and four temporal checkpoints. Seven frontier LLMs paired with up to seven memory configurations remain far from saturation: full-context baselines reach at most roughly 0.47 overall accuracy and 39 percent on multiple-choice items. Attribution analysis reveals that summary-based memory frequently retains surface facts while discarding the preference signals required for personalization, allowing simple retrieval to outperform purpose-built architectures, with the performance gap widening after shocks.
The work releases the full corpus, rule engine, model identifiers, prompts, and seeds to support reproducible comparison of recall, updating, and consolidation capabilities.
Why it matters
Provides a novel, technically rigorous benchmark for event-driven preference adaptation in LLM agents, directly usable by Dutch researchers building reliable personalized AI systems in finance and advisory domains.



