AI News selected for Professionals and Decision Makers
Primary Research Stream

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

06:00 · August 6, 2026 · arXiv cs.AI RSS

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored. We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model. On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions. Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.

Summary

FinPerMA addresses the challenge of evaluating whether LLM agents can sustain and update individualized user models across extended interactions in financial advising. Existing benchmarks tend to focus on factual recall or loosely generated dialogues, leaving event-driven shifts in preferences underexplored. The new benchmark constructs frozen longitudinal investor trajectories anchored in verifiable 2020–2026 macro, industry, and personal events drawn from behavioral-finance principles, then measures how well memory systems adapt recommendations after material changes.

Its generation pipeline relies on a deterministic three-layer Impact Model. Layer 1 produces an auditable ImpactConstraint from theory-informed rules, Layer 2 generates a narrated candidate reaction within that constraint, and Layer 3 applies automated validation and retry. Personas are sampled from calibrated demographic, financial, and psychometric distributions and mapped to Behavioral Investor Types, while timelines enforce minimum spacing and category balance. A dedicated Post-Shock checkpoint isolates whether an agent has incorporated a consequential event into its persistent model rather than relying on stale information.

Evaluation covers 2,994 questions across 276 personas and four temporal checkpoints. Seven frontier LLMs paired with up to seven memory configurations remain far from saturation: full-context baselines reach at most roughly 0.47 overall accuracy and 39 percent on multiple-choice items. Attribution analysis reveals that summary-based memory frequently retains surface facts while discarding the preference signals required for personalization, allowing simple retrieval to outperform purpose-built architectures, with the performance gap widening after shocks.

The work releases the full corpus, rule engine, model identifiers, prompts, and seeds to support reproducible comparison of recall, updating, and consolidation capabilities.

Why it matters

Provides a novel, technically rigorous benchmark for event-driven preference adaptation in LLM agents, directly usable by Dutch researchers building reliable personalized AI systems in finance and advisory domains.

More in this beat
agent-memoryevaluation-benchmarksfinancial-advisingFinPerMAlarge-language-modelsllm-agentspersonalized-memory
Object-Centric Environment Modeling for Agentic Tasks

06:00 · July 7, 2026

Object-Centric Environment Modeling for Agentic Tasks

This research is highly relevant for Dutch AI researchers and developers working on autonomous LLM agents. It provides a structured, programmatic approach to agent memory and environment modeling, which can be directly applied by technical teams in the Netherlands to build more robust and reliable AI systems.

Relevance 75 · Audience 90

MobileMem: Learning from a Year of Mobile Experiences

06:00 · August 17, 2026

MobileMem: Learning from a Year of Mobile Experiences

This research is highly relevant for Dutch AI researchers and developers focusing on edge AI and personal assistants. Its emphasis on on-device, local-first memory processing aligns perfectly with the EU's strict GDPR privacy standards, offering a practical framework for building compliant, personalized AI systems.

Relevance 85 · Audience 95

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

06:00 · August 3, 2026

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

The research is highly relevant for Dutch AI practitioners as it provides a reproducible, privacy-preserving framework using local inference that aligns with strict EU data sovereignty and governance standards. It offers actionable architectural blueprints for researchers building trustworthy, scalable autonomous agents.

Relevance 85 · Audience 95

Accurate and Efficient Long-Term Memory for LLM Agents

06:00 · July 21, 2026

Accurate and Efficient Long-Term Memory for LLM Agents

Provides novel, reproducible graph-based memory methods directly applicable to reliable LLM agent development; aligns with Dutch/EU emphasis on ethical, transparent AI and supports SME adoption of robust agent systems.

Relevance 75 · Audience 90

L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

06:00 · July 13, 2026

L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

This research is highly relevant for Dutch AI researchers and LegalTech developers building multi-agent systems for high-stakes, regulatory, or compliance domains. It provides actionable insights into preventing hallucination and over-deliberation, aligning with the Netherlands' strong focus on transparent, ethical, and reliable AI.

Relevance 85 · Audience 95

APeB: Benchmarking Personalization Ability of Large Language Model Agents

06:00 · July 7, 2026

APeB: Benchmarking Personalization Ability of Large Language Model Agents

This research provides a valuable benchmark and methodology for Dutch AI researchers and enterprises, particularly in e-commerce and customer service, developing personalized LLM agents. Improving intent discovery from user histories directly impacts the effectiveness of AI-driven consumer applications prevalent in the Netherlands.

Relevance 75 · Audience 90

AGI Maze as a Benchmark Framework for World-Modeling Agents

06:00 · July 2, 2026

AGI Maze as a Benchmark Framework for World-Modeling Agents

This research is highly relevant for Dutch AI researchers and developers focusing on autonomous agents and LLM reasoning capabilities. It provides a novel benchmarking tool to test and improve the robustness and world-modeling skills of AI systems, aligning with the Netherlands' strong academic focus on advanced, reliable AI.

Relevance 75 · Audience 90

Darwin Mobile Agent: A Roadmap for Self-Evolution

06:00 · June 23, 2026

Darwin Mobile Agent: A Roadmap for Self-Evolution

This research provides a novel, open-source infrastructure for developing autonomous, self-evolving GUI agents, which is highly actionable for Dutch AI researchers and developers working on reinforcement learning and automation. The focus on removing human priors aligns with advanced AI development goals within the Netherlands' strong technical ecosystem.

Relevance 75 · Audience 95