MatrAIx: Simulating the World with 8.3 Billion Persona Agents
06:00 · August 6, 2026 · arXiv cs.AI RSS

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Summary
MatrAIx addresses the persistent gap between scalable but abstract offline benchmarks and expensive, slow human studies by offering a population-scale simulated-user evaluation framework. The system models heterogeneous users through Persona 8B, a collection of 8.3 billion persona records defined across 1,290 categorical dimensions that cover background, psychology, capability, behavior, and lifestyle. Records are generated either by sampling from a dependency graph that maintains realistic attribute correlations and compatibility constraints, or by extracting structured profiles from human-authored sources such as Wikipedia biographies, review histories, and survey data. A quality-filtered coreset of roughly one million personas, split between 599,847 human-grounded and 400,000 synthetic entries, has been released for research use.
The second component is the MatrAIx Playground, which supplies four standardized interaction environments—Survey, AI Chatbot, Web, and App—in which persona agents can evaluate digital products. These environments are paired with 1,010 application tasks distributed across more than 25 domains, including commerce, software, finance, and healthcare. In a series of 18,189 evaluation trials conducted on eight representative tasks, agents powered by Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5 produced feedback that reflected systematic differences in user behavior, such as price sensitivity, tolerance for AI failures, and latency expectations.
Validation experiments support the reliability of the approach. A controlled study of 400 trials measured adherence to ten behavioral attributes across all four environments and found that declared persona traits were correctly expressed or suppressed in 91.5 percent of cases. Separate human and LLM judging assessed the fidelity of attribute extraction from source material. By combining large-scale synthetic populations with grounded profiles and controlled interaction settings, MatrAIx enables repeated, subgroup-aware testing that captures variation across user cohorts while remaining far less resource-intensive than traditional human-subject studies.
Why it matters
Directly actionable for Dutch AI researchers and advanced practitioners developing ethical, transparent evaluation methods aligned with EU priorities; high technical depth, reproducibility via open code/dataset, and novelty in large-scale persona simulation support SME adoption and regulatory-compliant testing in the Netherlands.










