Synthetic Consumer Insight Generation with Large Language Models
06:00 · July 8, 2026 · arXiv cs.AI RSS

Modern data-driven marketing relies on large amounts of consumer data, yet collecting such data can be costly, time-consuming, and difficult to scale. This research examines whether large language models (LLMs) can be used to generate synthetic consumer data for projective techniques, a set of methods designed to elicit consumer associations, emotions, wants, and needs. We test LLM-generated responses across multiple projective tasks, LLMs, prompting strategies, and temperature settings, and compare them with human responses from a primary research study on perceptions of city tourism destinations. Human and LLM responses were analyzed using linguistic measures, diversity and concentration metrics, topic models, and top-term analyses. The results show substantial overlap between human and LLM responses in broad topics and associations, but also important differences in style, linguistic structure, and the way diversity is generated. Recommendations are given on how to best utilize LLMs for generating synthetic consumer data, how model and prompt choices shape response quality, and on recognizing the limitations of LLM synthetic consumer data generation.
Summary
Modern data-driven marketing depends on extensive consumer datasets, yet primary collection remains expensive, slow, and difficult to scale. Researchers therefore examined whether large language models can produce usable synthetic responses for projective techniques—indirect questioning methods that surface consumer associations, emotions, wants, and needs without direct interrogation.
The study compared LLM outputs against human data collected in a primary investigation of city-tourism perceptions. Multiple projective tasks were tested across several models, prompting strategies, and temperature settings. Linguistic metrics, diversity and concentration indices, topic modeling, and top-term analysis revealed substantial overlap in broad thematic content and associations. At the same time, the generated text differed noticeably from human replies in stylistic choices, syntactic patterns, and the mechanisms that produce variety.
The authors conclude with practical guidance on model and prompt selection, the conditions under which synthetic data can supplement or replace human samples, and the persistent limitations that still require human oversight in consumer-research applications.
Why it matters
This article is highly relevant for researchers and advanced readers in the Dutch AI market as it addresses the growing need for synthetic data generation, which is crucial for navigating strict EU GDPR privacy regulations. The methodological insights into prompt engineering and model evaluation provide valuable frameworks for Dutch AI practitioners in marketing and consumer analytics.



