Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
06:00 · August 26, 2026 · arXiv cs.AI RSS

When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day "page-a-day" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.
Summary
A recent arXiv paper reports the first quantified scene-level audit of LLM-generated autobiography, measuring how much of the generated narrative can be corroborated against an independent, subject-specific ground-truth corpus. The study covers 366 daily first-person entries produced by prompting a conversational model with only a template, two exemplar days, and a daily quote, without access to the author’s actual records. Using a four-level rubric fixed before analysis, the audit finds that 354 days fail positive verification, for a 96.7 percent verification-failure rate.
The dominant failure mode is grounded drift: the model places real people, employers, and settings inside invented scenes rather than fabricating entirely new entities. Nineteen days contain claims directly contradicted by the record, while only twelve days include a scene that can be positively corroborated. When the same inputs are supplied to current named models, verification failure remains at 100 percent, indicating that the problem is not limited to older model versions.
Independent re-rating by two additional reviewers reproduces the headline failure rate, confirming that the original measurement was not inflated. However, agreement on the finer taxonomy is only fair to moderate, largely because the boundary between “weak” and “unverified” scenes hinges on an imprecise notion of setting. Supplying the generator with excerpts from the subject’s own corpus improves the verification rate, yet leaves a residual failure of 83.3 percent.
The work releases the full audit instrument, the nine-theme false-premise taxonomy used for screening, the generation prompts, and all reproducibility artifacts, allowing other researchers to apply the same scene-level method to additional corpora or models.
Why it matters
Directly relevant for Dutch AI researchers studying LLM evaluation, hallucination mitigation, ethical person simulation, and digital twins; methods are actionable for EU-compliant trustworthy AI development.










