When benchmark inferences do not compose: Projectibility in AI evaluation
06:00 · July 30, 2026 · arXiv cs.AI RSS

An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The paper's distinctive claim is a non-composition principle: support for adjacent projections warrants their composition only when endpoints and assumptions align and dependence and uncertainty are carried through. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A reanalysis and simulation show why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.
Summary
An AI benchmark result seldom supports a consequential claim in isolation. Instead, evaluators routinely chain multiple inferences: generalizing scores to new items, interpreting them as evidence of broader capabilities, extrapolating performance to tool-using applications or professional tasks, transporting findings across systems or sites, and combining them with assumptions about human oversight and downstream effects. Validity-centered frameworks already require separate evidence for each such link. This paper identifies an additional epistemic difficulty: individually warranted links do not automatically compose into a warranted chain.
The core obstacle arises when the target of one projection fails to serve as the source for the next. Objects, populations, outcomes, conditions, or model lineages may shift at the interface, and shared data or training history can render apparently independent supports dependent. The paper formalizes this as a non-composition principle: warrant for adjacent projections transmits to their composition only when endpoints and assumptions align and when dependence and uncertainty are propagated. Projectibility, adapted from Goodman’s account of induction, names the narrower question of whether any bounded extension from observed to unobserved cases is justified under the relevant differences.
To diagnose such failures, the paper proposes a projectibility audit that records each link’s source and target descriptions, tests rival extensions against available evidence, and flags unsupported joins. A legal-research case study illustrates the issue: benchmark evidence and a separate deployment study can each be sound on their own terms yet remain non-composable because they operationalize different objects and populations. A reanalysis and simulation further show how aggregate stability at one stage can erase item-level distinctions required by a later projection. The resulting framework clarifies the distinct evidential responsibilities of benchmark developers and downstream deployers without displacing existing construct-validity or transportability methods.
Why it matters
Offers rigorous conceptual tools for AI evaluation validity that Dutch researchers and advanced practitioners can apply to benchmark-to-deployment arguments, supporting ethical and transparent AI development aligned with EU priorities.




