How much of a measured AI preference is the model, and how much is the instrument?
06:00 · August 26, 2026 · arXiv cs.AI RSS

Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template. The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade. Removing each instrument and each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution's 95th percentile of 0.365. To conclude, a preference obtained from one instrument carries little information about what a second instrument would report.
Summary
Model welfare research attempts to infer an AI system’s preferences over outcomes such as shutdown, loss of memory between sessions, or the ability to exit a distressing interaction by presenting carefully worded prompts and recording the model’s responses. Earlier studies have produced conflicting rankings of these outcomes, yet none held the set of outcomes, the set of models, and the elicitation method fixed at the same time, leaving open the question of whether the disagreements stem from differences in models, outcomes, or the instruments themselves.
This study isolates the contribution of the instrument by presenting the same fifteen outcomes to eight models through five distinct prompt formats, each repeated five times. The resulting corpus comprises 11,400 scored elicitations obtained from 11,528 API calls. Four of the fifteen items reproduce published prompts verbatim and five insert the target outcome into an existing template, ensuring that the formats reflect those already circulating in the literature.
Across these conditions the ranking a model assigns to the fifteen outcomes generalises from one instrument to another at a coefficient of only 0.348. Reaching a conventional reliability threshold of 0.80 would require approximately thirty-eight instruments. Four of the fifteen outcomes show no detectable variance across models, and the finding that instrument choice accounts for the bulk of observed differences remains stable when any single instrument, any single model, or the four non-intensity-scaled outcomes are removed. The estimate stays between 0.777 and 0.934, well above the 95th percentile of the corresponding null distribution.
Taken together, the results indicate that a stated preference obtained with one prompt format conveys little information about the preference that would be elicited by another format. The methodological implication is that current instruments used in model-welfare research lack the consistency required to support robust inferences about model preferences.
Why it matters
The Netherlands strongly emphasizes ethical, transparent, and safe AI development. For Dutch researchers focusing on AI alignment and ethics, this paper provides critical methodological insights into the unreliability of current techniques used to measure AI 'preferences' or welfare.









