AI Evaluation Should Work With Humans
06:00 · August 17, 2026 · arXiv cs.AI RSS

This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction. Instead, the AI community should pivot to evaluating the performance of human--AI teams. We argue that this collaborative shift will foster AI systems that act as true complements to human capabilities and therefore lead to far better societal outcomes than will the current process.
Summary
The dominant approach to AI evaluation measures systems against the benchmark of autonomous, superhuman performance on narrowly defined tasks. This framing, the authors contend, steers development toward the implicit objective of replacing human labor rather than supporting it. By rewarding models that operate independently and outperform people, current metrics risk producing tools that are poorly aligned with the ways humans actually work and make decisions.
The position paper therefore calls for a shift in evaluation practice toward joint human–AI performance. Instead of isolating model capabilities, assessments would examine how effectively teams of people and systems solve problems together. Such metrics would reward designs that fill gaps in human expertise, reduce cognitive load, or improve decision quality under realistic constraints.
The authors argue that this collaborative orientation would encourage the creation of AI systems that function as genuine complements to human abilities. The resulting technologies, they maintain, are more likely to deliver positive societal outcomes than systems optimized solely for standalone superiority.
Why it matters
This paper aligns strongly with the Dutch and EU focus on ethical, human-centric AI and human oversight. It provides researchers with a conceptual foundation to develop new evaluation frameworks that prioritize human-AI collaboration over autonomous replacement, which is highly actionable for Dutch AI policy and enterprise deployment.










