Measuring Intelligence Beyond Human Scale
06:00 · July 9, 2026 · arXiv cs.AI RSS

How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and propose a new paradigm based on relative measurement in which models generate public challenges that separate other systems. Aggregating these outcomes yields an adversarial psychometric rating system that can scale with the systems being measured. We describe practical protocols that reduce incentives for private-information attacks, support judge-free adjudication, and naturally scale with agent capabilities. We instantiate the framework across verifiable and open-ended, non-verifiable domains, illustrating how model-generated evaluation can continue to measure systems beyond the human frontier.
Summary
Human-authored benchmarks for artificial intelligence have begun to saturate, leaving researchers without reliable ways to assess systems that exceed human performance. Once models surpass the capabilities of their evaluators, it becomes difficult to identify tasks that are simultaneously difficult, verifiable, and meaningful. The paper argues that this limitation is intrinsic to any absolute-scale approach that relies on fixed, human-designed tests.
Instead, the authors outline a relative-measurement framework in which AI systems themselves generate public challenges for other models. By pitting systems against one another through these model-authored tasks, outcomes can be aggregated into an adversarial psychometric rating that ranks capabilities without requiring human judges. The approach draws on established ideas from psychometrics but adapts them to an open, scalable setting where stronger agents continually produce harder tests.
The framework includes concrete protocols designed to limit opportunities for private-information attacks, ensure transparent adjudication, and operate across both verifiable domains, such as formal mathematics or coding problems, and open-ended domains where correctness is harder to define. Because the difficulty of the generated challenges can grow with the agents under test, the rating system remains informative even as performance moves further beyond the human frontier.
Why it matters
This research is highly relevant for Dutch AI researchers and auditors developing robust evaluation frameworks for advanced AI systems, aligning with the EU AI Act's focus on rigorous model benchmarking. It offers a scalable solution to the saturation of current human-authored benchmarks.






