Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks
06:00 · July 9, 2026 · arXiv cs.AI RSS

Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging. We present ImagingBench, a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration. ImagingBench evaluates three complementary settings: Expert, fixed expert-guided inverse reconstruction; Planner, planner-guided inverse reconstruction; and Forward, forward-system simulation for consistency checking. We benchmark leading proprietary and open-source image-centric multimodal systems, including Gemini, GPT, and Qwen, and compare them with representative task-specific non-agentic baselines. Across tasks, agentic models remain consistently weaker than specialized methods, especially on computational sensing problems such as lensless imaging, event-based reconstruction, time-of-flight imaging, and holography. Planner guidance provides only modest and inconsistent gains over the fixed-prompt Expert baseline. Although the models often generate visually plausible outputs, their reference-based fidelity remains poor, revealing a substantial gap between semantic visual competence and physically grounded imaging performance. ImagingBench provides a unified testbed for measuring this gap and tracking progress in agentic AI for computational imaging.
Summary
The paper presents ImagingBench, a benchmark that tests whether vision-language models and agentic systems can address the physics and inverse problems central to computational imaging. It comprises 20 tasks grouped into five categories—ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration—and evaluates models under three distinct protocols: a fixed Expert setting that supplies expert-guided prompts for reconstruction, a Planner setting that adds dynamic planner guidance, and a Forward setting that checks consistency by simulating the forward imaging process.
Proprietary and open-source systems, including Gemini, GPT, and Qwen, are compared against representative task-specific, non-agentic baselines. Across the suite, agentic models trail the specialized methods, with the largest shortfalls appearing in computational sensing problems such as lensless imaging, event-based reconstruction, time-of-flight imaging, and holography. Although the generated images frequently look plausible, reference-based fidelity metrics remain low, indicating that semantic visual competence does not translate into accurate recovery of the underlying physical quantities.
Planner guidance produces only modest and inconsistent gains relative to the simpler Expert baseline. The benchmark therefore supplies a unified testbed for measuring the gap between current agentic performance and the requirements of physically grounded imaging tasks, while also providing a means to track future progress in this area.
Why it matters
This research is highly relevant for the Dutch AI market, particularly for its strong medical imaging and high-tech optics sectors. It provides researchers with a crucial benchmark to understand the limitations of current VLMs in solving complex, physics-based inverse problems.




