DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
06:00 · July 20, 2026 · arXiv cs.AI RSS

We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows. DrawingVQA bridges this gap with 33 "Issued for Construction" drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning. To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction-engineering and four MLLM capability dimensions -- the first to explicitly map engineering workflows to AI reasoning competencies. Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths. This benchmark lays a foundation for domain-specialized multimodal reasoning to allow for advancement on integration of AI-driven understanding and real-world engineering workflows.
Summary
DrawingVQA addresses the absence of benchmarks that test multimodal large language models on the dense visual-textual artifacts actually used by civil and structural engineers. Unlike prior VQA collections built on natural images, schematic floor plans, or exam-style questions, the new benchmark draws exclusively from 33 “Issued for Construction” structural drawings—legally binding documents that combine precise geometry, engineering symbols, cross-sheet references, tables, and annotations.
The dataset contains 92 expert-authored question-answer pairs, each labeled with one of three reasoning depths: perceptual recognition, contextual interpretation, or expert-level domain judgment. A dual categorization scheme further situates every pair within seven construction-engineering tasks (such as quantity take-off and code compliance) and four core MLLM capabilities (visual perception, OCR and text understanding, knowledge, and reasoning). This mapping supplies the first explicit diagnostic link between professional engineering workflows and the internal competencies of current models.
When state-of-the-art MLLMs are evaluated on the benchmark, their accuracy falls well below that of practicing engineers, with the largest shortfalls appearing at the expert-reasoning depth and on quantity-take-off questions. The authors release the full set of drawings, questions, and annotations to enable reproducible measurement of progress toward multimodal systems that can support real engineering practice.
Why it matters
Directly relevant to Dutch researchers developing or evaluating multimodal models for engineering and construction workflows; offers actionable benchmark, metrics, and failure analysis that can be applied by NL teams in AEC and AI.



