cd /news/artificial-intelligence/drawingvqa-a-real-world-benchmark-fo… · home topics artificial-intelligence article
[ARTICLE · art-65472] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

Researchers introduced DrawingVQA, the first benchmark for evaluating multimodal large language models on real-world construction drawings, which combine abstract geometry, symbolic notation, tabular data, and domain-specific text. The benchmark includes 33 construction drawings and 92 question-answer pairs across three reasoning depths, revealing a substantial gap between model and expert performance, particularly at higher reasoning levels. This work lays a foundation for domain-specialized multimodal reasoning in engineering workflows.

read1 min views1 publishedJul 20, 2026

arXiv:2607.15418v1 Announce Type: new Abstract: We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows. DrawingVQA bridges this gap with 33 "Issued for Construction" drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning. To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction-engineering and four MLLM capability dimensions -- the first to explicitly map engineering workflows to AI reasoning competencies. Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths. This benchmark lays a foundation for domain-specialized multimodal reasoning to allow for advancement on integration of AI-driven understanding and real-world engineering workflows.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @drawingvqa 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/drawingvqa-a-real-wo…] indexed:0 read:1min 2026-07-20 ·