{"slug": "drawingvqa-a-real-world-benchmark-for-multi-depth-visual-textual-reasoning-on", "title": "DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings", "summary": "Researchers introduced DrawingVQA, the first benchmark for evaluating multimodal large language models (MLLMs) on real-world construction drawings, featuring 33 'Issued for Construction' drawings and 92 expert-curated question-answer pairs across three reasoning depths. Evaluations of state-of-the-art MLLMs revealed a substantial gap between model and expert performance, particularly at higher reasoning depths, highlighting the need for domain-specialized multimodal reasoning in engineering workflows.", "body_md": "arXiv:2607.15418v1 Announce Type: new\nAbstract: We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows. DrawingVQA bridges this gap with 33 \"Issued for Construction\" drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning. To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction-engineering and four MLLM capability dimensions -- the first to explicitly map engineering workflows to AI reasoning competencies. Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths. This benchmark lays a foundation for domain-specialized multimodal reasoning to allow for advancement on integration of AI-driven understanding and real-world engineering workflows.", "url": "https://wpnews.pro/news/drawingvqa-a-real-world-benchmark-for-multi-depth-visual-textual-reasoning-on", "canonical_source": "https://arxiv.org/abs/2607.15418", "published_at": "2026-07-20 04:00:00+00:00", "updated_at": "2026-07-20 13:52:39.989327+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "computer-vision", "natural-language-processing", "ai-research"], "entities": ["DrawingVQA"], "alternates": {"html": "https://wpnews.pro/news/drawingvqa-a-real-world-benchmark-for-multi-depth-visual-textual-reasoning-on", "markdown": "https://wpnews.pro/news/drawingvqa-a-real-world-benchmark-for-multi-depth-visual-textual-reasoning-on.md", "text": "https://wpnews.pro/news/drawingvqa-a-real-world-benchmark-for-multi-depth-visual-textual-reasoning-on.txt", "jsonld": "https://wpnews.pro/news/drawingvqa-a-real-world-benchmark-for-multi-depth-visual-textual-reasoning-on.jsonld"}}