{"slug": "scaffold-a-large-scale-structured-dataset-of-computer-science-research-figures", "title": "SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces", "summary": "Researchers released SCAFFOLD, a large-scale structured dataset pairing computer science research figures with captions, context, question-answer pairs, and chain-of-thought reasoning traces, built from arXiv papers. The dataset includes SCAFFOLD-157K (29,887 figures, 157,387 pairs from 3,058 papers), SCAFFOLD-37K (36,797 pairs), and SCAFFOLD-12K (12,000 pairs), with baseline experiments run on Qwen2.5-VL-3B-Instruct using the 12K subset. The dataset aims to fill a gap in training vision-language models to understand diagrams in academic papers.", "body_md": "arXiv:2609.00018v1 Announce Type: new\nAbstract: Computer science papers rely heavily on diagrams: architecture drawings, system flowcharts, and pipeline schematics that often carry more information than the text around them. There is currently no public dataset that pairs this specific kind of figure with captions, context, questions, answers, and step-by-step reasoning, which is exactly what is needed to train a vision-language model to understand them. We present \\textbf{SCAFFOLD}\\footnote{https://github.com/theranjitraut/scaffold}, a large-scale structured dataset of computer science research figures with diagram QA and Chain-of-Thought reasoning traces. This dataset consists of (image, caption, context, question-answer, chain-of-thought) tuples from arXiv computer science papers prepared using layout detection and PDF parsing, with an AI-assisted question-generation step. The resulting large-sized SCAFFOLD-157K dataset spans 3,058 papers with 29,887 figures (157,387 pairs), a medium-sized SCAFFOLD-37K dataset (36,797 pairs), and a small-sized SCAFFOLD-12K dataset (12,000 pairs). We used SCAFFOLD-12K for baseline experiments on Qwen2.5-VL-3B-Instruct.", "url": "https://wpnews.pro/news/scaffold-a-large-scale-structured-dataset-of-computer-science-research-figures", "canonical_source": "https://arxiv.org/abs/2609.00018", "published_at": "2026-09-02 04:00:00+00:00", "updated_at": "2026-09-02 04:26:28.572945+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "ai-research", "ai-tools"], "entities": ["SCAFFOLD", "arXiv", "Qwen2.5-VL-3B-Instruct"], "alternates": {"html": "https://wpnews.pro/news/scaffold-a-large-scale-structured-dataset-of-computer-science-research-figures", "markdown": "https://wpnews.pro/news/scaffold-a-large-scale-structured-dataset-of-computer-science-research-figures.md", "text": "https://wpnews.pro/news/scaffold-a-large-scale-structured-dataset-of-computer-science-research-figures.txt", "jsonld": "https://wpnews.pro/news/scaffold-a-large-scale-structured-dataset-of-computer-science-research-figures.jsonld"}}