{"slug": "refigbench-benchmarking-scientific-figure-reconstruction-as-editable-powerpoint", "title": "ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts", "summary": "A new benchmark called ReFigBench, built on 1,000 real overview figures retrieved from arXiv papers, evaluates multimodal coding agents on reconstructing scientific figures as editable PowerPoint slides. Testing ten configurations across four model families and two commercial harnesses, the researchers found perception remains a bottleneck, the same model can gain from a specialized PPTX workflow inside one harness and lose inside another, and the specialized workflow erases native connectors in every configuration even as human judges prefer its renderings in most matchups. Even the strongest agent fell short of the rubric ceiling, exposing a tension between fidelity and editability for practical multimodal document agents.", "body_md": "arXiv:2609.18844v1 Announce Type: new \nAbstract: Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often isolate short tool calls, API traces, or screenshot resemblance, and a low score under these proxies cannot say whether the model saw poorly, planned poorly, or was failed by its harness. We study scientific overview figure reconstruction, an agent task in which a source image must become an editable PowerPoint slide that preserves text, topology, layout, and native document structure. We introduce ReFigBench, a benchmark and evaluation framework built on 1,000 real overview figures retrieved from arXiv papers with full provenance. Coding agents from four model families reconstruct every figure under two workflows, direct code generation and a specialized PPTX workflow, and the strongest model runs inside two commercial harnesses, yielding ten configurations. Evaluation combines deterministic artifact checks, repeated automated scoring by judges from two model families, and blinded human comparisons. Perception remains a bottleneck that iterative rendering only partly repays. Whether workflow effort converts into quality depends on the model together with its harness, since the same model gains from the specialized workflow inside one harness and loses inside the other, and the harness shifts scores even under an identical direct prompt. The specialized workflow erases native connectors in every configuration, human judges still prefer its renderings in most matchups, and even the strongest agent falls short of the rubric ceiling. These results expose the tension between fidelity and editability as the central challenge for practical multimodal document agents.", "url": "https://wpnews.pro/news/refigbench-benchmarking-scientific-figure-reconstruction-as-editable-powerpoint", "canonical_source": "https://www.machinebrief.com/news/refigbench-benchmarking-scientific-figure-reconstruction-as-7ak5", "published_at": "2026-09-17 04:00:00+00:00", "updated_at": "2026-09-17 06:55:01.298381+00:00", "lang": "en", "topics": ["ai-agents", "ai-research", "computer-vision", "ai-tools"], "entities": ["ReFigBench", "arXiv", "PowerPoint", "PPTX"], "alternates": {"html": "https://wpnews.pro/news/refigbench-benchmarking-scientific-figure-reconstruction-as-editable-powerpoint", "markdown": "https://wpnews.pro/news/refigbench-benchmarking-scientific-figure-reconstruction-as-editable-powerpoint.md", "text": "https://wpnews.pro/news/refigbench-benchmarking-scientific-figure-reconstruction-as-editable-powerpoint.txt", "jsonld": "https://wpnews.pro/news/refigbench-benchmarking-scientific-figure-reconstruction-as-editable-powerpoint.jsonld"}}