{"slug": "gemini-3-6-flash-reads-charts-14-points-worse-than-the-model-it-replaced-ran-the", "title": "Gemini 3.6 Flash Reads Charts 14 Points Worse Than the Model It Replaced — LlamaIndex Ran the…", "summary": "LlamaIndex CEO Jerry Liu ran ParseBench on Google's Gemini 3.6 Flash and found it scores 14 points worse on chart reading than the model it replaced, dropping from 45.0 to 31.0, with overall document scores falling from 69.9 to 66.8, revealing that the upgrade is a downgrade for document understanding despite gains in agentic coding.", "body_md": "Member-only story\n\n# Gemini 3.6 Flash Reads Charts 14 Points Worse Than the Model It Replaced — LlamaIndex Ran the Numbers\n\nYesterday everyone cheered Gemini 3.6 Flash for topping the computer-use leaderboards. Then Jerry Liu, the CEO of LlamaIndex, ran it through a document-understanding benchmark and found the opposite story hiding underneath: on charts, the shiny new Flash model scores **14 points worse than the model it replaced** — 45.0 down to 31.0. If you use Gemini Flash to read documents, the “upgrade” is a downgrade.\n\nThis is the part of a model launch nobody puts in the blog post. Gemini 3.6 Flash is genuinely better at agentic coding and driving a desktop. But Liu’s ParseBench run shows that the same post-training that bought those agentic gains quietly cost the model its eyesight for documents. Charts collapsed, tables slipped, and the overall document score fell from 69.9 to 66.8. If your pipeline OCRs invoices, parses tables, or reads charts, bumping the version number is exactly the wrong move — and here’s the data, plus how to check it against your own documents before you ship the regression to production.\n\n## The benchmark that flips the launch narrative\n\nParseBench scores models on real document-parsing tasks — extracting tables, reading charts, preserving layout, staying faithful to the text — rather than…", "url": "https://wpnews.pro/news/gemini-3-6-flash-reads-charts-14-points-worse-than-the-model-it-replaced-ran-the", "canonical_source": "https://pub.towardsai.net/gemini-3-6-flash-reads-charts-14-points-worse-than-the-model-it-replaced-llamaindex-ran-the-f05abb598287?source=rss----98111c9905da---4", "published_at": "2026-07-23 03:42:23+00:00", "updated_at": "2026-07-23 03:57:23.447070+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-research"], "entities": ["LlamaIndex", "Jerry Liu", "Google", "Gemini 3.6 Flash", "ParseBench"], "alternates": {"html": "https://wpnews.pro/news/gemini-3-6-flash-reads-charts-14-points-worse-than-the-model-it-replaced-ran-the", "markdown": "https://wpnews.pro/news/gemini-3-6-flash-reads-charts-14-points-worse-than-the-model-it-replaced-ran-the.md", "text": "https://wpnews.pro/news/gemini-3-6-flash-reads-charts-14-points-worse-than-the-model-it-replaced-ran-the.txt", "jsonld": "https://wpnews.pro/news/gemini-3-6-flash-reads-charts-14-points-worse-than-the-model-it-replaced-ran-the.jsonld"}}