cd /news/large-language-models/gemini-3-6-flash-reads-charts-14-poi… · home topics large-language-models article
[ARTICLE · art-69559] src=pub.towardsai.net ↗ pub= topic=large-language-models verified=true sentiment=↓ negative

Gemini 3.6 Flash Reads Charts 14 Points Worse Than the Model It Replaced — LlamaIndex Ran the…

LlamaIndex CEO Jerry Liu ran ParseBench on Google's Gemini 3.6 Flash and found it scores 14 points worse on chart reading than the model it replaced, dropping from 45.0 to 31.0, with overall document scores falling from 69.9 to 66.8, revealing that the upgrade is a downgrade for document understanding despite gains in agentic coding.

read1 min views1 publishedJul 23, 2026
Gemini 3.6 Flash Reads Charts 14 Points Worse Than the Model It Replaced — LlamaIndex Ran the…
Image: Pub (auto-discovered)

Member-only story

Yesterday everyone cheered Gemini 3.6 Flash for topping the computer-use leaderboards. Then Jerry Liu, the CEO of LlamaIndex, ran it through a document-understanding benchmark and found the opposite story hiding underneath: on charts, the shiny new Flash model scores 14 points worse than the model it replaced — 45.0 down to 31.0. If you use Gemini Flash to read documents, the “upgrade” is a downgrade.

This is the part of a model launch nobody puts in the blog post. Gemini 3.6 Flash is genuinely better at agentic coding and driving a desktop. But Liu’s ParseBench run shows that the same post-training that bought those agentic gains quietly cost the model its eyesight for documents. Charts collapsed, tables slipped, and the overall document score fell from 69.9 to 66.8. If your pipeline OCRs invoices, parses tables, or reads charts, bumping the version number is exactly the wrong move — and here’s the data, plus how to check it against your own documents before you ship the regression to production.

The benchmark that flips the launch narrative #

ParseBench scores models on real document-parsing tasks — extracting tables, reading charts, preserving layout, staying faithful to the text — rather than…

── more in #large-language-models 4 stories · sorted by recency
── more on @llamaindex 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gemini-3-6-flash-rea…] indexed:0 read:1min 2026-07-23 ·