I showed 30 AI models 336 monitoring dashboards. Most of them can't tell time. A developer built a Kaggle benchmark of 336 synthetic Grafana-style monitoring panels to test whether vision models can read dashboards the way an on-call engineer does, grading each model on five fields: event type, start time, implicated deploy, recovery status, and peak value. Across 30 vision models from four vendors, the top models read the charts almost perfectly, while most others identified the peak but failed to read the clock and made the same error when uncertain; the best models also surfaced three bugs in the answer key. Canary fields for the y-axis label and dashed-marker count exposed models that accept an image but cannot actually see it. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 Most incidents start with a graph. Not a log line, but a panel: a latency line that jumped, a CPU series pinned at the ceiling, a gap where a metric stopped. The first thirty seconds of on-call are spent reading a picture . AI copilots for on-call usually get the evidence as text. Real incidents hand you a screenshot. So I built a benchmark for the step the demos skip: can a model read a monitoring panel the way an on-call engineer does? I ran it against 30 vision models on Kaggle. The best read these charts almost perfectly. Most of the rest find the peak but cannot read the clock, and when unsure they make the same mistake. Along the way, the best models also found three bugs in my answer key. The capability: turning one monitoring panel into five facts an engineer would act on. | Field | Question | Graded as | |---|---|---| | What happened | step change, ramp, spike, flapping, plateau, gap, or none | exact match | | When | the clock time it started, read off the x-axis | within ±2 min ±6 for slow ramps | | Which deploy | the dashed deploy marker A/B/C that lines up with the start, or none | exact match | | Recovering? | is the series back to its earlier level by the right edge | exact match | | How bad | the peak value, in the axis units | within ±10% | A model's score is the mean of those five fields over every panel. The panels are synthetic, and that is the point. A generator renders Grafana-style panels and records the ground truth, so nothing is hand-labelled and no model has seen them before. 336 panels cover 7 event shapes × dark/light × 75/200 dpi × 12 variants, randomising the metric, 1–4 series, 0–3 deploy markers, and everyday traps: Here is one panel. Latency on auth-gw climbs from 15:26, right at deploy A, and pins flat at 0.29 s. That is a plateau saturation , not a step change. | Model | What happened | When | Right? | |---|---|---|---| | Gemini 3.7 Flash | plateau | 15:25 | ✅ | | GPT-6 Astra | plateau | 15:25 | ✅ | | Claude Opus 5 | plateau | 15:25 | ✅ | | Claude Sonnet 4.5 | step change | 15:30 | ❌ | | Claude Haiku 4.5 | step change | 15:20 | ❌ | | GPT-5.4 nano | step change | 15:30 | ❌ | | Grok 4.20 non-reasoning | step change | 15:20 | ❌ | Canary fields catch models that cannot see. Every answer also reports the y-axis label and the number of dashed markers, which are trivial if you can see the image. That turned out to matter see "Blind but confident" . How it runs on Kaggle. It is one task built with the kaggle-benchmarks https://github.com/Kaggle/kaggle-benchmarks SDK. The PNGs and their answer key live in a Kaggle dataset attached to the task. A per-panel sub-task sends one image and asks for a typed answer, in its own isolated chat: @dataclass class Reading: event type: str step change, ramp, spike, flapping, plateau, gap, none start time: str "HH:MM" read off the x-axis, or "none" implicated deploy: str "A", "B", "C" or "none" recovering: bool peak value: float y axis label: str canary dashed marker count: int canary with kbench.chats.new f"panel-{id}" : r = llm.prompt PROMPT, image=images.from path path , schema=Reading, extra api params={"max tokens": 16384} The parent task runs all 336 panels with .evaluate and returns the composite score, with six assertions for the per-field breakdown. Each panel gets a 16,384-token output budget, above every model's 99th percentile. 30 vision models from four vendors, chosen to walk each vendor's price ladder rather than just crown a winner: A few catalog models could not be scored: Claude Opus 4.1, Claude Sonnet 4 and two Grok versions returned "model not found", DeepSeek-R1 rejects images, and three models accept the image but cannot see it more below . Every model ran all 336 panels, for 10,080 readings in the final version. Twenty-one models also ran two or three times on earlier versions with the same prompt; between runs a score moved by a median of 0.005 and never more than 0.015. Gemini 3.7 Flash leads at 0.978. Gemini 3.8 Flash 0.972 , Gemini 3.5 Flash 0.963 and GPT-6 Astra 0.961 sit within 0.02 behind it, with overlapping confidence intervals. The bills do not overlap: one full 336-panel run cost $1.51 on Gemini 3.7 Flash and $7.23 on GPT-6 Astra. Price does not buy chart reading. The most expensive run, Gemini 2.5 Pro at $9.80, scored 0.790 , and Gemini 3.1 Pro $7.75, 0.921 also trails every current Gemini Flash. Claude Opus 5 0.871, $5.41 sits just below GPT-5.6 Luna 0.877, $0.31 . On a budget, Luna and Gemma 4 31B 0.855, $0.46 are the bargains. Split the score into its five fields and one column does almost all the work. Peak value is easy for every model 0.83–1.00 . The millisecond-labelled-as-seconds trap barely registered: the worst drop on its 37 panels was 0.08, for Claude Opus 4.5. Start time is where models separate. The top five land within two minutes 87–91% of the time; the rest range from 20% to 80%, and 17 of the 30 models get it right less than half the time. I expected weak models to snap to printed tick labels. They don't: they land on a tick about 4% of the time, the same as the true starts. They are simply imprecise. For the eight weakest, the median miss is 5–9 minutes, and one answer in ten is off by 15 to 45. At 3 a.m., "the deploy at 15:20 or the one at 15:30" is the whole question. The top five name the event correctly almost every time 97–100% for every shape . The bottom ten have a default answer: step change . They give it for 56% of plateaus, 34% of flapping series and 25% of gaps. And 44% of real gaps get called "none". A metric that stopped reporting is often the actual incident. This is the gap panel. payments stops reporting at about 04:33 and resumes about ten minutes later, right as deploy C lands. Gemini 3.7 Flash, GPT-6 Astra and Claude Opus 5 all said gap, recovering, no deploy to blame . Claude Sonnet 4.5, Claude Haiku 4.5 and GPT-5.4 nano said none . Grok 4.20 non-reasoning said none , and still blamed deploy C. OpenAI's ladder is the steepest. GPT-5.4 nano 0.561 misreads the axis label or marker count on 64% of panels, and GPT-6 Astra is in the top four. Anthropic's ladder climbs steadily, from Haiku 4.5 at 0.596 to Sonnet 5 at 0.801 and Opus 5 at 0.871. Its top model still sits below Google's mid-tier Flash. Within Google, generation matters more than tier: both current Flash-Lites beat the older 2.5 Flash, and every current Flash beats both Pro models. The thinking dial helps, at a price. The same Grok 4.20 scored 0.646 without reasoning and 0.694 with it , and the reasoning run cost 11 times as much. Half the panels are rendered at 75 dpi, about the size of a thumbnail pasted into chat, and half at 200 dpi. For every Gemini and Gemma model the difference is noise ±0.02 . For every Claude model it is 0.07 to 0.14 , and for OpenAI's models 0.02 to 0.09. At 75 dpi Claude Haiku 4.5 reads the canary fields correctly on 62% of panels; at 200 dpi, on 94%. At 200 dpi alone the leaders barely move, but Claude Opus 5 climbs to 0.905 and Sonnet 5 to 0.845. If you send screenshots to a Claude or GPT model, send them big. The decoy series splits the same way: it costs Claude Haiku 4.5 and GPT-5.4 nano about 0.10, and current Gemini Flash models nothing measurable. This was the finding I did not expect. On an earlier run of the benchmark, GLM-5 accepted 99.7% of the images without complaint and got the canary fields right on under 1% of them. It reported axis labels like "requests per second" on CPU panels and counted dashed markers on panels that had none. It answered "step change" on 305 of 336 panels , each with a start time, a deploy letter and a peak value, all invented. Qwen3-235B, given the same images, answered "none" 319 times with an empty axis label. Same proxy, same images. Qwen answered as if every panel were empty; GLM-5 produced 336 confident, fabricated incident readings. If you wire a model into an on-call tool, check that it can see before you check whether it is smart. A two-field canary costs nothing. My first full run used an answer key built from what the generator intended . Then I looked for panels where four or five of the top five models agreed with each other and disagreed with my key. There were 32 of them , in three groups: I rebuilt the generator so every label comes from the rendered data, with a self-check that rejects any panel whose labels cannot be read off the chart. Then I re-ran all 30 models. On the new key, event type, recovery and peak have zero disputed panels. Scores rose by up to 0.06; two models dipped by under 0.01, and no model moved more than two places. One ambiguity is real rather than a bug. When does a gradual ramp start? At the first bend, or when it is clearly above the noise? On 16 slow ramps the top models split between those two readings. I re-scored ramps to accept any answer between the two. Scores moved by 0.009 on average and the ranking held, except GPT-6 Astra moving from fourth to second inside the top cluster. The public leaderboard uses the stricter key. The lesson I'd keep: when the strongest models agree with each other and not with you, check your key first. The resolution split followed vendor lines, not size or price. The unit trap I expected to be hardest was the easiest. And the best reader cost a fifth of GPT-6 Astra, which it beat. | | Model | Score | What | When | Deploy | Recovering | Peak | Cost/run | |---|---|---|---|---|---|---|---|---| | 1 | Gemini 3.7 Flash | 0.978 | 1.00 | 0.91 | 0.99 | 1.00 | 1.00 | $1.51 | | 2 | Gemini 3.8 Flash | 0.972 | 1.00 | 0.90 | 0.99 | 1.00 | 0.99 | $4.15 | | 3 | Gemini 3.5 Flash | 0.963 | 0.98 | 0.88 | 0.97 | 0.99 | 0.99 | $4.53 | | 4 | GPT-6 Astra | 0.961 | 1.00 | 0.89 | 0.93 | 0.99 | 1.00 | $7.23 | | 5 | Gemini 3.6 Flash | 0.958 | 0.97 | 0.89 | 0.96 | 0.97 | 0.99 | $2.25 | | 6 | GPT-5.5 | 0.949 | 0.98 | 0.85 | 0.95 | 0.98 | 0.98 | $7.50 | | 7 | Gemini 3 Flash preview | 0.940 | 0.97 | 0.80 | 0.96 | 0.99 | 0.99 | $3.25 | | 8 | Gemini 3.1 Pro preview | 0.921 | 0.98 | 0.67 | 0.98 | 0.98 | 0.99 | $7.75 | | 9 | GPT-5.6 Terra | 0.901 | 0.95 | 0.73 | 0.90 | 0.96 | 0.97 | $2.03 | | 10 | GPT-5.6 Luna | 0.877 | 0.92 | 0.69 | 0.88 | 0.92 | 0.97 | $0.31 | | 11 | Claude Opus 5 | 0.871 | 0.97 | 0.47 | 0.93 | 0.99 | 0.99 | $5.41 | | 12 | Gemma 4 31B | 0.855 | 0.97 | 0.39 | 0.95 | 0.97 | 0.99 | $0.46 | | 13 | Gemini 3.1 Flash-Lite | 0.843 | 0.91 | 0.47 | 0.91 | 0.94 | 0.99 | $0.16 | | 14 | Claude Opus 4.8 | 0.840 | 0.84 | 0.59 | 0.87 | 0.92 | 0.98 | $3.50 | | 15 | Gemma 4 26B | 0.827 | 0.88 | 0.40 | 0.93 | 0.93 | 0.99 | $0.59 | | 16 | Claude Opus 4.7 | 0.823 | 0.86 | 0.55 | 0.88 | 0.85 | 0.98 | $3.50 | | 17 | Gemini 3.5 Flash-Lite | 0.823 | 0.89 | 0.43 | 0.88 | 0.92 | 0.99 | $0.21 | | 18 | Claude Sonnet 5 | 0.801 | 0.85 | 0.49 | 0.81 | 0.89 | 0.97 | $1.45 | | 19 | Gemini 2.5 Pro | 0.790 | 0.84 | 0.39 | 0.85 | 0.90 | 0.97 | $9.80 | | 20 | GPT-5.4 | 0.782 | 0.90 | 0.37 | 0.83 | 0.85 | 0.95 | $1.21 | | 21 | GPT-5.4 mini | 0.751 | 0.68 | 0.50 | 0.77 | 0.83 | 0.96 | $0.36 | | 22 | Gemini 2.5 Flash | 0.739 | 0.72 | 0.26 | 0.84 | 0.91 | 0.97 | $1.25 | | 23 | Claude Sonnet 4.6 | 0.736 | 0.78 | 0.34 | 0.75 | 0.89 | 0.92 | $1.83 | | 24 | Claude Opus 4.6 | 0.711 | 0.72 | 0.36 | 0.76 | 0.82 | 0.89 | $3.05 | | 25 | Grok 4.20 reasoning | 0.694 | 0.72 | 0.20 | 0.76 | 0.85 | 0.95 | $3.32 | | 26 | Claude Opus 4.5 | 0.677 | 0.63 | 0.33 | 0.73 | 0.82 | 0.88 | $3.05 | | 27 | Grok 4.20 non-reasoning | 0.646 | 0.56 | 0.30 | 0.65 | 0.81 | 0.90 | $0.30 | | 28 | Claude Haiku 4.5 | 0.596 | 0.50 | 0.24 | 0.59 | 0.79 | 0.87 | $0.61 | | 29 | Claude Sonnet 4.5 | 0.596 | 0.37 | 0.25 | 0.67 | 0.82 | 0.88 | $1.83 | | 30 | GPT-5.4 nano | 0.561 | 0.47 | 0.27 | 0.53 | 0.71 | 0.83 | $0.10 | Reproducibility. The generator is seeded seed 2026 and regenerates the identical 336 images. The grader is deterministic, with no LLM judge. Limitations, honestly. The panels are synthetic and single-panel. The event vocabulary is mine "plateau" vs "saturation" . "Recovering" mostly follows from the event type. On 5 panels two deploy labels overlap, though the marker lines stay distinct. Each model saw each panel once in the final run; earlier repeats barely moved the ranking. I built this with an AI coding assistant Claude helping with the generator, the Kaggle harness and the analysis. The benchmark design, the audit decisions and every number here were checked against the run data.