# I showed 30 AI models 336 monitoring dashboards. Most of them can't tell time.

> Source: <https://dev.to/uptimearchitect/i-showed-30-ai-models-336-monitoring-dashboards-most-of-them-cant-tell-time-h0m>
> Published: 2026-10-02 11:29:39+00:00

*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*

Most incidents start with a graph.

Not a log line, but a panel: a latency line that jumped, a CPU series pinned at the ceiling, a gap where a metric stopped. The first thirty seconds of on-call are spent *reading a picture*.

AI copilots for on-call usually get the evidence as text. Real incidents hand you a screenshot. So I built a benchmark for the step the demos skip: **can a model read a monitoring panel the way an on-call engineer does?**

I ran it against 30 vision models on Kaggle. The best read these charts almost perfectly. Most of the rest find the peak but cannot read the clock, and when unsure they make the same mistake. Along the way, the best models also found three bugs in my answer key.

**The capability:** turning one monitoring panel into five facts an engineer would act on.

| Field | Question | Graded as | 
|---|---|---|
| What happened | step change, ramp, spike, flapping, plateau, gap, or none | exact match | 
| When | the clock time it started, read off the x-axis | within ±2 min (±6 for slow ramps) | 
| Which deploy | the dashed deploy marker (A/B/C) that lines up with the start, or none | exact match | 
| Recovering? | is the series back to its earlier level by the right edge | exact match | 
| How bad | the peak value, in the axis units | within ±10% | 

A model's score is the mean of those five fields over every panel.

**The panels are synthetic, and that is the point.** A generator renders Grafana-style panels and records the ground truth, so nothing is hand-labelled and no model has seen them before. **336 panels** cover 7 event shapes × dark/light × 75/200 dpi × 12 variants, randomising the metric, 1–4 series, 0–3 deploy markers, and everyday traps:

Here is one panel. Latency on `auth-gw` climbs from 15:26, right at deploy A, and pins flat at 0.29 s. That is a **plateau** (saturation), not a step change.

| Model | What happened | When | Right? | 
|---|---|---|---|
| Gemini 3.7 Flash | plateau | 15:25 | ✅ | 
| GPT-6 Astra | plateau | 15:25 | ✅ | 
| Claude Opus 5 | plateau | 15:25 | ✅ | 
| Claude Sonnet 4.5 | step change | 15:30 | ❌ | 
| Claude Haiku 4.5 | step change | 15:20 | ❌ | 
| GPT-5.4 nano | step change | 15:30 | ❌ | 
| Grok 4.20 (non-reasoning) | step change | 15:20 | ❌ | 

**Canary fields catch models that cannot see.** Every answer also reports the y-axis label and the number of dashed markers, which are trivial if you can see the image. That turned out to matter (see "Blind but confident").

**How it runs on Kaggle.** It is one task built with the [`kaggle-benchmarks`](https://github.com/Kaggle/kaggle-benchmarks) SDK. The PNGs and their answer key live in a Kaggle dataset attached to the task. A per-panel sub-task sends one image and asks for a typed answer, in its own isolated chat:

```
@dataclass
class Reading:
    event_type: str           # step_change, ramp, spike, flapping, plateau, gap, none
    start_time: str           # "HH:MM" read off the x-axis, or "none"
    implicated_deploy: str    # "A", "B", "C" or "none"
    recovering: bool
    peak_value: float
    y_axis_label: str         # canary
    dashed_marker_count: int  # canary

with kbench.chats.new(f"panel-{id}"):
    r = llm.prompt(PROMPT, image=images.from_path(path), schema=Reading,
                   extra_api_params={"max_tokens": 16384})
```

The parent task runs all 336 panels with `.evaluate()` and returns the composite score, with six assertions for the per-field breakdown. Each panel gets a 16,384-token output budget, above every model's 99th percentile.

**30 vision models from four vendors,** chosen to walk each vendor's price ladder rather than just crown a winner:

A few catalog models could not be scored: Claude Opus 4.1, Claude Sonnet 4 and two Grok versions returned "model not found", DeepSeek-R1 rejects images, and three models accept the image but cannot see it (more below).

Every model ran all 336 panels, for 10,080 readings in the final version. Twenty-one models also ran two or three times on earlier versions with the same prompt; between runs a score moved by a median of **0.005** and never more than 0.015.

**Gemini 3.7 Flash leads at 0.978.** Gemini 3.8 Flash (0.972), Gemini 3.5 Flash (0.963) and GPT-6 Astra (0.961) sit within 0.02 behind it, with overlapping confidence intervals. The bills do not overlap: one full 336-panel run cost **$1.51** on Gemini 3.7 Flash and **$7.23** on GPT-6 Astra.

Price does not buy chart reading. The most expensive run, **Gemini 2.5 Pro at $9.80, scored 0.790**, and Gemini 3.1 Pro ($7.75, 0.921) also trails every current Gemini Flash. Claude Opus 5 (0.871, $5.41) sits just below GPT-5.6 Luna (0.877, **$0.31**). On a budget, Luna and Gemma 4 31B (0.855, $0.46) are the bargains.

Split the score into its five fields and one column does almost all the work. **Peak value** is easy for every model (0.83–1.00). The millisecond-labelled-as-seconds trap barely registered: the worst drop on its 37 panels was 0.08, for Claude Opus 4.5. **Start time** is where models separate. The top five land within two minutes 87–91% of the time; the rest range from 20% to 80%, and 17 of the 30 models get it right less than half the time.

I expected weak models to snap to printed tick labels. They don't: they land on a tick about 4% of the time, the same as the true starts. They are simply imprecise. For the eight weakest, the median miss is 5–9 minutes, and one answer in ten is off by 15 to 45. At 3 a.m., "the deploy at 15:20 or the one at 15:30" is the whole question.

The top five name the event correctly almost every time (97–100% for every shape). The bottom ten have a default answer: **step change**. They give it for 56% of plateaus, 34% of flapping series and 25% of gaps. And 44% of real gaps get called "none". A metric that *stopped reporting* is often the actual incident.

This is the gap panel. `payments` stops reporting at about 04:33 and resumes about ten minutes later, right as deploy C lands.

Gemini 3.7 Flash, GPT-6 Astra and Claude Opus 5 all said *gap, recovering, no deploy to blame*. Claude Sonnet 4.5, Claude Haiku 4.5 and GPT-5.4 nano said *none*. Grok 4.20 (non-reasoning) said *none*, and still blamed deploy C.

OpenAI's ladder is the steepest. GPT-5.4 nano (0.561) misreads the axis label or marker count on 64% of panels, and GPT-6 Astra is in the top four. Anthropic's ladder climbs steadily, from Haiku 4.5 at 0.596 to Sonnet 5 at 0.801 and Opus 5 at 0.871. Its top model still sits below Google's mid-tier Flash. Within Google, generation matters more than tier: both current Flash-Lites beat the older 2.5 Flash, and every current Flash beats both Pro models.

The thinking dial helps, at a price. The same Grok 4.20 scored **0.646 without reasoning and 0.694 with it**, and the reasoning run cost 11 times as much.

Half the panels are rendered at 75 dpi, about the size of a thumbnail pasted into chat, and half at 200 dpi. For every Gemini and Gemma model the difference is noise (±0.02). For **every Claude model it is 0.07 to 0.14**, and for OpenAI's models 0.02 to 0.09. At 75 dpi Claude Haiku 4.5 reads the canary fields correctly on 62% of panels; at 200 dpi, on 94%.

At 200 dpi alone the leaders barely move, but Claude Opus 5 climbs to 0.905 and Sonnet 5 to 0.845. **If you send screenshots to a Claude or GPT model, send them big.** The decoy series splits the same way: it costs Claude Haiku 4.5 and GPT-5.4 nano about 0.10, and current Gemini Flash models nothing measurable.

This was the finding I did not expect. On an earlier run of the benchmark, **GLM-5 accepted 99.7% of the images without complaint and got the canary fields right on under 1% of them.** It reported axis labels like "requests per second" on CPU panels and counted dashed markers on panels that had none. It answered **"step change" on 305 of 336 panels**, each with a start time, a deploy letter and a peak value, all invented. Qwen3-235B, given the same images, answered "none" 319 times with an empty axis label.

Same proxy, same images. Qwen answered as if every panel were empty; GLM-5 produced 336 confident, fabricated incident readings. If you wire a model into an on-call tool, **check that it can see before you check whether it is smart.** A two-field canary costs nothing.

My first full run used an answer key built from what the generator *intended*. Then I looked for panels where four or five of the top five models agreed with each other and disagreed with my key. There were **32 of them**, in three groups:

I rebuilt the generator so every label comes from the rendered data, with a self-check that rejects any panel whose labels cannot be read off the chart. Then I re-ran all 30 models. On the new key, event type, recovery and peak have **zero** disputed panels. Scores rose by up to 0.06; two models dipped by under 0.01, and no model moved more than two places.

One ambiguity is real rather than a bug. **When does a gradual ramp start?** At the first bend, or when it is clearly above the noise? On 16 slow ramps the top models split between those two readings. I re-scored ramps to accept any answer between the two. Scores moved by 0.009 on average and the ranking held, except GPT-6 Astra moving from fourth to second inside the top cluster. The public leaderboard uses the stricter key.

The lesson I'd keep: **when the strongest models agree with each other and not with you, check your key first.**

The resolution split followed vendor lines, not size or price. The unit trap I expected to be hardest was the easiest. And the best reader cost a fifth of GPT-6 Astra, which it beat.

| # | Model | Score | What | When | Deploy | Recovering | Peak | Cost/run | 
|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.7 Flash | **0.978** | 1.00 | 0.91 | 0.99 | 1.00 | 1.00 | $1.51 | 
| 2 | Gemini 3.8 Flash | **0.972** | 1.00 | 0.90 | 0.99 | 1.00 | 0.99 | $4.15 | 
| 3 | Gemini 3.5 Flash | **0.963** | 0.98 | 0.88 | 0.97 | 0.99 | 0.99 | $4.53 | 
| 4 | GPT-6 Astra | **0.961** | 1.00 | 0.89 | 0.93 | 0.99 | 1.00 | $7.23 | 
| 5 | Gemini 3.6 Flash | **0.958** | 0.97 | 0.89 | 0.96 | 0.97 | 0.99 | $2.25 | 
| 6 | GPT-5.5 | **0.949** | 0.98 | 0.85 | 0.95 | 0.98 | 0.98 | $7.50 | 
| 7 | Gemini 3 Flash (preview) | **0.940** | 0.97 | 0.80 | 0.96 | 0.99 | 0.99 | $3.25 | 
| 8 | Gemini 3.1 Pro (preview) | **0.921** | 0.98 | 0.67 | 0.98 | 0.98 | 0.99 | $7.75 | 
| 9 | GPT-5.6 Terra | **0.901** | 0.95 | 0.73 | 0.90 | 0.96 | 0.97 | $2.03 | 
| 10 | GPT-5.6 Luna | **0.877** | 0.92 | 0.69 | 0.88 | 0.92 | 0.97 | $0.31 | 
| 11 | Claude Opus 5 | **0.871** | 0.97 | 0.47 | 0.93 | 0.99 | 0.99 | $5.41 | 
| 12 | Gemma 4 31B | **0.855** | 0.97 | 0.39 | 0.95 | 0.97 | 0.99 | $0.46 | 
| 13 | Gemini 3.1 Flash-Lite | **0.843** | 0.91 | 0.47 | 0.91 | 0.94 | 0.99 | $0.16 | 
| 14 | Claude Opus 4.8 | **0.840** | 0.84 | 0.59 | 0.87 | 0.92 | 0.98 | $3.50 | 
| 15 | Gemma 4 26B | **0.827** | 0.88 | 0.40 | 0.93 | 0.93 | 0.99 | $0.59 | 
| 16 | Claude Opus 4.7 | **0.823** | 0.86 | 0.55 | 0.88 | 0.85 | 0.98 | $3.50 | 
| 17 | Gemini 3.5 Flash-Lite | **0.823** | 0.89 | 0.43 | 0.88 | 0.92 | 0.99 | $0.21 | 
| 18 | Claude Sonnet 5 | **0.801** | 0.85 | 0.49 | 0.81 | 0.89 | 0.97 | $1.45 | 
| 19 | Gemini 2.5 Pro | **0.790** | 0.84 | 0.39 | 0.85 | 0.90 | 0.97 | $9.80 | 
| 20 | GPT-5.4 | **0.782** | 0.90 | 0.37 | 0.83 | 0.85 | 0.95 | $1.21 | 
| 21 | GPT-5.4 mini | **0.751** | 0.68 | 0.50 | 0.77 | 0.83 | 0.96 | $0.36 | 
| 22 | Gemini 2.5 Flash | **0.739** | 0.72 | 0.26 | 0.84 | 0.91 | 0.97 | $1.25 | 
| 23 | Claude Sonnet 4.6 | **0.736** | 0.78 | 0.34 | 0.75 | 0.89 | 0.92 | $1.83 | 
| 24 | Claude Opus 4.6 | **0.711** | 0.72 | 0.36 | 0.76 | 0.82 | 0.89 | $3.05 | 
| 25 | Grok 4.20 (reasoning) | **0.694** | 0.72 | 0.20 | 0.76 | 0.85 | 0.95 | $3.32 | 
| 26 | Claude Opus 4.5 | **0.677** | 0.63 | 0.33 | 0.73 | 0.82 | 0.88 | $3.05 | 
| 27 | Grok 4.20 (non-reasoning) | **0.646** | 0.56 | 0.30 | 0.65 | 0.81 | 0.90 | $0.30 | 
| 28 | Claude Haiku 4.5 | **0.596** | 0.50 | 0.24 | 0.59 | 0.79 | 0.87 | $0.61 | 
| 29 | Claude Sonnet 4.5 | **0.596** | 0.37 | 0.25 | 0.67 | 0.82 | 0.88 | $1.83 | 
| 30 | GPT-5.4 nano | **0.561** | 0.47 | 0.27 | 0.53 | 0.71 | 0.83 | $0.10 | 

**Reproducibility.** The generator is seeded (seed 2026) and regenerates the identical 336 images. The grader is deterministic, with no LLM judge.

**Limitations, honestly.** The panels are synthetic and single-panel. The event vocabulary is mine ("plateau" vs "saturation"). "Recovering" mostly follows from the event type. On 5 panels two deploy labels overlap, though the marker lines stay distinct. Each model saw each panel once in the final run; earlier repeats barely moved the ranking.

*I built this with an AI coding assistant (Claude) helping with the generator, the Kaggle harness and the analysis. The benchmark design, the audit decisions and every number here were checked against the run data.*
