{"slug": "benchmarking-frontier-ai-on-curved-tire-sidewall-ocr", "title": "Benchmarking Frontier AI on Curved Tire Sidewall OCR", "summary": "A developer built a Kaggle benchmark testing 11 frontier vision-language models on zero-shot OCR of ISO metric tire size codes from 51 curved tire sidewall images across 22 dimensions. Google Gemini 3.8 Flash and Anthropic Claude Opus 5.5 tied for first at 46/51 exact matches (90.20%), while Alibaba Qwen 3 Next 80B Thinking (5.88%), Zhipu AI GLM-5 (1.96%) and OpenAI gpt-oss-20b (0%) failed largely by violating the strict 9-character output constraint. The benchmark also found that naive horizontal X-coordinate sorting of YOLO character boxes reverses curved tire codes, such as reading 225/60R18 as 81R06/522.", "body_md": "*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*\n\nIn computer vision pipelines for automotive manufacturing, fleet maintenance, and tire recycling, reading tire specifications is a notorious challenge.\n\nTire sidewalls feature embossed, black-on-black rubber text following the wheel's circular arc. When training custom object detection models (such as YOLO) to detect individual character boxes (`/`, `0-9`, `R`), traditional post-processing relies on sorting bounding boxes by horizontal X-coordinates to reconstruct the tire code (e.g., `205/55R16`). \n\nHowever, this naive approach fails on real-world tires:\n\n`225/60R18` is read as `81R06/522`).\nCan frontier general-purpose Vision-Language Models (VLMs) perform **zero-shot OCR and structured extraction** of ISO metric tire size specifications (`[Width]/[Aspect]R[Rim]`) directly from raw tire sidewall photos without any fine-tuning, and act as automated auditors for human-annotated YOLO datasets?\n\nTo measure this, I created a curated test suite of **51 tire sidewall images** across **22 distinct tire dimensions** (covering standard horizontal text, steep curved arcs, and rotated positions), backed by verified ground-truth labels.\n\nThe live Kaggle Benchmark evaluates **11 frontier models** across 5 AI providers and open-weight architectures:\n\n*(Note: During task exploration, Gemini 3.7 Flash [46/51], Gemini 3 Flash Preview [45/51], and Claude Sonnet 5 [44/51] were also evaluated on the underlying task. DeepSeek-R1 was correctly rejected by the benchmark runner as a text-only reasoning model lacking a visual encoder, and Grok 4.6 encountered upstream proxy connection timeouts).*\n\nEach model was prompted with an isolated context per image:\n\n*\"Inspect the tire sidewall in this image carefully. Locate the standard ISO metric tire size specification (format: WWW/AARDI, e.g. 205/55R16, 225/60R18, 235/65R17). Return ONLY the exact 9-character tire size code.\"*\n\nThe benchmark recorded:\n\nThe table below reflects the official leaderboard from the live Kaggle Benchmark:\n\n| Rank | Model | Provider | Exact Matches | Exact Accuracy | Key Behavior | \n|---|---|---|---|---|---|\n| **#1** | **Google Gemini 3.8 Flash** |  | **46 / 51** | **90.20%** | Tied #1; top multimodal speed & curved text accuracy | \n| **#1** | **Anthropic Claude Opus 5.5** | Anthropic | **46 / 51** | **90.20%** | Tied #1; flagship frontier reasoning & flawless ISO extraction | \n| **#3** | **Google Gemini 3.5 Flash-Lite** |  | **45 / 51** | **88.24%** | Exceptional efficiency and accuracy for a lightweight model | \n| **#4** | **Google Gemini 2.5 Pro** |  | **44 / 51** | **86.27%** | Deep reasoning; minor slip on worn edge numerals | \n| **#5** | **Google Gemma 4 31B** | Google (OSS) | **41 / 51** | **80.39%** | Best open-weights model; matched Claude 5.5 models | \n| **#5** | **Anthropic Claude Sonnet 5.5** | Anthropic | **41 / 51** | **80.39%** | Solid reasoning; missed low-contrast rim digits | \n| **#5** | **Anthropic Claude Haiku 5.5** | Anthropic | **41 / 51** | **80.39%** | Tied with Sonnet 5.5 at a fraction of latency | \n| **#8** | **OpenAI GPT-6.1 Sol** | OpenAI | **39 / 51** | **76.47%** | Tended to confuse scuffed numbers ( `3` vs`2` ,`4` vs`6` ) | \n| **#9** | **Alibaba Qwen 3 Next 80B Thinking** | Alibaba | **3 / 51** | **5.88%** | Overthought reasoning; violated strict 9-char constraint | \n| **#10** | **Zhipu AI GLM-5** | Zhipu AI | **1 / 51** | **1.96%** | Low parsing accuracy on low-contrast embossments | \n| **#11** | **OpenAI gpt-oss-20b** | OpenAI (OSS) | **0 / 51** | **0.00%** | Failed strict 9-character formatting constraint | \n\n**Note on Kaggle UI Score Display**: On the Kaggle Benchmarks leaderboard UI, raw numeric scores appear with an automatic × 100% display badge (e.g. `46.00` renders as `4600.0%` and `41.00` renders as `4100.0%`, representing 46 and 41 correct items out of 51 total samples, or 90.20% and 80.39%). The table above provides the normalized, exact percentage rates.\n\nWhen deploying vision models in production—such as scanning thousands of tire sidewalls daily across an automotive fleet or recycling facility—raw accuracy must be balanced against inference cost. Kaggle Benchmarks plotted the **Score vs. Total Cost Pareto Frontier** across all evaluated models:\n\nIn our YOLO dataset, human annotations on rotated or curved tire arcs were often scrambled by bounding-box sorting algorithms (e.g. producing `81R06/522` or `71R05/522`). \n\n**All top-tier frontier VLMs effortlessly read these rotated, curved tires in the correct semantic human reading order**, outputting `225/60R18` and `225/50R17` flawlessly. This demonstrates that multimodal LLMs perceive continuous text streams gestalt-style rather than through brittle geometric axis projections.\n\nGoogle's **Gemini 3.8 Flash** and Anthropic's flagship **Claude Opus 5.5** shared top honors with identical scores of **46 / 51 (90.20%)**. While Claude Opus 5.5 demonstrated immense precision in distinguishing faint embossed characters, Gemini 3.8 Flash delivered that same accuracy with lightning-fast inference times.\n\nRight behind them, Google's lightweight **Gemini 3.5 Flash-Lite** scored **45 / 51 (88.24%)**, outperforming heavyweights like Gemini 2.5 Pro (44 / 51) and GPT-6.1 Sol (39 / 51)—proving that specialized vision distillation can outperform pure model parameter scale.\n\nOne of the most exciting results is **Gemma 4 31B** scoring **80.39% (41 / 51)**. It matched proprietary frontier flagships **Claude Sonnet 5.5** and **Claude Haiku 5.5**, while beating **OpenAI GPT-6.1 Sol (76.47%)**. For manufacturing plants and tire depots requiring local, on-premise execution without cloud API dependency, Gemma 4 31B offers production-grade zero-shot OCR capability.\n\nAnthropic's **Claude Haiku 5.5** tied **Claude Sonnet 5.5** exactly at **41 / 51 (80.39%)**. Getting flagship-grade multimodal OCR accuracy at Haiku's speed and cost profile makes it an exceptional candidate for high-throughput batch auditing.\n\nA fascinating divergence occurred with **Qwen 3 Next 80B Thinking** (3 / 51, 5.88%):\n\nWhile reasoning and \"thinking\" architectures excel at multi-step mathematics and coding, they frequently failed this pure extraction benchmark because their outputs included internal reasoning preambles, chain-of-thought commentary, or Markdown wrappers (e.g.,\n\n```` ``` text\\n205/55R16\\n``` ````\n\n) rather than adhering strictly to the required bare 9-character string.\n\nAcross the ~10% of cases where top models failed, the errors were **not** random hallucinations:\n\n`235/45R18`), models predicted `225/45R18` (confusing a scuffed embossed `3` with a `2`).` 245/50R20`), models predicted `265/50R20` (confusing `4` with `6`).` 255/60R18`), models misread the width and rim on a heavily worn sidewall.\nNotice that in almost every failure case, the **Aspect Ratio** and **Construction ('R')** remained 100% correct. The error was almost exclusively an off-by-ten millimeter width confusion caused by low-contrast black rubber embossments.\n\nDuring benchmark preparation, we discovered a human labeling typo in our original dataset where an annotator typed `3` instead of `R` (`255/60318`). During inference, the models naturally output `255/60R18` based on semantic domain awareness of tire codes, proving that VLMs can serve as effective automated validation auditors for industrial labeling pipelines.\n\nYou can inspect the full benchmark, code, and live evaluation runs directly on Kaggle:", "url": "https://wpnews.pro/news/benchmarking-frontier-ai-on-curved-tire-sidewall-ocr", "canonical_source": "https://dev.to/inushathathsara/benchmarking-frontier-ai-on-curved-tire-sidewall-ocr-145c", "published_at": "2026-10-11 03:08:01+00:00", "updated_at": "2026-10-11 03:19:45.731433+00:00", "lang": "en", "topics": ["computer-vision", "large-language-models", "ai-research", "ai-tools"], "entities": ["Kaggle", "Google", "Gemini 3.8 Flash", "Anthropic", "Claude Opus 5.5", "OpenAI", "Alibaba", "Zhipu AI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/benchmarking-frontier-ai-on-curved-tire-sidewall-ocr", "markdown": "https://wpnews.pro/news/benchmarking-frontier-ai-on-curved-tire-sidewall-ocr.md", "text": "https://wpnews.pro/news/benchmarking-frontier-ai-on-curved-tire-sidewall-ocr.txt", "jsonld": "https://wpnews.pro/news/benchmarking-frontier-ai-on-curved-tire-sidewall-ocr.jsonld"}}