# Can an LLM name the WCAG criterion a snippet violates? A 6-case RGAA benchmark on Kaggle

> Source: <https://dev.to/yvoolab/can-an-llm-name-the-wcag-criterion-a-snippet-violates-a-6-case-rgaa-benchmark-on-kaggle-1cp3>
> Published: 2026-10-09 08:18:04+00:00

*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23)*

Accessibility audits against RGAA 4.1, the French profile of WCAG 2.1 AA, are a service I offer to French sites. The first pass of an audit is tedious and mechanical: look at a fragment, name the success criterion it breaks. If a model can do that pass reliably, the human hours go to the judgement calls instead.

So the benchmark asks exactly that question, six times. Each case is a small HTML fragment with one obvious defect, plus one compliant control:

| Case | Fragment | Primary criterion | 
|---|---|---|
| img-no-alt | `<img src="hero.jpg">` | 1.1.1 Non-text Content | 
| input-no-label | `<input type="text" id="q" placeholder="Rechercher">` | 1.3.1 Info and Relationships | 
| low-contrast | `<p style="color:#999;background:#fff">Mentions légales</p>` | 1.4.3 Contrast (Minimum) | 
| div-button | `<div onclick="go()">Valider</div>` | 2.1.1 Keyboard | 
| no-lang | `<html><head><title>Accueil</title></head><body>Bonjour</body></html>` | 3.1.1 Language of Page | 
| ok-link | `<a href="https://dev.to/contact">Nous contacter</a>` | NONE | 

The prompt is one paragraph: you are auditing a French public site, here is the fragment, answer with the WCAG 2.2 criterion number only, or NONE. (Every id above is unchanged between WCAG 2.1 and 2.2, so the RGAA mapping holds either way.)

Grading is deterministic, no judge model. A case passes when the answer contains any *defensible* criterion for that fragment. Two fragments genuinely break more than one rule: a clickable `div` fails both Keyboard (2.1.1) and Name, Role, Value (4.1.2); an unlabeled input fails 1.3.1, 3.3.2 Labels or Instructions, and 4.1.2. My first local draft accepted only the primary id, and the first model I ran lost two points for answers I would have accepted from a junior auditor. That draft was grading taste in criteria, not the ability to spot the defect, so I widened it before pushing to Kaggle. I kept the strict count as a second column because the gap between the two turned out to be the interesting part.

Score is cases passed out of 6. The whole task is one Python file: six cases, one prompt template, one loop.

Five models from the Kaggle model proxy, chosen to cover three labs, two price tiers, and two open-weight reasoning models:

I wanted to see whether chain-of-thought buys anything on a task where the answer is a bare id.

Results from task version 2 on Kaggle:

| Model | Lenient (any defensible criterion) | Strict (primary only) | Cost for 6 cases | Latency | 
|---|---|---|---|---|
| Claude Sonnet 5 | 6/6 | 4/6 | $0.007 | 10 s | 
| Gemini 3 Flash Preview | 6/6 | 5/6 | $0.026 | 44 s | 
| Gemini 3.7 Flash | 6/6 | 5/6 | $0.027 | 68 s | 
| gpt-oss-120b | 5/6 | 4/6 | $0.0005 | 149 s | 
| DeepSeek-R1-0528 | 6/6 * | 6/6 * | $0.059 | 65 s | 

* DeepSeek ignored "criterion number only" and answered in prose, naming every candidate criterion for the ambiguous cases. Substring grading cannot tell an answer from an enumeration, so its strict column is not comparable. More on that below.

Per case, what each model answered:

| Case | Expected | Sonnet 5 | Gemini 3 Flash | Gemini 3.7 Flash | gpt-oss-120b | DeepSeek-R1 | 
|---|---|---|---|---|---|---|
| img-no-alt | 1.1.1 | 1.1.1 | 1.1.1 | 1.1.1 | 1.1.1 | 1.1.1 | 
| input-no-label | 1.3.1 | 3.3.2 | 3.3.2 | 3.3.2 | 3.3.2 | 1.3.1, 3.3.2, 4.1.2 | 
| low-contrast | 1.4.3 | 1.4.3 | 1.4.3 | 1.4.3 | 1.4.3 | 1.4.3 | 
| div-button | 2.1.1 | 4.1.2 | 2.1.1 | 2.1.1 | 2.1.1 | 2.1.1, 4.1.2 | 
| no-lang | 3.1.1 | 3.1.1 | 3.1.1 | 3.1.1 | `3` | 3.1.1 | 
| ok-link | NONE | NONE | NONE | NONE | NONE | compliant | 

**Spotting the defect was the easy part on these six cases. Naming the canonical rule was not.** Every model found every defect and none flagged the compliant link. The single lenient miss in the table is not a knowledge miss: gpt-oss-120b answered `3` on the missing-lang page, a truncated `3.1.1` after 225 tokens of reasoning. The only real disagreement is *which* criterion to cite for the two multi-rule fragments. Every model that gave a single answer picked 3.3.2 over 1.3.1 for the unlabeled input, and Sonnet picked 4.1.2 over 2.1.1 for the `div` button.

**The pick is not stable within one model.** I ran the task twice on Kaggle, 21 minutes apart; the two versions differ only in how the score is returned, not in the prompt. Sonnet answered 1.3.1 for the input the first time and 3.3.2 the second. Gemini 3 Flash Preview answered 4.1.2 for the `div` the first time and 2.1.1 the second. If you grade on exact id, you are measuring sampling noise.

**The control case held.** Every run on the clean `<a>`, across both task versions and all five models, said it was compliant. For a first-pass audit tool, the false-positive rate on obviously good markup matters more than the hit rate on obviously bad markup, because false positives are what make a human stop trusting the tool.

**Format compliance is its own axis, and the reasoning models lost it.** DeepSeek-R1 got every case right and still broke the contract: it answered in paragraphs, listed all three candidate criteria for the input, and its visible output includes the full `<think>` block. gpt-oss-120b kept the format but truncated an answer. The three proprietary models returned exactly the bare id asked for, every time. For anything you plan to pipe into a script, that is the number that decides.

**Reasoning tokens are where the money goes.** Sonnet answered the six cases in 10 seconds for under a cent. The two Gemini Flash models took 4 to 7 times longer and cost 4 times more, almost all of it output tokens spent thinking before emitting a five-character id. DeepSeek-R1 was the most expensive run at $0.059, eight times Sonnet, for the same six verdicts. gpt-oss-120b was the cheapest at $0.0005, but one case alone sat on the proxy for two minutes, and the first run was killed by a 429 before it finished.

What surprised me: I expected the cheap models to miss the contrast case, since judging `#999` on white needs an actual ratio (2.85:1, below the 4.5:1 floor). Nobody missed it.

What I would measure next: fragments with *two* defects where the model has to list both; the RGAA numbering (criterion 1.1, 11.1, and so on) instead of WCAG ids, since French audits are delivered in that vocabulary; and real pages with axe-core as the ground truth, which is where a six-case set stops being enough.

[https://www.kaggle.com/benchmarks/tasks/delphine53303/rgaa-a11y-judgement](https://www.kaggle.com/benchmarks/tasks/delphine53303/rgaa-a11y-judgement)

Task source, cover and raw run files: [https://github.com/yvoolab/rgaa-a11y-benchmark](https://github.com/yvoolab/rgaa-a11y-benchmark)
