Can an LLM name the WCAG criterion a snippet violates? A 6-case RGAA benchmark on Kaggle A developer built a six-case Kaggle benchmark testing whether five LLMs can name the WCAG 2.2 success criterion violated by small HTML fragments, using deterministic substring grading with both lenient (any defensible criterion) and strict (primary criterion only) scoring. All five models spotted every defect and none flagged the compliant control, but naming the canonical rule was harder: Claude Sonnet 5 scored 6/6 lenient and 4/6 strict, Gemini 3 Flash and Gemini 3.7 Flash each scored 6/6 lenient and 5/6 strict, gpt-oss-120b scored 5/6 lenient, and DeepSeek-R1 answered in prose enumerating candidate criteria, making its strict score incomparable. This is a submission for the Kaggle Benchmarking Challenge https://dev.to/challenges/kaggle-2026-09-23 Accessibility audits against RGAA 4.1, the French profile of WCAG 2.1 AA, are a service I offer to French sites. The first pass of an audit is tedious and mechanical: look at a fragment, name the success criterion it breaks. If a model can do that pass reliably, the human hours go to the judgement calls instead. So the benchmark asks exactly that question, six times. Each case is a small HTML fragment with one obvious defect, plus one compliant control: | Case | Fragment | Primary criterion | |---|---|---| | img-no-alt |