Show HN: LayoutLens: AI-Powered Visual UI Testing LayoutLens, an AI-powered visual UI testing tool, catches layout and accessibility bugs using deterministic axe-core and geometry checks that run keyless and free in CI, with an optional vision-LLM tier that the deterministic layer can overrule. The tool's LLM tier measures 81.1% accuracy on its bundled benchmark (60/74 labeled queries with gpt-4o-mini, measured 2026-07-21). LayoutLens is available via pip and supports three tiers: deterministic, hybrid, and LLM-only. LayoutLens catches the layout and accessibility bugs your pixel baseline can't see and your LLM can't be trusted about — deterministic axe-core and geometry checks that run keyless and free in CI, with an optional vision-LLM tier that the deterministic layer is allowed to overrule. Three tiers, use what you need: | Tier | What runs | Needs | Reliability | |---|---|---|---| Deterministic | axe-core WCAG A/AA + geometry/contrast/occlusion scorers | no API key or model call | measured facts, reproducible; this is the CI gate | Hybrid default | deterministic scan grounds the vision LLM; measured violations force the verdict | an LLM API key | precision-preserving: the model can add findings, never erase measured ones | LLM | natural-language questions answered from a screenshot | an LLM API key any LiteLLM provider, incl. Ollama/vLLM via api base | honest numbers below | Keyless, deterministic — safe as a required check on any fork result = await lens.check accessibility "page.html", mode="axe" result = await lens.check layout "page.html", viewport="mobile", mode="deterministic" Natural-language, grounded by the deterministic scan hybrid result = await lens.analyze "https://example.com", "Is the navigation user-friendly?" Or from pytest — the deterministic assertions need no key, and assert ui skips never fails without one: python def test landing page layoutlens : layoutlens.assert a11y "landing.html" axe, keyless layoutlens.assert layout "landing.html", viewport="mobile" keyless layoutlens.assert ui "landing.html", "Is the CTA above the fold?" Honest numbers: the LLM tier measures 81.1% on the bundled benchmark 60/74 labeled queries, gpt-4o-mini , measured 2026-07-21 — artifact /gojiplus/layoutlens/blob/main/benchmarks/results/2026-07-21 gpt-4o-mini.json . See Limitations limitations for what vision models can and cannot reliably judge — the deterministic tier exists precisely because of those limits. pip install layoutlens playwright install chromium For screenshot capture LayoutLens's API is async — run it with asyncio.run ... , or await it directly if you're already inside an async def e.g. pytest-asyncio, FastAPI, a notebook cell . Every snippet below assumes one of those two contexts; only the first one spells out the asyncio.run ... wrapper. python import asyncio from layoutlens import LayoutLens async def main : Initialize uses OPENAI API KEY env var lens = LayoutLens Test any website or local HTML result = await lens.analyze "https://your-site.com", "Is the header properly aligned?" print f"Answer: {result.answer}" print f"Confidence: {result.confidence:.1%}" asyncio.run main That's it No selectors, no complex setup, just natural language questions. LayoutLens vendors axe-core https://github.com/dequelabs/axe-core 4.10.3 and runs it against a real Playwright-rendered page to catch actual WCAG 2.1 A/AA violations — not an LLM guess. This mode is fully keyless: no OPENAI API KEY , no network call to an AI provider, just deterministic, reproducible results. Deterministic axe-core scan only — no API key needed layoutlens page.html --a11y axe Hybrid: axe-core + LLM vision, axe overrides the verdict on violations needs an API key layoutlens https://example.com --a11y hybrid Legacy vision-only accessibility check needs an API key layoutlens page.html --a11y llm --a11y requires one of hybrid / axe / llm and is mutually exclusive with --query — accessibility mode always uses the built-in WCAG checks instead of a free-form question. python from layoutlens import LayoutLens, AxeAuditor Raw axe-core report — no LayoutLens instance or API key needed at all report = await AxeAuditor .audit "page.html" print report.summary print report.ok True if there are zero violations print report.violations list A11yFinding : rule id, impact, wcag refs, nodes, ... Via the LayoutLens API, restricted to WCAG A/AA tags, still keyless lens = LayoutLens no API key required at construction result = await lens.check accessibility "page.html", mode="axe" print result.answer "Yes — axe-core found no WCAG A/AA violations" or lists violated rules — deterministic axe-core only. No API key, no LLM call. mode="axe" confidence is always 1.0 . default for mode="hybrid" check accessibility / check accessibility — runs axe-core and the LLM vision analysis, injecting the axe findings into the LLM's prompt as grounding context. If axe finds any violation, the final verdict is deterministically forced to "no" confidence 1.0 , regardless of what the LLM says — axe overrides the model, not the other way around. If axe finds nothing, the LLM's own answer/confidence are kept it can still flag issues axe's automated rules can't catch, like poor color choices that pass contrast math or confusing visual hierarchy .— legacy vision-only analysis, no axe-core involved. Requires an API key. mode="llm" Hybrid: axe grounds the LLM and can force the verdict result = await lens.check accessibility "page.html", mode="hybrid" print result.metadata "a11y" full axe report dict print result.metadata "engine" "axe-core 4.10.3" Alongside axe-core, LayoutLens ships LayoutScorer — a keyless, LLM-free detector for geometric and contrast defects, measured directly off the rendered page with the browser's own layout engine. Foundational contrast and geometry measurements were ported from UIJudgeBench https://github.com/gojiplus/uijudge-bench ; newer WCAG and text-occlusion checks are independent LayoutLens implementations evaluated by that benchmark. It finds: contrast — text below the WCAG AA ratio 4.5:1 normal, 3.0:1 large , with the measured ratio overlap — sibling elements whose bounding boxes collide clipping — content cut off by a fixed-size box with hidden overflow viewport-protrusion — elements extending past the viewport width horizontal-scroll bugs target-size — undersized targets that also fail the machine-measurable WCAG 2.5.8 spacing, inline, and unmodified user-agent-control exceptions focus-obscured — keyboard-focused components entirely hidden by author DOM content the automatable geometric core of WCAG 2.4.11 text-occlusion — rendered text, including chart labels, covered by another painted DOM element; this is a visual-quality finding, not a WCAG criterion python from layoutlens.layout import LayoutScorer, contrast ratio, read computed styles Scan a page — no LayoutLens instance, no API key, deterministic. report = await LayoutScorer .scan "page.html", viewport="mobile" print report.ok True if no defects found print report.summary findings grouped by class, with measured receipts for f in report.findings: print f.defect class, f.selector, f.measured each finding carries the numbers behind it Or use the pure WCAG contrast math directly no browser : contrast ratio 0x76, 0x76, 0x76 , 0xFF, 0xFF, 0xFF - 4.54 Every finding is a receipt: the offending selector, its bounding box, the measured value, and the threshold it violated. scan viewport=... re-runs the geometry at any viewport, so protrusion/overlap that only appear on mobile are caught. Automated findings are not a site-wide WCAG conformance claim. In particular, target-size equivalent/essential exceptions and focus-obscuration interaction-history exceptions remain explicit manual-review fields. Installing layoutlens registers a pytest plugin entry point layoutlens . The layoutlens fixture gives you three assertions: python def test checkout layoutlens : layoutlens.assert a11y "checkout.html" keyless axe gate layoutlens.assert layout "checkout.html", viewport="mobile" keyless geometry gate layoutlens.assert ui "checkout.html", "Is the pay button the most prominent element?" assert a11y / assert layout are keyless and deterministic — they run on every fork and PR with no secrets, and failure messages carry the rule id, selector, and measured numbers. assert ui vision LLM skips instead of failing when no API key is configured, or always with --layoutlens-no-llm — so one suite serves both the free deterministic lane and the LLM lane. --layoutlens-model picks the model for assert ui . layoutlens-mcp exposes the checks as MCP https://modelcontextprotocol.io tools for Claude Code, Cursor, and friends: pip install "layoutlens mcp " register the stdio server in your agent config: command: layoutlens-mcp Tools: audit accessibility and scan layout keyless, deterministic — they return measured numbers, not model opinions , in compact summaries of a few hundred tokens , plus check ui and compare ui vision LLM . The deterministic tools cover visual facts accessibility-tree snapshots cannot see: contrast, geometry, target spacing, complete focus obscuration, and text occlusion such as a chart line painted over its label. Both deterministic engines emit SARIF 2.1.0 https://sarifweb.azurewebsites.net/ : layoutlens page.html --layout deterministic --output sarif layout.sarif layoutlens page.html --a11y axe --output sarif a11y.sarif Upload with github/codeql-action/upload-sarif and findings appear as PR annotations with stable rule ids layout/page-overflow , axe/color-contrast , ... tracked over time — keyless, so it works on every fork. Or use the packaged action — gojiplus/layoutlens-action https://github.com/gojiplus/layoutlens-action — which bundles install, scan, job summary, PR annotations, a sticky results comment, and the SARIF upload into one step: - uses: gojiplus/layoutlens-action@v1 with: sources: "dist/ .html" Test single pages with custom questions: Test local HTML files result = await lens.analyze "checkout.html", "Is the payment form user-friendly?" Test with expert context from layoutlens.prompts import Instructions, UserContext instructions = Instructions expert persona="conversion expert", user context=UserContext business goals= "reduce cart abandonment" , target audience="mobile shoppers" , result = await lens.analyze "checkout.html", "How can we optimize this checkout flow?", instructions=instructions, Perfect for A/B testing and redesign validation. compare accepts URLs, local HTML files, or screenshot images — every source is rendered and every screenshot is sent to the model: result = await lens.compare "https://old-design.example.com", "https://new-design.example.com" , "Which design is more accessible?", print f"Winner: {result.answer}" Domain expert knowledge with one line of code: Professional accessibility audit WCAG expert result = await lens.check accessibility "product-page.html", compliance level="AA" Conversion rate optimization CRO expert result = await lens.optimize conversions "landing.html", business goals= "increase signups" , industry="saas" Mobile UX analysis Mobile expert result = await lens.analyze mobile ux "app.html", performance focus=True E-commerce audit Retail expert result = await lens.audit ecommerce "checkout.html", page type="checkout" Legacy methods still work result = await lens.check accessibility "product-page.html" Backward compatible analyze handles single or multiple sources/queries — pass lists to either source or query and it fans out every combination concurrently: results = await lens.analyze source= "home.html", "about.html", "contact.html" , query= "Is it accessible?", "Is it mobile-friendly?" , Returns a BatchResult; processes 6 combinations concurrently print f"{results.successful queries}/{results.total queries} succeeded" Cap concurrent API calls with max concurrent result = await lens.analyze source= "page1.html", "page2.html", "page3.html" , query="Is it accessible?", max concurrent=5, All results provide clean, typed JSON for automation: result = await lens.analyze "page.html", "Is it accessible?" Export to clean JSON json data = result.to json Returns typed JSON string print json data { "source": "page.html", "query": "Is it accessible?", "answer": "Yes, the page follows accessibility standards...", "confidence": 0.85, "reasoning": "The page has proper heading structure...", "screenshot path": "/path/to/screenshot.png", "viewport": "desktop", "timestamp": "2024-01-15 10:30:00", "execution time": 2.3, "metadata": {} } Type-safe structured access from layoutlens.types import AnalysisResultJSON import json data: AnalysisResultJSON = json.loads result.to json confidence = data "confidence" Fully typed: float Choose from 6 built-in domain experts with specialized knowledge: Available experts: accessibility expert, conversion expert, mobile expert, ecommerce expert, healthcare expert, finance expert Use any expert with custom analysis result = await lens.analyze with expert source="healthcare-portal.html", query="How can we improve patient experience?", expert persona="healthcare expert", focus areas= "patient privacy", "health literacy" , user context={ "target audience": "elderly patients", "accessibility needs": "large text", "simple navigation" , "industry": "healthcare", }, Expert comparison analysis URLs, local HTML files, or screenshots result = await lens.compare with expert sources= "https://old.example.com", "https://new.example.com" , query="Which design converts better?", expert persona="conversion expert", focus areas= "cta prominence", "trust signals" , Test suites are declared in YAML/JSON and loaded into a UITestSuite . Breaking change v1.7.0 : every test case must declare expected results — an answer "yes"/"no", matched against the parsed leading yes/no token of the analysis answer and/or a contains list terms that must appear, case-insensitively, in the answer + reasoning . A case with no expected results now raises ValidationError at load time instead of silently grading on confidence alone. test suite.yaml name: "Homepage Suite" description: "Accessibility and layout checks" test cases: - name: "Navigation Alignment" html path: "pages/home.html" queries: - "Is the navigation menu properly centered?" viewports: "desktop" expected results: answer: "yes" contains: "centered" expected confidence: 0.7 optional, defaults to 0.7 python import yaml from layoutlens import LayoutLens, UITestSuite with open "test suite.yaml" as f: suite = UITestSuite.from dict yaml.safe load f lens = LayoutLens results = await lens.run test suite suite list UITestResult , one per test case for r in results: print f"{r.test case name}: {r.passed tests}/{r.total tests} passed" print r.to json includes per-assertion "assertion detail" There is no CLI subcommand for suites — run test suite is a Python API only. See examples/sample test suite.yaml /gojiplus/layoutlens/blob/main/examples/sample test suite.yaml for a complete, runnable example. For external evaluation harnesses e.g. UIJudgeBench , judge sends your prompt verbatim — no persona, no scaffolding, no appended JSON contract — alongside a single image, and returns a parsed, structured verdict. Your harness owns the entire prompt, including its own response contract and prompt versioning. python from layoutlens import LayoutLens lens = LayoutLens model="gpt-4o" or any vision model via provider/api base prompt = "You are a UI evaluation judge. Compare the layout in the image against the " "criteria below and respond ONLY as JSON: " '{"answer": "A" | "B", "confidence": 0.0-1.0, "rationale": "..."}.\n' "Criteria: which layout has clearer visual hierarchy?" result = await lens.judge "candidate.png", prompt, max tokens=300 result.answer parsed "answer" field, or "unknown" if unparseable result.confidence parsed 0-1, else 0.0 result.rationale parsed "rationale"/"reasoning", else "" result.raw full raw model text result.refused True if the model declined result.usage {"prompt tokens": ..., "completion tokens": ..., "total tokens": ...} result.parse mode "json" | "fallback" | "none" For bulk evaluation, judge batch uses provider-native asynchronous Batch APIs. Native OpenAI uses the official Responses Batch API, gemini/ models use the Google Gen AI inline Batch API, and other supported providers use LiteLLM's file-based Batch API. For example, a localization benchmark can preserve the input coordinate frame and explicitly cap reasoning: python from layoutlens import BatchRequest, LayoutLens lens = LayoutLens provider="openai", model="gpt-5.6-luna" results = await lens.judge batch BatchRequest "item-1", "target.jpg", prompt , max tokens=256, reasoning effort="low", image detail="original", Resume manifests are content-addressed by the exact prompts, images, model, backend, endpoint, token budget, reasoning effort, and image detail, so a changed request cannot reuse a stale response. A per-manifest lock prevents two processes from submitting the same exact batch concurrently. Manifests created before 2.1.1 fail closed with explicit migration details because they cannot attest their original prompts, images, or token budget. Changing an input creates a new fingerprint; if any prior same-model manifest records an overlapping submitted id, resume fails closed until the user explicitly migrates the job or authorizes a fresh billed run. An ungraceful process stop can leave a .json.lock file: confirm no matching run is active, then remove only that lock file to resume from the preserved manifest. Key guarantees: - Verbatim prompt — LayoutLens adds nothing to the text you provide. - No caching — every judge call hits the model, so a benchmark controls its own determinism. - Per-model parameter policy — models that reject non-default sampling params Claude Sonnet 5, Opus 4.6+ omit temperature automatically; others send temperature=0.0 . - Self-hosted endpoints — point at Ollama/vLLM via api base : lens = LayoutLens provider="litellm", model="ollama/qwen2.5vl", api base="http://localhost:11434", Analyze a single page layoutlens https://example.com "Is this accessible?" Analyze local files layoutlens page.html "Is the design professional?" Compare two designs URLs, local HTML files, or screenshot images layoutlens https://old.example.com https://new.example.com --compare Analyze with different viewport layoutlens site.com "Is it mobile-friendly?" --viewport mobile JSON output for automation layoutlens page.html "Is it accessible?" --output json Deterministic WCAG accessibility scan — no API key required see "Deterministic Accessibility Checks" above for hybrid/llm modes layoutlens page.html --a11y axe Choose model / pass an API key explicitly layoutlens page.html "Is it accessible?" --model gpt-4o --api-key sk-... Run layoutlens with no arguments or --help to see the full flag reference: --query/-q , --compare/-c , --viewport/-v {desktop,mobile,tablet} , --output/-o {text,json} , --api-key , --model/-m , --a11y {hybrid,axe,llm} . - name: Visual UI Test run: | pip install layoutlens playwright install chromium layoutlens ${{ env.PREVIEW URL }} "Is it accessible and mobile-friendly?" python import pytest from layoutlens import LayoutLens @pytest.mark.asyncio async def test homepage quality : lens = LayoutLens result = await lens.analyze "homepage.html", "Is this production-ready?" assert result.confidence 0.8 assert "yes" in result.answer.lower LayoutLens bundles a compact benchmark suite 18 fixtures / 74 labeled queries for smoke-testing AI performance. For a larger, paper-rigor benchmark of AI judges of web UI quality — 4,000+ machine-verified items across accessibility, layout, and referring tasks, built on LayoutLens's own axe/browser machinery — see UIJudgeBench dataset on Hugging Face https://huggingface.co/datasets/gojiberries/uijudge-bench . LayoutLens is a planned judge baseline there. Run LayoutLens against test data python benchmarks/run benchmark.py --api-key sk-your-key With custom settings python benchmarks/run benchmark.py \ --api-key sk-your-key \ --output benchmarks/my results \ --no-batch \ --filename custom results.json Evaluate results against ground truth python benchmarks/evaluation/evaluator.py \ --answer-keys benchmarks/answer keys \ --results benchmarks/layoutlens output \ --output evaluation report.json The evaluator scores every answer deterministically leading yes/no token vs the answer key; ambiguous answers count as incorrect and writes an artifact with per-category and overall accuracy. The committed benchmarks/results/2026-07-21 gpt-4o-mini.json /gojiplus/layoutlens/blob/main/benchmarks/results/2026-07-21 gpt-4o-mini.json is a real measured run: { "evaluation summary": { "date": "2026-07-21", "model": "gpt-4o-mini", "total queries": 74, "total correct": 60, "ambiguous answers": 7, "overall accuracy": 0.811, "evaluator version": "2.0", "evaluator method": "Deterministic structured yes/no; ambiguous answers count as incorrect." }, "category results": { "responsive design": {"total queries": 21, "correct predictions": 20, "accuracy": 0.952}, "layout alignment": {"total queries": 24, "correct predictions": 19, "accuracy": 0.792}, "accessibility": {"total queries": 21, "correct predictions": 16, "accuracy": 0.762}, "ui components": {"total queries": 8, "correct predictions": 5, "accuracy": 0.625} } } Create your own test data and answer keys: python Use the async API for custom benchmark workflows from layoutlens import LayoutLens async def run custom benchmark : lens = LayoutLens test cases = {"source": "page1.html", "query": "Is it accessible?"}, {"source": "page2.html", "query": "Is it mobile-friendly?"}, results = for case in test cases: result = await lens.analyze case "source" , case "query" results.append { "test": case, "result": result.to json , Clean JSON output "passed": result.confidence 0.7, } return results Simple configuration options: Via environment export OPENAI API KEY="sk-..." Via code lens = LayoutLens api key="sk-...", model="gpt-4o-mini", or "gpt-4o" for higher accuracy cache enabled=True, Reduce API costs cache type="memory", "memory" or "file" Calibrate your trust to the tier you use: Vision LLMs miss fine-grained UI differences. On DiffSpot arXiv 2605.29615 https://arxiv.org/abs/2605.29615 , a 2026 benchmark of fine-grained web-UI changes, the best frontier model scored 47.2% overall and under 23% recall on the hard tier ; open models hallucinated differences on 18–24% of identical pairs. Do not use the LLM tier as a sole gate for subtle visual regressions — that is what the deterministic scorers are for. Passing axe-core is not WCAG conformance. Automated rules cover only a subset of WCAG; Microsoft's a11y LLM evaluation makes the same disclaimer for its own checks. axe passing means "no automated rule failed", not "accessible". The deterministic scorers measure rendered facts, not full intent. The WCAG 2.5.8 spacing, inline, and unmodified-user-agent-control exceptions are modeled. Equivalent-control and essential-presentation exceptions still require review, as do interaction-history cases under WCAG 2.4.11. General text occlusion is a visual-quality signal, not a WCAG conformance claim. Findings carry their measured numbers so you can judge. Our own benchmark is small 74 labeled queries and easier than DiffSpot-class tasks; the 81.1% figure is honest but narrow. The harness is model-agnostic benchmarks/run benchmark.py --model ... — re-run it rather than trusting ours. - 📖 - Comprehensive guides and API reference Full Documentation https://gojiplus.github.io/layoutlens/ - 🎯 - Real-world usage patterns Examples https://github.com/gojiplus/layoutlens/tree/main/examples - 🐛 - Report bugs, request features, get help Issues https://github.com/gojiplus/layoutlens/issues Natural Language - Write tests like you'd describe the UI to a colleague Domain Expert Knowledge - Built-in expertise in accessibility, CRO, mobile UX, and more Rich Context Support - Business goals, user personas, compliance standards, and technical constraints Zero Selectors - No more fragile XPath or CSS selectors Visual Understanding - AI sees what users see, not just code Async-by-Default - Concurrent processing for optimal performance Simple API - One analyze method handles single pages, batches, and comparisons Structured JSON Output - TypedDict schemas for full type safety in automation Honest Benchmarking - Compact built-in suite 81.1% measured accuracy, gpt-4o-mini, 74 queries ; see UIJudgeBench https://github.com/gojiplus/uijudge-bench for the full-scale external benchmark Deterministic Accessibility - Vendored axe-core WCAG 2.1 A/AA checks, no API key or LLM variance Making UI testing as simple as asking "Does this look right?"