cd /news/developer-tools/show-hn-layoutlens-ai-powered-visual… Β· home β€Ί topics β€Ί developer-tools β€Ί article
[ARTICLE Β· art-107513] src=github.com β†— pub= topic=developer-tools verified=true sentiment=↑ positive

Show HN: LayoutLens: AI-Powered Visual UI Testing

LayoutLens, an AI-powered visual UI testing tool, catches layout and accessibility bugs using deterministic axe-core and geometry checks that run keyless and free in CI, with an optional vision-LLM tier that the deterministic layer can overrule. The tool's LLM tier measures 81.1% accuracy on its bundled benchmark (60/74 labeled queries with gpt-4o-mini, measured 2026-07-21). LayoutLens is available via pip and supports three tiers: deterministic, hybrid, and LLM-only.

read16 min views6 publishedAug 23, 2026
Show HN: LayoutLens: AI-Powered Visual UI Testing
Image: Michielbdejong (auto-discovered)

LayoutLens catches the layout and accessibility bugs your pixel baseline can't see and your LLM can't be trusted about β€” deterministic axe-core and geometry checks that run keyless and free in CI, with an optional vision-LLM tier that the deterministic layer is allowed to overrule.

Three tiers, use what you need:

Tier What runs Needs Reliability
Deterministic
axe-core WCAG A/AA + geometry/contrast/occlusion scorers no API key or model call measured facts, reproducible; this is the CI gate
Hybrid (default)
deterministic scan grounds the vision LLM; measured violations force the verdict
an LLM API key precision-preserving: the model can add findings, never erase measured ones
LLM
natural-language questions answered from a screenshot an LLM API key (any LiteLLM provider, incl. Ollama/vLLM via api_base )
honest numbers below
result = await lens.check_accessibility("page.html", mode="axe")
result = await lens.check_layout("page.html", viewport="mobile", mode="deterministic")

result = await lens.analyze("https://example.com", "Is the navigation user-friendly?")

Or from pytest β€” the deterministic assertions need no key, and assert_ui

skips (never fails) without one:

def test_landing_page(layoutlens):
    layoutlens.assert_a11y("landing.html")  # axe, keyless
    layoutlens.assert_layout("landing.html", viewport="mobile")  # keyless
    layoutlens.assert_ui("landing.html", "Is the CTA above the fold?")

Honest numbers: the LLM tier measures 81.1% on the bundled benchmark (60/74 labeled queries, gpt-4o-mini

, measured 2026-07-21 β€” artifact). See Limitations for what vision models can and cannot reliably judge β€” the deterministic tier exists precisely because of those limits.

pip install layoutlens
playwright install chromium  # For screenshot capture

LayoutLens's API is async β€” run it with asyncio.run(...)

, or await

it directly if you're already inside an async def

(e.g. pytest-asyncio, FastAPI, a notebook cell). Every snippet below assumes one of those two contexts; only the first one spells out the asyncio.run(...)

wrapper.

import asyncio
from layoutlens import LayoutLens

async def main():
    lens = LayoutLens()

    result = await lens.analyze(
        "https://your-site.com", "Is the header properly aligned?"
    )
    print(f"Answer: {result.answer}")
    print(f"Confidence: {result.confidence:.1%}")

asyncio.run(main())

That's it! No selectors, no complex setup, just natural language questions.

LayoutLens vendors axe-core 4.10.3 and runs it against a real Playwright-rendered page to catch actual WCAG 2.1 A/AA violations β€” not an LLM guess. This mode is fully keyless: no OPENAI_API_KEY

, no network call to an AI provider, just deterministic, reproducible results.

layoutlens page.html --a11y axe

layoutlens https://example.com --a11y hybrid

layoutlens page.html --a11y llm

--a11y

requires one of hybrid

/axe

/llm

and is mutually exclusive with --query

β€” accessibility mode always uses the built-in WCAG checks instead of a free-form question.

from layoutlens import LayoutLens, AxeAuditor

report = await AxeAuditor().audit("page.html")
print(report.summary())
print(report.ok)  # True if there are zero violations
print(report.violations)  # list[A11yFinding]: rule_id, impact, wcag_refs, nodes, ...

lens = LayoutLens()  # no API key required at construction
result = await lens.check_accessibility("page.html", mode="axe")
print(
    result.answer
)  # "Yes β€” axe-core found no WCAG A/AA violations" (or lists violated rules)

β€” deterministic axe-core only. No API key, no LLM call.mode="axe"

confidence

is always1.0

.(default formode="hybrid"

check_accessibility

/check_accessibility

) β€” runs axe-coreandthe LLM vision analysis, injecting the axe findings into the LLM's prompt as grounding context. If axe finds any violation, the final verdict is deterministically forced to "no" (confidence1.0

), regardless of what the LLM says β€” axe overrides the model, not the other way around. If axe finds nothing, the LLM's own answer/confidence are kept (it can still flag issues axe's automated rules can't catch, like poor color choices that pass contrast math or confusing visual hierarchy).β€” legacy vision-only analysis, no axe-core involved. Requires an API key.mode="llm"

result = await lens.check_accessibility("page.html", mode="hybrid")
print(result.metadata["a11y"])  # full axe report dict
print(result.metadata["engine"])  # "axe-core 4.10.3"

Alongside axe-core, LayoutLens ships LayoutScorer

β€” a keyless, LLM-free detector for geometric and contrast defects, measured directly off the rendered page with the browser's own layout engine. Foundational contrast and geometry measurements were ported from UIJudgeBench; newer WCAG and text-occlusion checks are independent LayoutLens implementations evaluated by that benchmark. It finds:

contrastβ€” text below the WCAG AA ratio (4.5:1 normal, 3.0:1 large), with the measured ratio** overlap**β€” sibling elements whose bounding boxes collide** clipping**β€” content cut off by a fixed-size box with hidden overflow** viewport-protrusion**β€” elements extending past the viewport width (horizontal-scroll bugs)** target-size**β€” undersized targets that also fail the machine-measurable WCAG 2.5.8 spacing, inline, and unmodified user-agent-control exceptionsfocus-obscuredβ€” keyboard-focused components entirely hidden by author DOM content (the automatable geometric core of WCAG 2.4.11)** text-occlusion**β€” rendered text, including chart labels, covered by another painted DOM element; this is a visual-quality finding, not a WCAG criterion

from layoutlens.layout import LayoutScorer, contrast_ratio, read_computed_styles

report = await LayoutScorer().scan("page.html", viewport="mobile")
print(report.ok)  # True if no defects found
print(report.summary())  # findings grouped by class, with measured receipts
for f in report.findings:
    print(
        f.defect_class, f.selector, f.measured
    )  # each finding carries the numbers behind it

contrast_ratio((0x76, 0x76, 0x76), (0xFF, 0xFF, 0xFF))  # -> 4.54

Every finding is a receipt: the offending selector, its bounding box, the measured value, and the threshold it violated. scan(viewport=...)

re-runs the geometry at any viewport, so protrusion/overlap that only appear on mobile are caught. Automated findings are not a site-wide WCAG conformance claim. In particular, target-size equivalent/essential exceptions and focus-obscuration interaction-history exceptions remain explicit manual-review fields.

Installing layoutlens registers a pytest plugin (entry point layoutlens

). The layoutlens

fixture gives you three assertions:

def test_checkout(layoutlens):
    layoutlens.assert_a11y("checkout.html")  # keyless axe gate
    layoutlens.assert_layout(
        "checkout.html", viewport="mobile"
    )  # keyless geometry gate
    layoutlens.assert_ui(
        "checkout.html", "Is the pay button the most prominent element?"
    )

assert_a11y

/assert_layout

arekeyless and deterministicβ€” they run on every fork and PR with no secrets, and failure messages carry the rule id, selector, and measured numbers.assert_ui

(vision LLM)skips instead of failing when no API key is configured, or always with--layoutlens-no-llm

β€” so one suite serves both the free deterministic lane and the LLM lane.--layoutlens-model

picks the model forassert_ui

.

layoutlens-mcp

exposes the checks as MCP tools for Claude Code, Cursor, and friends:

pip install "layoutlens[mcp]"

Tools: audit_accessibility

and scan_layout

(keyless, deterministic β€” they return measured numbers, not model opinions, in compact summaries of a few hundred tokens), plus check_ui

and compare_ui

(vision LLM). The deterministic tools cover visual facts accessibility-tree snapshots cannot see: contrast, geometry, target spacing, complete focus obscuration, and text occlusion such as a chart line painted over its label.

Both deterministic engines emit SARIF 2.1.0:

layoutlens page.html --layout deterministic --output sarif > layout.sarif
layoutlens page.html --a11y axe --output sarif > a11y.sarif

Upload with github/codeql-action/upload-sarif

and findings appear as PR annotations with stable rule ids (layout/page-overflow

, axe/color-contrast

, ...) tracked over time β€” keyless, so it works on every fork.

Or use the packaged action β€” gojiplus/layoutlens-action β€” which bundles install, scan, job summary, PR annotations, a sticky results comment, and the SARIF upload into one step:

- uses: gojiplus/layoutlens-action@v1
  with:
    sources: "dist/*.html"

Test single pages with custom questions:

result = await lens.analyze("checkout.html", "Is the payment form user-friendly?")

from layoutlens.prompts import Instructions, UserContext

instructions = Instructions(
    expert_persona="conversion_expert",
    user_context=UserContext(
        business_goals=["reduce_cart_abandonment"], target_audience="mobile_shoppers"
    ),
)

result = await lens.analyze(
    "checkout.html",
    "How can we optimize this checkout flow?",
    instructions=instructions,
)

Perfect for A/B testing and redesign validation. compare()

accepts URLs, local HTML files, or screenshot images β€” every source is rendered and every screenshot is sent to the model:

result = await lens.compare(
    ["https://old-design.example.com", "https://new-design.example.com"],
    "Which design is more accessible?",
)
print(f"Winner: {result.answer}")

Domain expert knowledge with one line of code:

result = await lens.check_accessibility("product-page.html", compliance_level="AA")

result = await lens.optimize_conversions(
    "landing.html", business_goals=["increase_signups"], industry="saas"
)

result = await lens.analyze_mobile_ux("app.html", performance_focus=True)

result = await lens.audit_ecommerce("checkout.html", page_type="checkout")

result = await lens.check_accessibility("product-page.html")  # Backward compatible

analyze()

handles single or multiple sources/queries β€” pass lists to either source

or query

and it fans out every combination concurrently:

results = await lens.analyze(
    source=["home.html", "about.html", "contact.html"],
    query=["Is it accessible?", "Is it mobile-friendly?"],
)
print(f"{results.successful_queries}/{results.total_queries} succeeded")
result = await lens.analyze(
    source=["page1.html", "page2.html", "page3.html"],
    query="Is it accessible?",
    max_concurrent=5,
)

All results provide clean, typed JSON for automation:

result = await lens.analyze("page.html", "Is it accessible?")

json_data = result.to_json()  # Returns typed JSON string
print(json_data)

from layoutlens.types import AnalysisResultJSON
import json

data: AnalysisResultJSON = json.loads(result.to_json())
confidence = data["confidence"]  # Fully typed: float

Choose from 6 built-in domain experts with specialized knowledge:


result = await lens.analyze_with_expert(
    source="healthcare-portal.html",
    query="How can we improve patient experience?",
    expert_persona="healthcare_expert",
    focus_areas=["patient_privacy", "health_literacy"],
    user_context={
        "target_audience": "elderly_patients",
        "accessibility_needs": ["large_text", "simple_navigation"],
        "industry": "healthcare",
    },
)

result = await lens.compare_with_expert(
    sources=["https://old.example.com", "https://new.example.com"],
    query="Which design converts better?",
    expert_persona="conversion_expert",
    focus_areas=["cta_prominence", "trust_signals"],
)

Test suites are declared in YAML/JSON and loaded into a UITestSuite

. Breaking change (v1.7.0): every test case must declare expected_results

β€” an answer

("yes"/"no", matched against the parsed leading yes/no token of the analysis answer) and/or a contains

list (terms that must appear, case-insensitively, in the answer + reasoning). A case with no expected_results

now raises ValidationError

at load time instead of silently grading on confidence alone.

name: "Homepage Suite"
description: "Accessibility and layout checks"
test_cases:
  - name: "Navigation Alignment"
    html_path: "pages/home.html"
    queries:
      - "Is the navigation menu properly centered?"
    viewports: ["desktop"]
    expected_results:
      answer: "yes"
      contains: ["centered"]
    expected_confidence: 0.7   # optional, defaults to 0.7
python
import yaml
from layoutlens import LayoutLens, UITestSuite

with open("test_suite.yaml") as f:
    suite = UITestSuite.from_dict(yaml.safe_load(f))

lens = LayoutLens()
results = await lens.run_test_suite(suite)  # list[UITestResult], one per test case
for r in results:
    print(f"{r.test_case_name}: {r.passed_tests}/{r.total_tests} passed")
    print(r.to_json())  # includes per-assertion "assertion_detail"

There is no CLI subcommand for suites β€” run_test_suite

is a Python API only. See examples/sample_test_suite.yaml for a complete, runnable example.

For external evaluation harnesses (e.g. UIJudgeBench), judge()

sends your prompt verbatim β€” no persona, no scaffolding, no appended JSON contract β€” alongside a single image, and returns a parsed, structured verdict. Your harness owns the entire prompt, including its own response contract and prompt versioning.

from layoutlens import LayoutLens

lens = LayoutLens(model="gpt-4o")  # or any vision model via provider/api_base

prompt = (
    "You are a UI evaluation judge. Compare the layout in the image against the "
    "criteria below and respond ONLY as JSON: "
    '{"answer": "A" | "B", "confidence": 0.0-1.0, "rationale": "..."}.\n'
    "Criteria: which layout has clearer visual hierarchy?"
)

result = await lens.judge("candidate.png", prompt, max_tokens=300)

result.answer  # parsed "answer" field, or "unknown" if unparseable
result.confidence  # parsed 0-1, else 0.0
result.rationale  # parsed "rationale"/"reasoning", else ""
result.raw  # full raw model text
result.refused  # True if the model declined
result.usage  # {"prompt_tokens": ..., "completion_tokens": ..., "total_tokens": ...}
result.parse_mode  # "json" | "fallback" | "none"

For bulk evaluation, judge_batch()

uses provider-native asynchronous Batch APIs. Native OpenAI uses the official Responses Batch API, gemini/*

models use the Google Gen AI inline Batch API, and other supported providers use LiteLLM's file-based Batch API. For example, a localization benchmark can preserve the input coordinate frame and explicitly cap reasoning:

from layoutlens import BatchRequest, LayoutLens

lens = LayoutLens(provider="openai", model="gpt-5.6-luna")
results = await lens.judge_batch(
    [BatchRequest("item-1", "target.jpg", prompt)],
    max_tokens=256,
    reasoning_effort="low",
    image_detail="original",
)

Resume manifests are content-addressed by the exact prompts, images, model, backend, endpoint, token budget, reasoning effort, and image detail, so a changed request cannot reuse a stale response. A per-manifest lock prevents two processes from submitting the same exact batch concurrently. Manifests created before 2.1.1 fail closed with explicit migration details because they cannot attest their original prompts, images, or token budget. Changing an input creates a new fingerprint; if any prior same-model manifest records an overlapping submitted id, resume fails closed until the user explicitly migrates the job or authorizes a fresh billed run. An ungraceful process stop can leave a .json.lock

file: confirm no matching run is active, then remove only that lock file to resume from the preserved manifest.

Key guarantees:

Verbatim promptβ€” LayoutLens adds nothing to the text you provide. - No cachingβ€” every judge call hits the model, so a benchmark controls its own determinism. - Per-model parameter policyβ€” models that reject non-default sampling params (Claude Sonnet 5, Opus 4.6+) omittemperature

automatically; others sendtemperature=0.0

. - Self-hosted endpointsβ€” point at Ollama/vLLM viaapi_base

:

lens = LayoutLens(
    provider="litellm",
    model="ollama/qwen2.5vl",
    api_base="http://localhost:11434",
)
layoutlens https://example.com "Is this accessible?"

layoutlens page.html "Is the design professional?"

layoutlens https://old.example.com https://new.example.com --compare

layoutlens site.com "Is it mobile-friendly?" --viewport mobile

layoutlens page.html "Is it accessible?" --output json

layoutlens page.html --a11y axe

layoutlens page.html "Is it accessible?" --model gpt-4o --api-key sk-...

Run layoutlens

with no arguments (or --help

) to see the full flag reference: --query/-q

, --compare/-c

, --viewport/-v {desktop,mobile,tablet}

, --output/-o {text,json}

, --api-key

, --model/-m

, --a11y {hybrid,axe,llm}

.

- name: Visual UI Test
  run: |
    pip install layoutlens
    playwright install chromium
    layoutlens ${{ env.PREVIEW_URL }} "Is it accessible and mobile-friendly?"
python
import pytest
from layoutlens import LayoutLens

@pytest.mark.asyncio
async def test_homepage_quality():
    lens = LayoutLens()
    result = await lens.analyze("homepage.html", "Is this production-ready?")
    assert result.confidence > 0.8
    assert "yes" in result.answer.lower()

LayoutLens bundles a compact benchmark suite (18 fixtures / 74 labeled queries) for smoke-testing AI performance. For a larger, paper-rigor benchmark of AI judges of web UI quality β€” 4,000+ machine-verified items across accessibility, layout, and referring tasks, built on LayoutLens's own axe/browser machinery β€” see ** UIJudgeBench** (

dataset on Hugging Face). LayoutLens is a planned judge baseline there.

python benchmarks/run_benchmark.py --api-key sk-your-key

python benchmarks/run_benchmark.py \
  --api-key sk-your-key \
  --output benchmarks/my_results \
  --no-batch \
  --filename custom_results.json
python benchmarks/evaluation/evaluator.py \
  --answer-keys benchmarks/answer_keys \
  --results benchmarks/layoutlens_output \
  --output evaluation_report.json

The evaluator scores every answer deterministically (leading yes/no token vs the answer key; ambiguous answers count as incorrect) and writes an artifact with per-category and overall accuracy. The committed benchmarks/results/2026-07-21_gpt-4o-mini.json is a real measured run:

{
  "evaluation_summary": {
    "date": "2026-07-21",
    "model": "gpt-4o-mini",
    "total_queries": 74,
    "total_correct": 60,
    "ambiguous_answers": 7,
    "overall_accuracy": 0.811,
    "evaluator_version": "2.0",
    "evaluator_method": "Deterministic structured yes/no; ambiguous answers count as incorrect."
  },
  "category_results": {
    "responsive_design": {"total_queries": 21, "correct_predictions": 20, "accuracy": 0.952},
    "layout_alignment":  {"total_queries": 24, "correct_predictions": 19, "accuracy": 0.792},
    "accessibility":     {"total_queries": 21, "correct_predictions": 16, "accuracy": 0.762},
    "ui_components":      {"total_queries": 8,  "correct_predictions": 5,  "accuracy": 0.625}
  }
}

Create your own test data and answer keys:

from layoutlens import LayoutLens

async def run_custom_benchmark():
    lens = LayoutLens()

    test_cases = [
        {"source": "page1.html", "query": "Is it accessible?"},
        {"source": "page2.html", "query": "Is it mobile-friendly?"},
    ]

    results = []
    for case in test_cases:
        result = await lens.analyze(case["source"], case["query"])
        results.append(
            {
                "test": case,
                "result": result.to_json(),  # Clean JSON output
                "passed": result.confidence > 0.7,
            }
        )

    return results

Simple configuration options:

export OPENAI_API_KEY="sk-..."

lens = LayoutLens(
    api_key="sk-...",
    model="gpt-4o-mini",  # or "gpt-4o" for higher accuracy
    cache_enabled=True,   # Reduce API costs
    cache_type="memory",  # "memory" or "file"
)

Calibrate your trust to the tier you use:

Vision LLMs miss fine-grained UI differences. On DiffSpot (arXiv 2605.29615), a 2026 benchmark of fine-grained web-UI changes, the best frontier model scored 47.2% overall andunder 23% recall on the hard tier; open models hallucinated differences on 18–24% of identical pairs. Do not use the LLM tier as a sole gate for subtle visual regressions β€” that is what the deterministic scorers are for.Passing axe-core is not WCAG conformance. Automated rules cover only a subset of WCAG; Microsoft's a11y LLM evaluation makes the same disclaimer for its own checks. axe passing means "no automated rule failed", not "accessible".The deterministic scorers measure rendered facts, not full intent. The WCAG 2.5.8 spacing, inline, and unmodified-user-agent-control exceptions are modeled. Equivalent-control and essential-presentation exceptions still require review, as do interaction-history cases under WCAG 2.4.11. General text occlusion is a visual-quality signal, not a WCAG conformance claim. Findings carry their measured numbers so you can judge.Our own benchmark is small(74 labeled queries) and easier than DiffSpot-class tasks; the 81.1% figure is honest but narrow. The harness is model-agnostic (benchmarks/run_benchmark.py --model ...

) β€” re-run it rather than trusting ours.

  • πŸ“–
  • Comprehensive guides and API referenceFull Documentation - 🎯
  • Real-world usage patternsExamples - πŸ›
  • Report bugs, request features, get helpIssues

Natural Language- Write tests like you'd describe the UI to a colleague** Domain Expert Knowledge**- Built-in expertise in accessibility, CRO, mobile UX, and more** Rich Context Support**- Business goals, user personas, compliance standards, and technical constraints** Zero Selectors**- No more fragile XPath or CSS selectors** Visual Understanding**- AI sees what users see, not just code** Async-by-Default**- Concurrent processing for optimal performance** Simple API**- One analyze method handles single pages, batches, and comparisons** Structured JSON Output**- TypedDict schemas for full type safety in automation** Honest Benchmarking**- Compact built-in suite (81.1% measured accuracy, gpt-4o-mini, 74 queries); seeUIJudgeBenchfor the full-scale external benchmarkDeterministic Accessibility- Vendored axe-core WCAG 2.1 A/AA checks, no API key or LLM variance

Making UI testing as simple as asking "Does this look right?"

── more in #developer-tools 4 stories Β· sorted by recency
── more on @layoutlens 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/show-hn-layoutlens-a…] indexed:0 read:16min 2026-08-23 Β· β€”