{"slug": "show-hn-litigationbench-a-litigation-task-based-ai-benchmark", "title": "Show HN: LitigationBench. A Litigation Task-Based AI Benchmark", "summary": "Litco released LitigationBench, a benchmark for language models on litigation work, showing that models with Litco's safeguards active report zero fabricated authorities on ordinary tasks, while models without safeguards invent case law at varying rates. The benchmark runs 59 tasks twice per model, once without safeguards and once inside Litco's production agent, publishing both scores, the gap, and every failure. Litco's results indicate that several models' rewritten Supreme Court brief introductions beat the filed brief in blind preference tests by a three-judge panel.", "body_md": "# LitigationBench\n\nby Litco\n\nA benchmark for language models on litigation work. Every model runs the same tasks twice: once without Litco’s safeguards, once with them active inside Litco’s production agent. Litco publishes both numbers, the gap between them, and every failure, including failures by the models inside the product.\n\n## Leaderboard\n\nClick a column to sort (arrows mark the direction that counts as better); click a row for both settings side by side and the row’s serving pins. Models missing a setting in this preview sit below the scored rows with the reason stated.\n\n**The composite** is the run’s quality score for the\nsetting shown. Litco publishes the score before penalties beside\nthe composite, so readers can see the effect of those penalties for themselves. The fabricated-\nauthorities and false-premises columns are shaded per cell: green for zero, amber for an\nadopted false premise, red for a fabrication; the full flag detail is in each row’s expanded\npanel and in the candor matrix below. Cost is the metered provider bill for the tasks shown.\nThe self-hosted row runs on Litco’s own hardware; its cost column is priced at OpenRouter’s\ncurrent rates for the same open-weight model, qwen/qwen3.6-35b-a3b, so every row’s cost\nrests on the same basis.\n\n* Without safeguards, the cost column shows the bill for all 59 tasks rather than a fraction of a cent per task.\n\n## Fabricated authorities, without Litco’s safeguards and with them.\n\nEvery fabricating model reports zero fabricated authorities on its ordinary\nwork with Litco’s safeguards active. The handful that remain surfaced only when the run\ndeliberately pressed a model with a false premise or an invented authority against the live\ncorpus. The chart pairs each model’s count without Litco’s safeguards against the same\nmodel’s count with them active. A green tick marks a measured zero, and an amber bar counts\nfabrications that happened on those provoking tasks. **Lower is better, and zero is the goal.**\n\n## Each model’s cost and score.\n\nEach dot is one model. Farther right costs more per task, and higher up scores better on the tasks. The self-hosted model runs on Litco’s own hardware and is priced at OpenRouter’s rates for the same model, so every dot rests on the same cost basis. Click a dot to compare two models side by side.\n\n## Which models invent case law, and how often.\n\nCounts of fabricated authorities per model over the 59-task battery without Litco’s safeguards. These are inventions of legal authority by the model running alone.\n\n## Failures under pressure, model by model.\n\nThe candor tasks press each model with invented premises, questions that lack a real legal response, and pressure to produce authority for a position no authority supports. The chart counts three failures for each model: fabricated authorities, adopted false premises, and hedges under pressure. Red bars show the model without safeguards, and blue bars show the same model inside Litco.\n\n### Per-category skill scores, inside Litco\n\nSix skills, graded per task and rolled up per model: candor under pressure, faithful case reading, treatment awareness, procedural competence, appellate issue framing, and reasoning and drafting.\n\nSpecial task set · practitioner indistinguishability\n\n## Platform introductions against filed ones, blind.\n\nThis task set takes the introduction section of a real OT2025 Supreme Court merits brief and has each model rewrite it from the same record. A blind three-judge panel, cross-vendor and in both presentation orders, first tries to spot which one a machine wrote, then, with no AI framing at all, simply says which introduction it prefers. Detection sits close to a coin flip for most models. On preference, several models’ rewrites beat the filed brief most of the time.\n\nSpecial task set · cert-QP framing\n\n## Ranking question-presented drafts, pairwise.\n\nThis task set hands each model a granted certiorari petition and has it draft the question presented. One petition was granted after every model’s training cutoff and counts double. A judge then compares every pair of drafts head to head rather than scoring each draft alone, and the bars show how often each model’s draft won its comparisons.\n\nSpecial task set · AI-isms\n\n## How machine-written each draft reads.\n\nThe benchmark scores each model’s drafting output against a public list of\nrecognizable AI writing tells: tic words, spaced em-dashes, stacked verbless fragments, the\n“It’s not X. It’s Y.” reveal, and the rest of the list, reported as tells per 1,000 words\nrather than pass or fail. Every model wrote the same three pieces: a brief argument, a\nclient letter, and a research memo, each with the same length target and all of the law\nsupplied in the prompt, so every rate rests on a comparable amount of text. These rates\nnow count against the leaderboard: a model’s drafting score is reduced by its rate of AI\nwriting tells, up to fifteen points.\n**Lower is better, and zero is the goal.**\n\nSome models wrote too little in this task set to measure fairly. Their rows show “not enough text to score” instead of a number.\n\n### What gives each model away\n\nOne card per model, heaviest accent first. Each chip is one tell counted in that model’s three drafts, with how often it fired; the chips that name a construction, such as the “It’s not X. It’s Y.” reveal, quote the pattern being counted, not the model’s sentence.\n\nSpecial task set · case characterization\n\n## Reading an opinion: holding, disposition, facts.\n\nThis task set gives each model an opinion and grades three things: separating the holding from dicta, stating the disposition correctly, and getting the facts right. Each bar is the share of those calls the model got right, graded against a reading of each opinion that a person checked by hand. Each dot places one model inside Litco, best to worst, on an axis zoomed to where the scores actually sit.\n\nSpecial task set · calendaring, FRCP/local rules\n\n## Deadlines computed under the Federal Rules.\n\nCalendaring tasks hand each model a trigger date and a filing deadline to compute under the Federal Rules. Every deadline in the battery is checked against a computer program that applies Rule 6(a)’s counting rules directly, so the grading never depends on another model’s arithmetic. Each dot places one model inside Litco, best to worst, on an axis zoomed to where the scores actually sit.\n\nSpecial task set · candor, across three settings\n\n## The candor matrix, three settings.\n\nThe candor tasks press each model with invented premises, questions the record cannot resolve, and pressure to cite authority, across three settings: without Litco at all, with Litco active but no real corpus mounted, and with Litco active against the live 3.6-million-opinion litlex corpus under deliberate pressure. Each cell below counts fabricated authorities and adopted false premises for that model in that setting. Hedge counts ride along for information only and never turn a cell red.\n\n## Head to head\n\nAny two models, axis by axis, in the mode selected on the leaderboard. Pick from the dropdowns or click points on the scatter.\n\n## Method\n\n### What a run is\n\nA model is dropped into the production Litco matter agent (the same agent loop, tools, and verification stack Litco’s customers run) and given the task battery against a seeded litigation matter and the production case-law corpus. Nothing about the harness is model-specific: every model sees identical prompts, tools, and iteration budgets. Spend is metered per call.\n\n### The two modes\n\nBattery v2 measures every model twice. **With Litco safeguards** is the production\nconfiguration: the agent with its verification stack active, over a 42-task battery.\n**Without safeguards** is the model on its own over a 59-task battery that includes\nthe fabrication probes: the stack observes and records but never intervenes. The\nper-model difference between the two composites is printed on the chart above. The two\nbatteries share their core tasks; composites are compared across modes as the run\nreports them, and the battery sizes are stated wherever the numbers appear.\n\n### Candor penalties and routing eligibility\n\nA model that invents case law, citing an authority that does not exist, wears a red flag on every surface of this page and loses twenty points from its composite for each fabrication. A model that adopts a false premise embedded in a question loses ten points for each adoption. Litco floors every composite at twenty, however many penalties a model accumulates in a run. A model that fabricates also loses Litco routing eligibility for that mode, whatever its score, because an invented citation is the one failure a legal tool may never let through. A model’s drafting score is also reduced by its rate of AI writing tells, measured in this page’s AI-isms task set, up to fifteen points. Litco publishes the score before penalties beside the composite, so readers can see the effect of each penalty for themselves.\n\n### Infrastructure failures\n\nA task whose model calls moved zero tokens is an infrastructure failure: it lands in an infra count, never in a quality axis. A row where most tasks failed that way publishes as quarantined with the failure stated, rather than as a low score. One assisted row is quarantined in the current preview, and the leaderboard says so on the row.\n\n### Serving policy and pinning\n\nPublic rows run through OpenRouter pinned to a named provider at a disclosed precision, with zero-data-retention routing requested at call time; two models were not ZDR-routable through OpenRouter and ran on the vendor’s direct API, which their rows disclose. Self-hosted rows run on Litco-operated hardware and are labeled as such. Because scored traffic goes through public endpoints rather than Litco infrastructure, a model vendor can independently verify the serving conditions of its own row. Lane or precision changes count as new rows, never silent edits.\n\n### Hold-out policy\n\nLitco publishes the methodology and keeps the task set private, because published tasks invite tuning to them, and a tuned-to benchmark stops predicting real work. The battery is versioned; when a battery version is retired, its tasks are disclosed and a fresh version replaces them. Scores are only ever compared within a battery version.\n\n### Preview status\n\nThis page currently renders the run’s own reported numbers from its live progress feed, marked PREVIEW. Judging is still in progress for the per-category skill breakdowns, the question-presented pairwise comparisons, and the brief-introduction indistinguishability test; latency percentiles and the final uniform rescore land with the final ingest, which replaces this data, and the changelog will say so.\n\n### What the assisted mode is measuring\n\nThe verification stack the assisted rows run under is the shipped product: byte-level\nquote checks, retrieval fingerprints, citation resolution, the citator, verify-on-stop, and\nthe document commit gate. [The verification page describes each\nmechanism](/verification); this page measures what they change.\n\n## Changelog\n\n## FAQ\n\n## Is this related to Stanford’s LegalBench or to Benchmark Litigation?\n\nNo on both counts. LitigationBench is unrelated to LegalBench, the academic legal-reasoning benchmark from Stanford, and unrelated to Benchmark Litigation, the directory that ranks litigation firms. This page benchmarks AI models on litigation work inside the Litco platform.\n\n## Why does a strong model score lower with Litco’s safeguards active?\n\nThe frontier models that never fabricate never trigger the candor penalty, and Litco’s checks add steps to every task, so a model that never fabricated can score a few points lower with the safeguards active than without them. The two settings run different batteries, too: the with-safeguards battery adds adversarial pressure tasks the without-safeguards battery does not run, so the two scores are not directly comparable. The safeguards exist for the failures a firm cannot predict in advance, and the models that do fabricate show what changes: every fabricating model reports zero fabricated authorities on its ordinary work with Litco’s safeguards active.\n\n## Which lane did each model run on, and what about data retention?\n\nThe lane and precision are pinned on every row and in each row’s detail panel, under the serving policy above. Self-hosted rows run on Litco-operated hardware, and their traffic never leaves that hardware.\n\n## How do new models get onto the board?\n\nWhen a model worth testing ships, Litco runs the current battery version against it and republishes. The changelog records every addition, re-run, and battery rotation, and each row keeps its run date.\n\n## Can I see the tasks, or run the battery myself?\n\nThe task set is private while its battery version is live; the hold-out policy above says why. Retired battery versions are disclosed. If you want a model evaluated, or want to reproduce the method against your own task set, get in touch.", "url": "https://wpnews.pro/news/show-hn-litigationbench-a-litigation-task-based-ai-benchmark", "canonical_source": "https://litco.ai/litigationbench", "published_at": "2026-07-23 01:11:18+00:00", "updated_at": "2026-07-23 01:22:13.904897+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-products", "ai-agents"], "entities": ["Litco", "LitigationBench", "OpenRouter", "qwen/qwen3.6-35b-a3b", "Supreme Court"], "alternates": {"html": "https://wpnews.pro/news/show-hn-litigationbench-a-litigation-task-based-ai-benchmark", "markdown": "https://wpnews.pro/news/show-hn-litigationbench-a-litigation-task-based-ai-benchmark.md", "text": "https://wpnews.pro/news/show-hn-litigationbench-a-litigation-task-based-ai-benchmark.txt", "jsonld": "https://wpnews.pro/news/show-hn-litigationbench-a-litigation-task-based-ai-benchmark.jsonld"}}