Building an RL environment and dataset for legal DecoverAI built a legal-agent evaluation benchmark around a fictional matter, USA v. Cascade Timber Holdings, Inc., comprising 1,486 RFC 5322 .eml documents spanning November 2021 to December 2023 across 26 mailboxes, with 51 distinct tasks split into 10 flagship, 18 atomic/retrieval/join, 6 RLVR pilots and 17 litigation skill episodes. DecoverAI's internal assessment in suggestions.md scores the dataset Q ≈ 49/100, in the band marked "blocked for training", citing answer leakage through headers, few supervised targets, contradictory facts and only one matter. The dataset is positioned as sound for demos and as an evaluation seed, with labels limited to seed documents (one review tag each plus privileged and PII flags). A legal-agent evaluation benchmark built around a Matter, USA v. Cascade Timber Holdings, Inc. , built by DecoverAI. It is an email corpus with planted evidence chains, designed for training and evaluating models on legal evidence reasoning : responsiveness, privilege, chronology, knowledge analysis and joins across several documents. All companies, people and events are fictional. Not ready for training. The internal assessment suggestions.md https://github.com/decover-tech/frontier-legal/blob/main/suggestions.md scores it Q ≈ 49/100, in the band marked "blocked for training". It's sound as a demo and an evaluation seed, but the answers leak through headers, there are few supervised targets, some facts contradict each other, and there's only one matter. Read Known issues known-issues before you use it. | Documents rendered | 1,486 RFC 5322 .eml 250 seed, 1,150 generated in batches, 86 in 7 threads from thread kit | | Documents planned | 1,486 DOCUMENT MANIFEST.csv ; thread rows have status=thread-expansion | | Time span | Nov 2021 – Dec 2023 | | Custodians | 26 mailboxes | | Organizations | Cascade Timber client , Alder Point Partners administrator , Bellhaven Advisory broker , L&L Associates outside counsel , GreenAcre surveyor , Moss & Lane auditors , IRS, consultants, buyers | | Labels | Seed documents only: 1 review tag each, plus privileged and PII flags | | Benchmark | 51 distinct tasks: 10 flagship, 18 atomic/retrieval/join, 6 RLVR pilots and 17 litigation skill episodes | | Matters | 1 | The catalog follows the linked task-table format in FrontierSWE v2 https://github.com/Proximal-Labs/frontier-swe-v2 . It covers 51 distinct tasks in this repository. The 19 litigation skill family IDs include two aliases for existing pilots: CTH-LIT-09 maps to CTH-PRESERVATION-001 and CTH-LIT-10 maps to CTH-CHRONOLOGY-001 ; those episodes are counted once. Agent-visible prompts are in the flagship task pack https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl . These use the evidence-based scoring rubric https://github.com/decover-tech/frontier-legal/blob/main/benchmark/eval/scoring rubric.md . | Task ID | Task | Category | Split | |---|---|---|---| | CTH-AGENT-001 | Acreage knowledge chain https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl L1 | knowledge chronology | train | | CTH-AGENT-002 | Northwest parcel reconstruction https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl L2 | transaction reconstruction | train | | CTH-AGENT-003 | Broker registration and authority https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl L3 | registration authority | train | | CTH-AGENT-004 | Post-warning program continuation https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl L4 | autonomous investigation | test ood | | CTH-AGENT-005 | Clearwater transaction sequencing https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl L5 | temporal join | test ood | | CTH-AGENT-006 | Privilege and work-product review https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl L6 | privilege review | test id | | CTH-AGENT-007 | Whistleblower complaint and insider pool https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl L7 | whistleblower credibility | test ood | | CTH-AGENT-008 | Examination and preservation timeline https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl L8 | preservation analysis | test id | | CTH-AGENT-009 | Financial trail and exposure https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl L9 | financial reconstruction | test id | | CTH-AGENT-010 | Matter theory memo https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship prompts.jsonl L10 | case theory memo | test ood | Agent-visible prompts are in the atomic task pack https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl . These use the same evidence-based rubric; difficulty and split are recorded per prompt. | Task ID | Task | Category | Split | |---|---|---|---| | CTH-T1-001 | Responsiveness: transaction email https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L1 | responsiveness | train | | CTH-T1-002 | Responsiveness: disputed program scope https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L2 | responsiveness | train | | CTH-T1-003 | Privilege: counsel risk memo https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L3 | privilege single | train | | CTH-T1-004 | Privilege: bare forward https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L4 | privilege single | train | | CTH-T1-005 | Kovel retention contrast https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L5 | kovel contrast | train | | CTH-T0-006 | Entity and employer resolution https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L6 | entity resolution | train | | CTH-T0-007 | Event date versus document date https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L7 | date extraction | train | | CTH-R-008 | Acreage verification retrieval https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L8 | retrieval | train | | CTH-R-009 | Broker authority retrieval https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L9 | retrieval | train | | CTH-T2-010 | Registration-status contradiction https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L10 | contradiction | train | | CTH-T2-011 | Investor-description change https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L11 | contradiction | validation | | CTH-T2-012 | Parcel-scope competing explanations https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L12 | contradiction | train | | CTH-T2-013 | Broker-gap evidence join https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L13 | evidence join | train | | CTH-T2-014 | Acreage concern chronology https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L14 | chronology | train | | CTH-T2-015 | Notice and subsequent action https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L15 | knowledge | train | | CTH-T1-016 | Responsiveness versus privilege https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L16 | responsiveness | validation | | CTH-T2-017 | Abstention on examination closure https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L17 | negative control | validation | | CTH-T2-018 | Program-separation distractor control https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1 tier2 sample.jsonl L18 | distractor control | train | These tasks have deterministic verifiers. The first four are single-turn; preservation and chronology are multi-turn search/read/submit episodes. See the RLVR runner https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/README.md and evidence-agent interface https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/AGENT TASKS.md . | Task ID | Task | Category | Mode | |---|---|---|---| | CTH-DATE-001 | Sent-date extraction https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/tasks/CTH-DATE-001.json | Date extraction | Single-turn smoke test | | CTH-AUDIT-001 | Evidence audit 75 emails, 18 questions https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/tasks/CTH-AUDIT-001.json | Cross-document reasoning | Single-turn development | | CTH-INVENTORY-001 | Collection reconciliation 75 emails https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/tasks/CTH-INVENTORY-001.json | Collection inventory | Single-turn development | | CTH-INVENTORY-002 | Collection reconciliation 300 emails https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/tasks/CTH-INVENTORY-002.json | Collection inventory | Single-turn development | | CTH-PRESERVATION-001 | Preservation audit CTH-LIT-09 / legal-hold https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/preservation/CTH-PRESERVATION-001.json | Preservation | Multi-turn development | | CTH-CHRONOLOGY-001 | Evidence chronology CTH-LIT-10 / chronology https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/chronology/CTH-CHRONOLOGY-001.json | Chronology | Multi-turn development | These bounded development episodes have deterministic evidence and artifact checks. The registry https://github.com/decover-tech/frontier-legal/blob/main/benchmark/litigation skills/registry.json is authoritative for task packages, pilot aliases and dependencies. See suite usage https://github.com/decover-tech/frontier-legal/blob/main/benchmark/litigation skills/README.md and coverage and review status https://github.com/decover-tech/frontier-legal/blob/main/benchmark/litigation skills/COVERAGE.md . They evaluate supplied-policy exercises within one matter; the legacy preservation and chronology pilots still need separate independent oracle review. CTH-LIT-19 tests closure readiness with a missing trigger; its inclusion does not imply that closure is authorized. Completed October 9, 2026: 10 tasks, 10 runs per task per model, five models, 500 completed episodes , using Cascade Timber. The final ten-task dashboard https://cascade-timber-final-benchmark.decoverhq-in-8411.chatgpt.site private Sites access includes task-level means, scored counts, variability, and CSV export. Scores are deterministic verifier rewards, expressed as percentages. | Model | Completed / scored | Final composite | Observed task mean | Mean task SD pp | Worst–best episode score | Total cost USD | |---|---|---|---|---|---|---| | GPT-6 Astra | 100 / 100 | 44.01% | 44.01% | 5.60 | 5.00–86.40% | $86.02 | | Grok 4.7 | 100 / 99 | N/A | 26.76% | 8.49 | 0.00–82.55% | $137.25 | | Claude Opus 5.5 | 100 / 100 | 24.40% | 24.40% | 6.45 | 0.00–73.59% | $350.52 | | GLM 5.3 Prime | 100 / 99 | N/A | 18.63% | 9.43 | 0.00–70.35% | $106.64 | | Gemini 3.1 Pro Preview | 100 / 100 | 3.98% | 3.98% | 5.52 | 0.00–40.79% | $45.85 | The final composite equally weights the ten task means, each over ten scored runs. The observed task mean uses available scored runs and equally weights tasks; it is incomplete for Grok and GLM , rather than a final composite. Grok chronology has 9/10 scored runs after an invalid provider response; GLM triage has 9/10 after exhausting its episode output-token budget. These two episodes remain unscored; all scored zeros are retained. Mean task SD averages the ten within-task sample standard deviations in percentage points; it is not a composite confidence interval. Worst–best ranges span scored episodes across all ten tasks. All 500 episodes passed saved replay verification , covering 21,331 transitions . There were 498 scored episodes and 0 full passes . Total recorded provider cost was $726.28 , including all recorded attempts, with no unknown-cost calls. Each episode starts with fresh context and allows up to 300 actions; requested reasoning effort is low, with 16,384 output tokens per request and 131,072 output tokens per episode. Every model received corrected model-neutral search/read/submit instructions; provider fallbacks were disabled. These development results cover ten tasks in one synthetic matter. Preservation and chronology still need independent equivalent-evidence oracle review. Strict artifact formatting and accepted-evidence matching affect scores; these graders do not measure free-form prose quality. Equal requested reasoning effort does not imply equal compute across providers. The earlier 50-episode chronology pilot is separate from this leaderboard; its result summary https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/chronology/results/chronology-10x5-2026-10-09.json and individual episode scores https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/chronology/results/chronology-10x5-2026-10-09-episodes.csv remain available. Neither those pilot scores nor Gemini’s separate single validation are included in the final ten-task results. - You have to join documents. The corpus is built so that no single document settles an issue. Take the question of whether an acreage problem was known before the applications were filed: answering it takes option dates, application dates, ledger entries and one employee's role, all from different documents. - The past is never written from hindsight. Each batch is scanned for terms its authors couldn't have known yet, such as the subpoena, counsel's engagement or the pause. The results are in CONTINUITY BATCH0N.md . A model can't pick up the answer from a document that was written "too early". - Doctrines come as contrast pairs. Privilege, work product, Kovel consultants, bare forwards and responsiveness each appear as small chains: a clear positive, a clear negative and a genuinely ambiguous case. - Some readings are left open on purpose. Several issues are DISPUTED by design, for example EMAIL-117 a clerical error or a cover story? . The right output is to state the uncertainty, not to force a label. - The email looks like real collected mail. Quoted history nests one level per reply, quoting follows each sender's mail client Outlook or Gmail , each organization has its own signature, time zones follow daylight saving, and about a third of the corpus is routine noise, decoys and near-duplicates. - Long threads, not just pairs. The first four expanded threads THR-001–004 run 12–14 messages each, with forks, reply-alls that add or drop people, side forwards and unanswered questions. With the density extensions, the longest coherent reply chain is 14 deep. - Denser evidence chains. THR-005–007 add 32 emails and eight supporting text records for held Q4 packets, source-credit allocations to buyers, and a dated remediation pilot. See the expansion register https://github.com/decover-tech/frontier-legal/blob/main/documentation/DENSITY EXPANSION.md for new fictional facts, counterevidence and questions deliberately left unresolved. - Real contracts, versioned. KW-01/02 option agreements go from drafts DOCX to wet-signed scans PDF, no text layer . The Bellhaven credit purchase agreements and the pre-broker template also carry their version history. Each email carries the version that existed on its date, and the copies in data/contracts/ are byte-identical to the attachments, so hash deduplication links them. - Signature logos. Outlook-style orgs Cascade Timber, L&L, Moss & Lane, Whitaker carry an inline logo in an HTML part, and quoted signatures keep theirs, so long threads pile up image001.png , image002.png and so on. Inline logos aren't counted as attachments. The seed keeps its original structure. data/emails/ corpus tracked; new files need git add -f , see .gitignore Custodians/