A legal-agent evaluation benchmark built around a Matter, USA v. Cascade Timber Holdings, Inc., built by DecoverAI. It is an email corpus with planted evidence chains, designed for training and evaluating models on legal evidence reasoning: responsiveness, privilege, chronology, knowledge analysis and joins across several documents. All companies, people and events are fictional.
Not ready for training. The internal assessment (suggestions.md)
scores it Q ≈ 49/100, in the band marked "blocked for training". It's sound as a demo and
an evaluation seed, but the answers leak through headers, there are few supervised
targets, some facts contradict each other, and there's only one matter. Read
Known issues before you use it.
| Documents rendered | 1,486 RFC 5322 .eml (250 seed, 1,150 generated in batches, 86 in 7 threads fromthread_kit ) |
| Documents planned | 1,486 ( DOCUMENT_MANIFEST.csv ; thread rows havestatus=thread-expansion ) |
| Time span | Nov 2021 – Dec 2023 |
| Custodians | 26 mailboxes |
| Organizations | Cascade Timber (client), Alder Point Partners (administrator), Bellhaven Advisory (broker), L&L Associates (outside counsel), GreenAcre (surveyor), Moss & Lane (auditors), IRS, consultants, buyers |
| Labels | Seed documents only: 1 review tag each, plus privileged and PII flags |
| Benchmark | 51 distinct tasks: 10 flagship, 18 atomic/retrieval/join, 6 RLVR pilots and 17 litigation skill episodes |
| Matters | 1 |
The catalog follows the linked task-table format in FrontierSWE v2.
It covers 51 distinct tasks in this repository. The 19 litigation skill family IDs
include two aliases for existing pilots: CTH-LIT-09 maps to CTH-PRESERVATION-001
and CTH-LIT-10 maps to CTH-CHRONOLOGY-001; those episodes are counted once.
Agent-visible prompts are in the flagship task pack. These use the evidence-based scoring rubric.
| Task ID | Task | Category | Split |
|---|---|---|---|
| CTH-AGENT-001 | Acreage knowledge chain | knowledge chronology | train |
| CTH-AGENT-002 | Northwest parcel reconstruction | transaction reconstruction | train |
| CTH-AGENT-003 | Broker registration and authority | registration authority | train |
| CTH-AGENT-004 | Post-warning program continuation | autonomous investigation | test_ood |
| CTH-AGENT-005 | Clearwater transaction sequencing | temporal join | test_ood |
| CTH-AGENT-006 | Privilege and work-product review | privilege review | test_id |
| CTH-AGENT-007 | Whistleblower complaint and insider pool | whistleblower credibility | test_ood |
| CTH-AGENT-008 | Examination and preservation timeline | preservation analysis | test_id |
| CTH-AGENT-009 | Financial trail and exposure | financial reconstruction | test_id |
| CTH-AGENT-010 | Matter theory memo | case theory memo | test_ood |
Agent-visible prompts are in the atomic task pack. These use the same evidence-based rubric; difficulty and split are recorded per prompt.
| Task ID | Task | Category | Split |
|---|---|---|---|
| CTH-T1-001 | Responsiveness: transaction email | responsiveness | train |
| CTH-T1-002 | Responsiveness: disputed program scope | responsiveness | train |
| CTH-T1-003 | Privilege: counsel risk memo | privilege single | train |
| CTH-T1-004 | Privilege: bare forward | privilege single | train |
| CTH-T1-005 | Kovel retention contrast | kovel contrast | train |
| CTH-T0-006 | Entity and employer resolution | entity resolution | train |
| CTH-T0-007 | Event date versus document date | date extraction | train |
| CTH-R-008 | Acreage verification retrieval | retrieval | train |
| CTH-R-009 | Broker authority retrieval | retrieval | train |
| CTH-T2-010 | Registration-status contradiction | contradiction | train |
| CTH-T2-011 | Investor-description change | contradiction | validation |
| CTH-T2-012 | Parcel-scope competing explanations | contradiction | train |
| CTH-T2-013 | Broker-gap evidence join | evidence join | train |
| CTH-T2-014 | Acreage concern chronology | chronology | train |
| CTH-T2-015 | Notice and subsequent action | knowledge | train |
| CTH-T1-016 | Responsiveness versus privilege | responsiveness | validation |
| CTH-T2-017 | Abstention on examination closure | negative control | validation |
| CTH-T2-018 | Program-separation distractor control | distractor control | train |
These tasks have deterministic verifiers. The first four are single-turn; preservation and chronology are multi-turn search/read/submit episodes. See the RLVR runner and evidence-agent interface.
| Task ID | Task | Category | Mode |
|---|---|---|---|
| CTH-DATE-001 | Sent-date extraction | Date extraction | Single-turn smoke test |
| CTH-AUDIT-001 | Evidence audit (75 emails, 18 questions) | Cross-document reasoning | Single-turn development |
| CTH-INVENTORY-001 | Collection reconciliation (75 emails) | Collection inventory | Single-turn development |
| CTH-INVENTORY-002 | Collection reconciliation (300 emails) | Collection inventory | Single-turn development |
| CTH-PRESERVATION-001 | Preservation audit (CTH-LIT-09 / legal-hold) | Preservation | Multi-turn development |
| CTH-CHRONOLOGY-001 | Evidence chronology (CTH-LIT-10 / chronology) | Chronology | Multi-turn development |
These bounded development episodes have deterministic evidence and artifact checks. The registry is authoritative for task packages, pilot aliases and dependencies. See suite usage and coverage and review status. They evaluate supplied-policy exercises within one matter; the legacy preservation and chronology pilots still need separate independent oracle review.
CTH-LIT-19 tests closure readiness with a missing trigger; its inclusion does not imply that closure is authorized.
Completed October 9, 2026: 10 tasks, 10 runs per task per model, five models, 500 completed episodes, using Cascade Timber. The final ten-task dashboard (private Sites access) includes task-level means, scored counts, variability, and CSV export. Scores are deterministic verifier rewards, expressed as percentages.
| Model | Completed / scored | Final composite | Observed task mean | Mean task SD (pp) | Worst–best episode score | Total cost (USD) |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 100 / 100 | 44.01% | 44.01% | 5.60 | 5.00–86.40% | $86.02 |
| Grok 4.7 | 100 / 99 | N/A | 26.76% | 8.49 | 0.00–82.55% | $137.25 |
| Claude Opus 5.5 | 100 / 100 | 24.40% | 24.40% | 6.45 | 0.00–73.59% | $350.52 |
| GLM 5.3 Prime | 100 / 99 | N/A | 18.63% | 9.43 | 0.00–70.35% | $106.64 |
| Gemini 3.1 Pro Preview | 100 / 100 | 3.98% | 3.98% | 5.52 | 0.00–40.79% | $45.85 |
The final composite equally weights the ten task means, each over ten scored runs. The observed task mean uses available scored runs and equally weights tasks; it is incomplete for Grok and GLM, rather than a final composite. Grok chronology has 9/10 scored runs after an invalid provider response; GLM triage has 9/10 after exhausting its episode output-token budget. These two episodes remain unscored; all scored zeros are retained. Mean task SD averages the ten within-task sample standard deviations in percentage points; it is not a composite confidence interval. Worst–best ranges span scored episodes across all ten tasks.
All 500 episodes passed saved replay verification, covering 21,331 transitions. There were 498 scored episodes and 0 full passes. Total recorded provider cost was $726.28, including all recorded attempts, with no unknown-cost calls. Each episode starts with fresh context and allows up to 300 actions; requested reasoning effort is low, with 16,384 output tokens per request and 131,072 output tokens per episode. Every model received corrected model-neutral search/read/submit instructions; provider fallbacks were disabled.
These development results cover ten tasks in one synthetic matter. Preservation and chronology still need independent equivalent-evidence oracle review. Strict artifact formatting and accepted-evidence matching affect scores; these graders do not measure free-form prose quality. Equal requested reasoning effort does not imply equal compute across providers.
The earlier 50-episode chronology pilot is separate from this leaderboard; its result summary and individual episode scores remain available. Neither those pilot scores nor Gemini’s separate single validation are included in the final ten-task results.
- You have to join documents. The corpus is built so that no single document settles an issue. Take the question of whether an acreage problem was known before the applications were filed: answering it takes option dates, application dates, ledger entries and one employee's role, all from different documents.
- The past is never written from hindsight. Each batch is scanned for terms its authors
couldn't have known yet, such as the subpoena, counsel's engagement or the . The
results are in
CONTINUITY_BATCH0N.md. A model can't pick up the answer from a document that was written "too early". - Doctrines come as contrast pairs. Privilege, work product, Kovel consultants, bare forwards and responsiveness each appear as small chains: a clear positive, a clear negative and a genuinely ambiguous case.
- Some readings are left open on purpose. Several issues are
[DISPUTED]by design, for example EMAIL-117 (a clerical error or a cover story?). The right output is to state the uncertainty, not to force a label. - The email looks like real collected mail. Quoted history nests one level per reply, quoting follows each sender's mail client (Outlook or Gmail), each organization has its own signature, time zones follow daylight saving, and about a third of the corpus is routine noise, decoys and near-duplicates.
- Long threads, not just pairs. The first four expanded threads (THR-001–004) run 12–14 messages each, with forks, reply-alls that add or drop people, side forwards and unanswered questions. With the density extensions, the longest coherent reply chain is 14 deep.
- Denser evidence chains. THR-005–007 add 32 emails and eight supporting text records for held Q4 packets, source-credit allocations to buyers, and a dated remediation pilot. Seethe expansion register for new fictional facts, counterevidence and questions deliberately left unresolved.
- Real contracts, versioned. KW-01/02 option agreements go from drafts (DOCX) to wet-signed
scans (PDF, no text layer). The Bellhaven credit purchase agreements and the pre-broker
template also carry their version history. Each email carries the version that existed on
its date, and the copies in
data/contracts/are byte-identical to the attachments, so hash deduplication links them. - Signature logos. Outlook-style orgs (Cascade Timber, L&L, Moss & Lane, Whitaker) carry
an inline logo in an HTML part, and quoted signatures keep theirs, so long threads pile up
image001.png,image002.pngand so on. Inline logos aren't counted as attachments. The seed keeps its original structure.
data/emails/ # corpus (tracked; new files need `git add -f`, see .gitignore)
Custodians/<Name>/EMAIL-NNN_<subject>.eml
Loadfile_Cascade_Timber.{csv,dat} # load file: all 1,486; newest 32 have blank review labels
README.md # corpus build notes (v3 realism pass)
data/emails/Exhibits/ # 38 standalone exhibit PDFs + ../Exhibit_Manifest.csv (not attached to emails)
definitions/ # labeling protocols: Responsiveness, ACP, Work Product, Subpoena (summons)
logs/ # batch ledgers, continuity reports, thread logs, demo privilege log (untracked)
benchmark/ # tasks, hidden gold, splits, schema, rubric, GOLD_LABELS.csv
documentation/ # AUTHORING ONLY:
CASE_BIBLE.md EVIDENCE_ARCS.md # ground-truth world model; arcs A–J, planned joins and open questions
DOCUMENT_MANIFEST.csv # 1,454-row plan with arc, event and evidentiary role
CONTRADICTIONS_RESOLVED.md # decisions and edits for the four known contradictions
MATTER_AGENT_TASKS.md suggestions.md Cascade_Timber_EML_Dataset_Plan.md
output/dataset_inconsistency_report.md # document-consistency review + resolution log
data/contracts/ # standalone contract collection (every version) + INDEX.csv → carrying emails
tools/exhibit_kit/ # rebuilds the standalone exhibits from one spec
tools/doc_kit/ # attachment builder: versioned library docs → PDF/DOCX/XLSX, scans, apply to emails, publish
library/ plans/ # document content + version history; email→version attach plans (ATT-00N)
tools/thread_kit/ # thread expander: context, validate, render, rollback, scan, logos
rules.json # knowledge cutoffs, participant windows, logo orgs (AUTHORING ONLY)
logos/ # org logos (full size + signature size)
| Source | Coverage | Form |
|---|---|---|
benchmark/hidden_gold/seed_header_labels.csv and the load file'sTAG /PRIVILEGED columns |
Seed EMAIL-001–250 | One mixed tag per document, such as Not Responsive ,Routine ,Privileged Legal Advice ,Knowledge/Scienter orRed Flag . This isn't multi-label, and privileged documents carry no responsiveness label. |
documentation/DOCUMENT_MANIFEST.csv |
All 1,454 rows | arc ,event_id andintended_evidentiary_role . These are authoring intent, not reviewed gold. |
benchmark/hidden_gold/ |
28 tasks, all answerable from seed documents | Required, counter, distractor and context evidence IDs; gold facts and inferences; must_include ,must_not_claim andmust_qualify lists; unknowns. The schema is inschemas/task.schema.json . |
CASE_BIBLE.md /EVIDENCE_ARCS.md |
The whole matter | Prose ground truth with [ESTABLISHED] ,[PROPOSED] ,[DISPUTED] and[INFERENCE] tags. This is the fastest source for writing new (task, evidence, target) examples. |
The emails carry no label headers: the X-Decover-* headers were stripped from the original 1,454 files and are absent from the 32 additions. DocIDs
live in the filenames (EMAIL-NNN_<subject>.eml) and the load file. Document-level labels for the whole
original corpus are in benchmark/GOLD_LABELS.csv (machine-drafted, expert review pending).
The newest 32 documents are unreviewed and not included in those labels or existing
hash-pinned benchmark snapshots; their load-file label columns are blank.
- Don't split at random. Documents in a thread or evidence chain answer each other.
benchmark/splits/evidence_packages_sample.csvdefines packages by family that must stay together, currently 17 of them. - Task splits are
train,test_id(a skill seen in training, on an unseen family) andtest_ood(held-out flagship compositions). The mapping is in thebenchmark README . - All of it is one matter. A held-out split measures generalization within this matter, not transfer to a new one. Building a second, independent matter for evaluation is still open.
- Keep the model away from: the load file's
TAGandPRIVILEGEDcolumns (withhold them; the emails themselves no longer carry label headers),CASE_BIBLE.md,EVIDENCE_ARCS.md, the manifest'sarc/event_id/role columns, the ledgers andhidden_gold/. The model may see the EML bodies, the Definitions files and ordinary load-file metadata.
These come from suggestions.md and from output/dataset_inconsistency_report.md. The benchmark tasks
don't depend on them.
- Resolved (2026-09-28): the § 7602 summons/subpoena framing, the two Kevin Tran identities,
the two Cascade addresses and the deadline chain. See
documentation/CONTRADICTIONS_RESOLVED.md. - Resolved (2026-09-29): the standalone exhibits in
data/emails/Exhibits/now agree with the executed contracts and seed emails (Clearwater option dates and installments, buyer lot amounts, OR/WA filing-receipt dates, status footers, placeholder numbers). Rebuilt bytools/exhibit_kit/build_exhibits.py; the resolution log is at the end of the inconsistency report. - Still open, by design or header-locked: - The seed parcel CSV and GreenAcre survey (EMAIL-037/039) use NW-1042…NW-1093 and different
acreage pairs from the NW-01…NW-08 table in
CASE_BIBLE.md§2. No crosswalk exists; don't merge them by position.- "Batch N" and "buyer lot N" in generated email subjects are templated labels that don't track the exhibits or the executed agreements. Join buyer records on buyer, date and amount, not on batch or lot number.
- EMAIL-793 (9/8/22) has a subject that cites the 9/15 notice; four generated "extension chain" notes predate the summons. Both are Subject-header anachronisms, kept because headers are locked.
- The seed parcel CSV and GreenAcre survey (EMAIL-037/039) use NW-1042…NW-1093 and different acreage pairs from the NW-01…NW-08 table in
- The privilege log is out of step with the corpus. Most of its 1,009 rows name people who don't appear in the corpus.
- Workbook search queries and chronology are stale. They cover only EMAIL-001–030.
- Emails are short and convenient. Replies average about 32 new words, and some
admissions are too tidy (e.g. "keep this between us"). The
thread_kitthreads push back on this: they have longer replies, a validator that flags tidy admissions, and hedged readings that stay open. - Generation isn't reproducible. No generator or model versions or file hashes are recorded.
For the single-turn, multi-model RLVR tasks (date extraction, evidence audit, and
collection reconciliation), see
benchmark/rlvr/README.md. It includes an offline
self-test, OpenAI/Anthropic/Gemini/OpenRouter adapters, deterministic scoring, and saved
JSON run records.
- Parse the
.emlfiles with any MIME library. UseDATESENT, or theDate:header, for chronology. The DocID order follows collection order, not time. - Withhold the load file's
TAG/PRIVILEGEDcolumns andbenchmark/GOLD_LABELS.csvfrom model inputs. The.emlfiles are already clean: there are noX-Decover-*headers. - Build examples by evidence package, not by document, and hold out whole packages.
- For any label beyond the seed tags, generate targets from
CASE_BIBLE.mdand the benchmark gold, and have an expert review them. Treat[PROPOSED]facts as unknowns until the documents that support them exist. - Score with
benchmark/eval/scoring_rubric.md. It covers evidence recall, join correctness, counter-evidence, restraint on uncertain points, citation accuracy and stopping efficiency.
- All 1,400 planned documents are rendered. Thread expansion has added 54 more: THR-001
Clearwater scouting (EMAIL-1401–1412), THR-002 Bellhaven registration (1421–1434), THR-003
Q4 push-through (1441–1454) and THR-004 the $2.4M estimate (1461–1474). Their specs and
reports are in
Logs/threads/. - The authorized density pass adds THR-005–007 (EMAIL-1475–1506), bringing the total to
1,486 messages. All original emails remain byte-identical. Story specs, a metadata
synchronization/checking script and validation results are in
tools/thread_kit/expansions/;the authoring register explains the additions and their limits. - The top fixes for training readiness, from
suggestions.md:- Move labels out of the headers.
- Write 150–200 examples reviewed by experts, with multi-label responsiveness, privilege basis, chronology and grounded QA, including examples where the right answer is to abstain.
- Resolve the contradictions listed above.
- Fix the summons/subpoena mislabel.
- Add a second matter for evaluation, plus hashes.