cd /news/ai-agents/building-an-rl-environment-and-datas… · home › topics › ai-agents › article
[ARTICLE · art-148522] src=github.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Building an RL environment and dataset for legal

DecoverAI built a legal-agent evaluation benchmark around a fictional matter, USA v. Cascade Timber Holdings, Inc., comprising 1,486 RFC 5322 .eml documents spanning November 2021 to December 2023 across 26 mailboxes, with 51 distinct tasks split into 10 flagship, 18 atomic/retrieval/join, 6 RLVR pilots and 17 litigation skill episodes. DecoverAI's internal assessment in suggestions.md scores the dataset Q ≈ 49/100, in the band marked "blocked for training", citing answer leakage through headers, few supervised targets, contradictory facts and only one matter. The dataset is positioned as sound for demos and as an evaluation seed, with labels limited to seed documents (one review tag each plus privileged and PII flags).

read14 min views2 publishedOct 9, 2026
Building an RL environment and dataset for legal
Image: Michielbdejong (auto-discovered)

A legal-agent evaluation benchmark built around a Matter, USA v. Cascade Timber Holdings, Inc., built by DecoverAI. It is an email corpus with planted evidence chains, designed for training and evaluating models on legal evidence reasoning: responsiveness, privilege, chronology, knowledge analysis and joins across several documents. All companies, people and events are fictional.

Not ready for training. The internal assessment (suggestions.md) scores it Q ≈ 49/100, in the band marked "blocked for training". It's sound as a demo and an evaluation seed, but the answers leak through headers, there are few supervised targets, some facts contradict each other, and there's only one matter. Read Known issues before you use it.

| Documents rendered | 1,486 RFC 5322 .eml (250 seed, 1,150 generated in batches, 86 in 7 threads fromthread_kit ) | | Documents planned | 1,486 ( DOCUMENT_MANIFEST.csv ; thread rows havestatus=thread-expansion ) | | Time span | Nov 2021 – Dec 2023 | | Custodians | 26 mailboxes | | Organizations | Cascade Timber (client), Alder Point Partners (administrator), Bellhaven Advisory (broker), L&L Associates (outside counsel), GreenAcre (surveyor), Moss & Lane (auditors), IRS, consultants, buyers | | Labels | Seed documents only: 1 review tag each, plus privileged and PII flags | | Benchmark | 51 distinct tasks: 10 flagship, 18 atomic/retrieval/join, 6 RLVR pilots and 17 litigation skill episodes | | Matters | 1 |

The catalog follows the linked task-table format in FrontierSWE v2. It covers 51 distinct tasks in this repository. The 19 litigation skill family IDs include two aliases for existing pilots: CTH-LIT-09 maps to CTH-PRESERVATION-001 and CTH-LIT-10 maps to CTH-CHRONOLOGY-001; those episodes are counted once.

Agent-visible prompts are in the flagship task pack. These use the evidence-based scoring rubric.

Task ID Task Category Split
CTH-AGENT-001 Acreage knowledge chain knowledge chronology train
CTH-AGENT-002 Northwest parcel reconstruction transaction reconstruction train
CTH-AGENT-003 Broker registration and authority registration authority train
CTH-AGENT-004 Post-warning program continuation autonomous investigation test_ood
CTH-AGENT-005 Clearwater transaction sequencing temporal join test_ood
CTH-AGENT-006 Privilege and work-product review privilege review test_id
CTH-AGENT-007 Whistleblower complaint and insider pool whistleblower credibility test_ood
CTH-AGENT-008 Examination and preservation timeline preservation analysis test_id
CTH-AGENT-009 Financial trail and exposure financial reconstruction test_id
CTH-AGENT-010 Matter theory memo case theory memo test_ood

Agent-visible prompts are in the atomic task pack. These use the same evidence-based rubric; difficulty and split are recorded per prompt.

Task ID Task Category Split
CTH-T1-001 Responsiveness: transaction email responsiveness train
CTH-T1-002 Responsiveness: disputed program scope responsiveness train
CTH-T1-003 Privilege: counsel risk memo privilege single train
CTH-T1-004 Privilege: bare forward privilege single train
CTH-T1-005 Kovel retention contrast kovel contrast train
CTH-T0-006 Entity and employer resolution entity resolution train
CTH-T0-007 Event date versus document date date extraction train
CTH-R-008 Acreage verification retrieval retrieval train
CTH-R-009 Broker authority retrieval retrieval train
CTH-T2-010 Registration-status contradiction contradiction train
CTH-T2-011 Investor-description change contradiction validation
CTH-T2-012 Parcel-scope competing explanations contradiction train
CTH-T2-013 Broker-gap evidence join evidence join train
CTH-T2-014 Acreage concern chronology chronology train
CTH-T2-015 Notice and subsequent action knowledge train
CTH-T1-016 Responsiveness versus privilege responsiveness validation
CTH-T2-017 Abstention on examination closure negative control validation
CTH-T2-018 Program-separation distractor control distractor control train

These tasks have deterministic verifiers. The first four are single-turn; preservation and chronology are multi-turn search/read/submit episodes. See the RLVR runner and evidence-agent interface.

Task ID Task Category Mode
CTH-DATE-001 Sent-date extraction Date extraction Single-turn smoke test
CTH-AUDIT-001 Evidence audit (75 emails, 18 questions) Cross-document reasoning Single-turn development
CTH-INVENTORY-001 Collection reconciliation (75 emails) Collection inventory Single-turn development
CTH-INVENTORY-002 Collection reconciliation (300 emails) Collection inventory Single-turn development
CTH-PRESERVATION-001 Preservation audit (CTH-LIT-09 / legal-hold) Preservation Multi-turn development
CTH-CHRONOLOGY-001 Evidence chronology (CTH-LIT-10 / chronology) Chronology Multi-turn development

These bounded development episodes have deterministic evidence and artifact checks. The registry is authoritative for task packages, pilot aliases and dependencies. See suite usage and coverage and review status. They evaluate supplied-policy exercises within one matter; the legacy preservation and chronology pilots still need separate independent oracle review.

CTH-LIT-19 tests closure readiness with a missing trigger; its inclusion does not imply that closure is authorized.

Completed October 9, 2026: 10 tasks, 10 runs per task per model, five models, 500 completed episodes, using Cascade Timber. The final ten-task dashboard (private Sites access) includes task-level means, scored counts, variability, and CSV export. Scores are deterministic verifier rewards, expressed as percentages.

Model Completed / scored Final composite Observed task mean Mean task SD (pp) Worst–best episode score Total cost (USD)
GPT-6 Astra 100 / 100 44.01% 44.01% 5.60 5.00–86.40% $86.02
Grok 4.7 100 / 99 N/A 26.76% 8.49 0.00–82.55% $137.25
Claude Opus 5.5 100 / 100 24.40% 24.40% 6.45 0.00–73.59% $350.52
GLM 5.3 Prime 100 / 99 N/A 18.63% 9.43 0.00–70.35% $106.64
Gemini 3.1 Pro Preview 100 / 100 3.98% 3.98% 5.52 0.00–40.79% $45.85

The final composite equally weights the ten task means, each over ten scored runs. The observed task mean uses available scored runs and equally weights tasks; it is incomplete for Grok and GLM, rather than a final composite. Grok chronology has 9/10 scored runs after an invalid provider response; GLM triage has 9/10 after exhausting its episode output-token budget. These two episodes remain unscored; all scored zeros are retained. Mean task SD averages the ten within-task sample standard deviations in percentage points; it is not a composite confidence interval. Worst–best ranges span scored episodes across all ten tasks.

All 500 episodes passed saved replay verification, covering 21,331 transitions. There were 498 scored episodes and 0 full passes. Total recorded provider cost was $726.28, including all recorded attempts, with no unknown-cost calls. Each episode starts with fresh context and allows up to 300 actions; requested reasoning effort is low, with 16,384 output tokens per request and 131,072 output tokens per episode. Every model received corrected model-neutral search/read/submit instructions; provider fallbacks were disabled.

These development results cover ten tasks in one synthetic matter. Preservation and chronology still need independent equivalent-evidence oracle review. Strict artifact formatting and accepted-evidence matching affect scores; these graders do not measure free-form prose quality. Equal requested reasoning effort does not imply equal compute across providers.

The earlier 50-episode chronology pilot is separate from this leaderboard; its result summary and individual episode scores remain available. Neither those pilot scores nor Gemini’s separate single validation are included in the final ten-task results.

  • You have to join documents. The corpus is built so that no single document settles an issue. Take the question of whether an acreage problem was known before the applications were filed: answering it takes option dates, application dates, ledger entries and one employee's role, all from different documents.
  • The past is never written from hindsight. Each batch is scanned for terms its authors couldn't have known yet, such as the subpoena, counsel's engagement or the . The results are inCONTINUITY_BATCH0N.md . A model can't pick up the answer from a document that was written "too early".
  • Doctrines come as contrast pairs. Privilege, work product, Kovel consultants, bare forwards and responsiveness each appear as small chains: a clear positive, a clear negative and a genuinely ambiguous case.
  • Some readings are left open on purpose. Several issues are[DISPUTED] by design, for example EMAIL-117 (a clerical error or a cover story?). The right output is to state the uncertainty, not to force a label.
  • The email looks like real collected mail. Quoted history nests one level per reply, quoting follows each sender's mail client (Outlook or Gmail), each organization has its own signature, time zones follow daylight saving, and about a third of the corpus is routine noise, decoys and near-duplicates.
  • Long threads, not just pairs. The first four expanded threads (THR-001–004) run 12–14 messages each, with forks, reply-alls that add or drop people, side forwards and unanswered questions. With the density extensions, the longest coherent reply chain is 14 deep.
  • Denser evidence chains. THR-005–007 add 32 emails and eight supporting text records for held Q4 packets, source-credit allocations to buyers, and a dated remediation pilot. Seethe expansion register for new fictional facts, counterevidence and questions deliberately left unresolved.
  • Real contracts, versioned. KW-01/02 option agreements go from drafts (DOCX) to wet-signed scans (PDF, no text layer). The Bellhaven credit purchase agreements and the pre-broker template also carry their version history. Each email carries the version that existed on its date, and the copies indata/contracts/ are byte-identical to the attachments, so hash deduplication links them.
  • Signature logos. Outlook-style orgs (Cascade Timber, L&L, Moss & Lane, Whitaker) carry an inline logo in an HTML part, and quoted signatures keep theirs, so long threads pile upimage001.png ,image002.png and so on. Inline logos aren't counted as attachments. The seed keeps its original structure.
data/emails/                          # corpus (tracked; new files need `git add -f`, see .gitignore)
  Custodians/<Name>/EMAIL-NNN_<subject>.eml
  Loadfile_Cascade_Timber.{csv,dat}   # load file: all 1,486; newest 32 have blank review labels
  README.md                           # corpus build notes (v3 realism pass)
data/emails/Exhibits/                 # 38 standalone exhibit PDFs + ../Exhibit_Manifest.csv (not attached to emails)
definitions/                          # labeling protocols: Responsiveness, ACP, Work Product, Subpoena (summons)
logs/                                 # batch ledgers, continuity reports, thread logs, demo privilege log (untracked)
benchmark/                            # tasks, hidden gold, splits, schema, rubric, GOLD_LABELS.csv
documentation/                        # AUTHORING ONLY:
  CASE_BIBLE.md  EVIDENCE_ARCS.md     #   ground-truth world model; arcs A–J, planned joins and open questions
  DOCUMENT_MANIFEST.csv               #   1,454-row plan with arc, event and evidentiary role
  CONTRADICTIONS_RESOLVED.md          #   decisions and edits for the four known contradictions
  MATTER_AGENT_TASKS.md  suggestions.md  Cascade_Timber_EML_Dataset_Plan.md
output/dataset_inconsistency_report.md  # document-consistency review + resolution log
data/contracts/                       # standalone contract collection (every version) + INDEX.csv → carrying emails
tools/exhibit_kit/                    # rebuilds the standalone exhibits from one spec
tools/doc_kit/                        # attachment builder: versioned library docs → PDF/DOCX/XLSX, scans, apply to emails, publish
  library/  plans/                    # document content + version history; email→version attach plans (ATT-00N)
tools/thread_kit/                     # thread expander: context, validate, render, rollback, scan, logos
  rules.json                          # knowledge cutoffs, participant windows, logo orgs (AUTHORING ONLY)
  logos/                              # org logos (full size + signature size)
Source Coverage Form
benchmark/hidden_gold/seed_header_labels.csv and the load file'sTAG /PRIVILEGED columns Seed EMAIL-001–250 One mixed tag per document, such as Not Responsive ,Routine ,Privileged Legal Advice ,Knowledge/Scienter orRed Flag . This isn't multi-label, and privileged documents carry no responsiveness label.
documentation/DOCUMENT_MANIFEST.csv All 1,454 rows arc ,event_id andintended_evidentiary_role . These are authoring intent, not reviewed gold.
benchmark/hidden_gold/ 28 tasks, all answerable from seed documents Required, counter, distractor and context evidence IDs; gold facts and inferences; must_include ,must_not_claim andmust_qualify lists; unknowns. The schema is inschemas/task.schema.json .
CASE_BIBLE.md /EVIDENCE_ARCS.md The whole matter Prose ground truth with [ESTABLISHED] ,[PROPOSED] ,[DISPUTED] and[INFERENCE] tags. This is the fastest source for writing new (task, evidence, target) examples.

The emails carry no label headers: the X-Decover-* headers were stripped from the original 1,454 files and are absent from the 32 additions. DocIDs live in the filenames (EMAIL-NNN_<subject>.eml) and the load file. Document-level labels for the whole original corpus are in benchmark/GOLD_LABELS.csv (machine-drafted, expert review pending). The newest 32 documents are unreviewed and not included in those labels or existing hash-pinned benchmark snapshots; their load-file label columns are blank.

  • Don't split at random. Documents in a thread or evidence chain answer each other.benchmark/splits/evidence_packages_sample.csv defines packages by family that must stay together, currently 17 of them.
  • Task splits are train ,test_id (a skill seen in training, on an unseen family) andtest_ood (held-out flagship compositions). The mapping is in thebenchmark README .
  • All of it is one matter. A held-out split measures generalization within this matter, not transfer to a new one. Building a second, independent matter for evaluation is still open.
  • Keep the model away from: the load file'sTAG andPRIVILEGED columns (withhold them; the emails themselves no longer carry label headers),CASE_BIBLE.md ,EVIDENCE_ARCS.md , the manifest'sarc /event_id /role columns, the ledgers andhidden_gold/ . The model may see the EML bodies, the Definitions files and ordinary load-file metadata.

These come from suggestions.md and from output/dataset_inconsistency_report.md. The benchmark tasks don't depend on them.

  • Resolved (2026-09-28): the § 7602 summons/subpoena framing, the two Kevin Tran identities, the two Cascade addresses and the deadline chain. Seedocumentation/CONTRADICTIONS_RESOLVED.md .
  • Resolved (2026-09-29): the standalone exhibits indata/emails/Exhibits/ now agree with the executed contracts and seed emails (Clearwater option dates and installments, buyer lot amounts, OR/WA filing-receipt dates, status footers, placeholder numbers). Rebuilt bytools/exhibit_kit/build_exhibits.py ; the resolution log is at the end of the inconsistency report.
  • Still open, by design or header-locked: - The seed parcel CSV and GreenAcre survey (EMAIL-037/039) use NW-1042…NW-1093 and different acreage pairs from the NW-01…NW-08 table in CASE_BIBLE.md §2. No crosswalk exists; don't merge them by position.
    • "Batch N" and "buyer lot N" in generated email subjects are templated labels that don't track the exhibits or the executed agreements. Join buyer records on buyer, date and amount, not on batch or lot number.
    • EMAIL-793 (9/8/22) has a subject that cites the 9/15 notice; four generated "extension chain" notes predate the summons. Both are Subject-header anachronisms, kept because headers are locked.
  • The seed parcel CSV and GreenAcre survey (EMAIL-037/039) use NW-1042…NW-1093 and different acreage pairs from the NW-01…NW-08 table in
  • The privilege log is out of step with the corpus. Most of its 1,009 rows name people who don't appear in the corpus.
  • Workbook search queries and chronology are stale. They cover only EMAIL-001–030.
  • Emails are short and convenient. Replies average about 32 new words, and some admissions are too tidy (e.g. "keep this between us"). Thethread_kit threads push back on this: they have longer replies, a validator that flags tidy admissions, and hedged readings that stay open.
  • Generation isn't reproducible. No generator or model versions or file hashes are recorded.

For the single-turn, multi-model RLVR tasks (date extraction, evidence audit, and collection reconciliation), see benchmark/rlvr/README.md. It includes an offline self-test, OpenAI/Anthropic/Gemini/OpenRouter adapters, deterministic scoring, and saved JSON run records.

  1. Parse the .eml files with any MIME library. UseDATESENT , or theDate: header, for chronology. The DocID order follows collection order, not time.
  2. Withhold the load file's TAG /PRIVILEGED columns andbenchmark/GOLD_LABELS.csv from model inputs. The.eml files are already clean: there are noX-Decover-* headers.
  3. Build examples by evidence package, not by document, and hold out whole packages.
  4. For any label beyond the seed tags, generate targets from CASE_BIBLE.md and the benchmark gold, and have an expert review them. Treat[PROPOSED] facts as unknowns until the documents that support them exist.
  5. Score with benchmark/eval/scoring_rubric.md . It covers evidence recall, join correctness, counter-evidence, restraint on uncertain points, citation accuracy and stopping efficiency.
  • All 1,400 planned documents are rendered. Thread expansion has added 54 more: THR-001 Clearwater scouting (EMAIL-1401–1412), THR-002 Bellhaven registration (1421–1434), THR-003 Q4 push-through (1441–1454) and THR-004 the $2.4M estimate (1461–1474). Their specs and reports are in Logs/threads/ .
  • The authorized density pass adds THR-005–007 (EMAIL-1475–1506), bringing the total to 1,486 messages. All original emails remain byte-identical. Story specs, a metadata synchronization/checking script and validation results are in tools/thread_kit/expansions/ ;the authoring register explains the additions and their limits.
  • The top fixes for training readiness, from suggestions.md :
    1. Move labels out of the headers.
    2. Write 150–200 examples reviewed by experts, with multi-label responsiveness, privilege basis, chronology and grounded QA, including examples where the right answer is to abstain.
    3. Resolve the contradictions listed above.
    4. Fix the summons/subpoena mislabel.
    5. Add a second matter for evaluation, plus hashes.
── more in #ai-agents 4 stories · sorted by recency
── more on @decoverai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/building-an-rl-envir…] indexed:0 read:14min 2026-10-09 · —