{"slug": "building-an-rl-environment-and-dataset-for-legal", "title": "Building an RL environment and dataset for legal", "summary": "DecoverAI built a legal-agent evaluation benchmark around a fictional matter, USA v. Cascade Timber Holdings, Inc., comprising 1,486 RFC 5322 .eml documents spanning November 2021 to December 2023 across 26 mailboxes, with 51 distinct tasks split into 10 flagship, 18 atomic/retrieval/join, 6 RLVR pilots and 17 litigation skill episodes. DecoverAI's internal assessment in suggestions.md scores the dataset Q ≈ 49/100, in the band marked \"blocked for training\", citing answer leakage through headers, few supervised targets, contradictory facts and only one matter. The dataset is positioned as sound for demos and as an evaluation seed, with labels limited to seed documents (one review tag each plus privileged and PII flags).", "body_md": "A legal-agent evaluation benchmark built around a Matter, *USA v. Cascade Timber Holdings, Inc.*, built by DecoverAI. It is\nan email corpus with planted evidence chains, designed for training and evaluating models on\n**legal evidence reasoning**: responsiveness, privilege, chronology, knowledge analysis and\njoins across several documents. All companies, people and events are fictional.\n\n**Not ready for training.** The internal assessment ([`suggestions.md`](https://github.com/decover-tech/frontier-legal/blob/main/suggestions.md))\nscores it Q ≈ 49/100, in the band marked \"blocked for training\". It's sound as a demo and\nan evaluation seed, but the answers leak through headers, there are few supervised\ntargets, some facts contradict each other, and there's only one matter. Read\n[Known issues](#known-issues) before you use it.\n\n| Documents rendered | 1,486 RFC 5322 `.eml` (250 seed, 1,150 generated in batches, 86 in 7 threads from`thread_kit` ) | \n| Documents planned | 1,486 ( `DOCUMENT_MANIFEST.csv` ; thread rows have`status=thread-expansion` ) | \n| Time span | Nov 2021 – Dec 2023 | \n| Custodians | 26 mailboxes | \n| Organizations | Cascade Timber (client), Alder Point Partners (administrator), Bellhaven Advisory (broker), L&L Associates (outside counsel), GreenAcre (surveyor), Moss & Lane (auditors), IRS, consultants, buyers | \n| Labels | Seed documents only: 1 review tag each, plus privileged and PII flags | \n| Benchmark | 51 distinct tasks: 10 flagship, 18 atomic/retrieval/join, 6 RLVR pilots and 17 litigation skill episodes | \n| Matters | 1 | \n\nThe catalog follows the linked task-table format in [FrontierSWE v2](https://github.com/Proximal-Labs/frontier-swe-v2).\nIt covers **51 distinct tasks** in this repository. The 19 litigation skill family IDs\ninclude two aliases for existing pilots: `CTH-LIT-09` maps to `CTH-PRESERVATION-001`\nand `CTH-LIT-10` maps to `CTH-CHRONOLOGY-001`; those episodes are counted once.\n\nAgent-visible prompts are in [the flagship task pack](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl).\nThese use the [evidence-based scoring rubric](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/eval/scoring_rubric.md).\n\n| Task ID | Task | Category | Split | \n|---|---|---|---|\n| CTH-AGENT-001 | [Acreage knowledge chain](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl#L1) | knowledge chronology | `train` | \n| CTH-AGENT-002 | [Northwest parcel reconstruction](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl#L2) | transaction reconstruction | `train` | \n| CTH-AGENT-003 | [Broker registration and authority](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl#L3) | registration authority | `train` | \n| CTH-AGENT-004 | [Post-warning program continuation](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl#L4) | autonomous investigation | `test_ood` | \n| CTH-AGENT-005 | [Clearwater transaction sequencing](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl#L5) | temporal join | `test_ood` | \n| CTH-AGENT-006 | [Privilege and work-product review](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl#L6) | privilege review | `test_id` | \n| CTH-AGENT-007 | [Whistleblower complaint and insider pool](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl#L7) | whistleblower credibility | `test_ood` | \n| CTH-AGENT-008 | [Examination and preservation timeline](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl#L8) | preservation analysis | `test_id` | \n| CTH-AGENT-009 | [Financial trail and exposure](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl#L9) | financial reconstruction | `test_id` | \n| CTH-AGENT-010 | [Matter theory memo](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/flagship_prompts.jsonl#L10) | case theory memo | `test_ood` | \n\nAgent-visible prompts are in [the atomic task pack](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl).\nThese use the same evidence-based rubric; difficulty and split are recorded per prompt.\n\n| Task ID | Task | Category | Split | \n|---|---|---|---|\n| CTH-T1-001 | [Responsiveness: transaction email](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L1) | responsiveness | `train` | \n| CTH-T1-002 | [Responsiveness: disputed program scope](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L2) | responsiveness | `train` | \n| CTH-T1-003 | [Privilege: counsel risk memo](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L3) | privilege single | `train` | \n| CTH-T1-004 | [Privilege: bare forward](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L4) | privilege single | `train` | \n| CTH-T1-005 | [Kovel retention contrast](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L5) | kovel contrast | `train` | \n| CTH-T0-006 | [Entity and employer resolution](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L6) | entity resolution | `train` | \n| CTH-T0-007 | [Event date versus document date](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L7) | date extraction | `train` | \n| CTH-R-008 | [Acreage verification retrieval](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L8) | retrieval | `train` | \n| CTH-R-009 | [Broker authority retrieval](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L9) | retrieval | `train` | \n| CTH-T2-010 | [Registration-status contradiction](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L10) | contradiction | `train` | \n| CTH-T2-011 | [Investor-description change](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L11) | contradiction | `validation` | \n| CTH-T2-012 | [Parcel-scope competing explanations](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L12) | contradiction | `train` | \n| CTH-T2-013 | [Broker-gap evidence join](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L13) | evidence join | `train` | \n| CTH-T2-014 | [Acreage concern chronology](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L14) | chronology | `train` | \n| CTH-T2-015 | [Notice and subsequent action](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L15) | knowledge | `train` | \n| CTH-T1-016 | [Responsiveness versus privilege](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L16) | responsiveness | `validation` | \n| CTH-T2-017 | [Abstention on examination closure](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L17) | negative control | `validation` | \n| CTH-T2-018 | [Program-separation distractor control](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/tasks/tier1_tier2_sample.jsonl#L18) | distractor control | `train` | \n\nThese tasks have deterministic verifiers. The first four are single-turn;\npreservation and chronology are multi-turn search/read/submit episodes.\nSee the [RLVR runner](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/README.md) and [evidence-agent interface](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/AGENT_TASKS.md).\n\n| Task ID | Task | Category | Mode | \n|---|---|---|---|\n| CTH-DATE-001 | [Sent-date extraction](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/tasks/CTH-DATE-001.json) | Date extraction | Single-turn smoke test | \n| CTH-AUDIT-001 | [Evidence audit (75 emails, 18 questions)](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/tasks/CTH-AUDIT-001.json) | Cross-document reasoning | Single-turn development | \n| CTH-INVENTORY-001 | [Collection reconciliation (75 emails)](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/tasks/CTH-INVENTORY-001.json) | Collection inventory | Single-turn development | \n| CTH-INVENTORY-002 | [Collection reconciliation (300 emails)](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/tasks/CTH-INVENTORY-002.json) | Collection inventory | Single-turn development | \n| CTH-PRESERVATION-001 | [Preservation audit (CTH-LIT-09 / legal-hold)](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/preservation/CTH-PRESERVATION-001.json) | Preservation | Multi-turn development | \n| CTH-CHRONOLOGY-001 | [Evidence chronology (CTH-LIT-10 / chronology)](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/chronology/CTH-CHRONOLOGY-001.json) | Chronology | Multi-turn development | \n\nThese bounded development episodes have deterministic evidence and artifact checks.\nThe [registry](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/litigation_skills/registry.json) is authoritative for task packages,\npilot aliases and dependencies. See [suite usage](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/litigation_skills/README.md)\nand [coverage and review status](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/litigation_skills/COVERAGE.md).\nThey evaluate supplied-policy exercises within one matter; the legacy preservation\nand chronology pilots still need separate independent oracle review.\n\n`CTH-LIT-19` tests closure readiness with a missing trigger; its inclusion does not imply that closure is authorized.\n\nCompleted October 9, 2026: **10 tasks, 10 runs per task per model, five models,\n500 completed episodes**, using Cascade Timber. The\n[final ten-task dashboard](https://cascade-timber-final-benchmark.decoverhq-in-8411.chatgpt.site)\n(private Sites access) includes task-level means, scored counts, variability, and CSV export.\nScores are deterministic verifier rewards, expressed as percentages.\n\n| Model | Completed / scored | Final composite | Observed task mean | Mean task SD (pp) | Worst–best episode score | Total cost (USD) | \n|---|---|---|---|---|---|---|\n| GPT-6 Astra | 100 / 100 | **44.01%** | 44.01% | 5.60 | 5.00–86.40% | $86.02 | \n| Grok 4.7 | 100 / 99 | **N/A** | 26.76% | 8.49 | 0.00–82.55% | $137.25 | \n| Claude Opus 5.5 | 100 / 100 | **24.40%** | 24.40% | 6.45 | 0.00–73.59% | $350.52 | \n| GLM 5.3 Prime | 100 / 99 | **N/A** | 18.63% | 9.43 | 0.00–70.35% | $106.64 | \n| Gemini 3.1 Pro Preview | 100 / 100 | **3.98%** | 3.98% | 5.52 | 0.00–40.79% | $45.85 | \n\nThe final composite equally weights the ten task means, each over ten scored runs.\nThe observed task mean uses available scored runs and equally weights tasks; it is\n**incomplete for Grok and GLM**, rather than a final composite. Grok chronology has\n9/10 scored runs after an invalid provider response; GLM triage has 9/10 after\nexhausting its episode output-token budget. These two episodes remain unscored;\nall scored zeros are retained. Mean task SD averages the ten within-task sample\nstandard deviations in percentage points; it is not a composite confidence interval.\nWorst–best ranges span scored episodes across all ten tasks.\n\nAll **500 episodes passed saved replay verification**, covering **21,331 transitions**.\nThere were **498 scored episodes and 0 full passes**. Total recorded provider cost was\n**$726.28**, including all recorded attempts, with no unknown-cost calls.\nEach episode starts with fresh context and allows up to 300 actions; requested\nreasoning effort is low, with 16,384 output tokens per request and 131,072 output\ntokens per episode. Every model received corrected model-neutral search/read/submit\ninstructions; provider fallbacks were disabled.\n\nThese development results cover ten tasks in one synthetic matter. Preservation and chronology still need independent equivalent-evidence oracle review. Strict artifact formatting and accepted-evidence matching affect scores; these graders do not measure free-form prose quality. Equal requested reasoning effort does not imply equal compute across providers.\n\nThe earlier 50-episode chronology pilot is separate from this leaderboard; its\n[result summary](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/chronology/results/chronology-10x5-2026-10-09.json) and\n[individual episode scores](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/chronology/results/chronology-10x5-2026-10-09-episodes.csv)\nremain available. Neither those pilot scores nor Gemini’s separate single validation\nare included in the final ten-task results.\n\n- **You have to join documents.** The corpus is built so that no single document settles an\nissue. Take the question of whether an acreage problem was known before the applications\nwere filed: answering it takes option dates, application dates, ledger entries and one\nemployee's role, all from different documents.\n- **The past is never written from hindsight.** Each batch is scanned for terms its authors\ncouldn't have known yet, such as the subpoena, counsel's engagement or the pause. The\nresults are in`CONTINUITY_BATCH0N.md` . A model can't pick up the answer from a document\nthat was written \"too early\".\n- **Doctrines come as contrast pairs.** Privilege, work product, Kovel consultants, bare\nforwards and responsiveness each appear as small chains: a clear positive, a clear negative\nand a genuinely ambiguous case.\n- **Some readings are left open on purpose.** Several issues are`[DISPUTED]` by design,\nfor example EMAIL-117 (a clerical error or a cover story?). The right output is to state the\nuncertainty, not to force a label.\n- **The email looks like real collected mail.** Quoted history nests one level per reply,\nquoting follows each sender's mail client (Outlook or Gmail), each organization has its own\nsignature, time zones follow daylight saving, and about a third of the corpus is routine\nnoise, decoys and near-duplicates.\n- **Long threads, not just pairs.** The first four expanded threads (THR-001–004) run 12–14 messages\neach, with forks, reply-alls that add or drop people, side forwards and unanswered\nquestions. With the density extensions, the longest coherent reply chain is 14 deep.\n- **Denser evidence chains.** THR-005–007 add 32 emails and eight supporting text records\nfor held Q4 packets, source-credit allocations to buyers, and a dated remediation pilot.\nSee[the expansion register](https://github.com/decover-tech/frontier-legal/blob/main/documentation/DENSITY_EXPANSION.md) for new fictional\nfacts, counterevidence and questions deliberately left unresolved.\n- **Real contracts, versioned.** KW-01/02 option agreements go from drafts (DOCX) to wet-signed\nscans (PDF, no text layer). The Bellhaven credit purchase agreements and the pre-broker\ntemplate also carry their version history. Each email carries the version that existed on\nits date, and the copies in`data/contracts/` are byte-identical to the attachments, so\nhash deduplication links them.\n- **Signature logos.** Outlook-style orgs (Cascade Timber, L&L, Moss & Lane, Whitaker) carry\nan inline logo in an HTML part, and quoted signatures keep theirs, so long threads pile up`image001.png` ,`image002.png` and so on. Inline logos aren't counted as attachments. The\nseed keeps its original structure.\n\n```\ndata/emails/                          # corpus (tracked; new files need `git add -f`, see .gitignore)\n  Custodians/<Name>/EMAIL-NNN_<subject>.eml\n  Loadfile_Cascade_Timber.{csv,dat}   # load file: all 1,486; newest 32 have blank review labels\n  README.md                           # corpus build notes (v3 realism pass)\ndata/emails/Exhibits/                 # 38 standalone exhibit PDFs + ../Exhibit_Manifest.csv (not attached to emails)\ndefinitions/                          # labeling protocols: Responsiveness, ACP, Work Product, Subpoena (summons)\nlogs/                                 # batch ledgers, continuity reports, thread logs, demo privilege log (untracked)\nbenchmark/                            # tasks, hidden gold, splits, schema, rubric, GOLD_LABELS.csv\ndocumentation/                        # AUTHORING ONLY:\n  CASE_BIBLE.md  EVIDENCE_ARCS.md     #   ground-truth world model; arcs A–J, planned joins and open questions\n  DOCUMENT_MANIFEST.csv               #   1,454-row plan with arc, event and evidentiary role\n  CONTRADICTIONS_RESOLVED.md          #   decisions and edits for the four known contradictions\n  MATTER_AGENT_TASKS.md  suggestions.md  Cascade_Timber_EML_Dataset_Plan.md\noutput/dataset_inconsistency_report.md  # document-consistency review + resolution log\ndata/contracts/                       # standalone contract collection (every version) + INDEX.csv → carrying emails\ntools/exhibit_kit/                    # rebuilds the standalone exhibits from one spec\ntools/doc_kit/                        # attachment builder: versioned library docs → PDF/DOCX/XLSX, scans, apply to emails, publish\n  library/  plans/                    # document content + version history; email→version attach plans (ATT-00N)\ntools/thread_kit/                     # thread expander: context, validate, render, rollback, scan, logos\n  rules.json                          # knowledge cutoffs, participant windows, logo orgs (AUTHORING ONLY)\n  logos/                              # org logos (full size + signature size)\n```\n\n| Source | Coverage | Form | \n|---|---|---|\n| `benchmark/hidden_gold/seed_header_labels.csv` and the load file's`TAG` /`PRIVILEGED` columns | Seed EMAIL-001–250 | One mixed tag per document, such as `Not Responsive` ,`Routine` ,`Privileged Legal Advice` ,`Knowledge/Scienter` or`Red Flag` . This isn't multi-label, and privileged documents carry no responsiveness label. | \n| `documentation/DOCUMENT_MANIFEST.csv` | All 1,454 rows | `arc` ,`event_id` and`intended_evidentiary_role` . These are authoring intent, not reviewed gold. | \n| `benchmark/hidden_gold/` | 28 tasks, all answerable from seed documents | Required, counter, distractor and context evidence IDs; gold facts and inferences; `must_include` ,`must_not_claim` and`must_qualify` lists; unknowns. The schema is in`schemas/task.schema.json` . | \n| `CASE_BIBLE.md` /`EVIDENCE_ARCS.md` | The whole matter | Prose ground truth with `[ESTABLISHED]` ,`[PROPOSED]` ,`[DISPUTED]` and`[INFERENCE]` tags. This is the fastest source for writing new (task, evidence, target) examples. | \n\nThe emails carry no label headers: the `X-Decover-*` headers were stripped from the original 1,454 files and are absent from the 32 additions. DocIDs\nlive in the filenames (`EMAIL-NNN_<subject>.eml`) and the load file. Document-level labels for the whole\noriginal corpus are in `benchmark/GOLD_LABELS.csv` (machine-drafted, expert review pending).\nThe newest 32 documents are unreviewed and not included in those labels or existing\nhash-pinned benchmark snapshots; their load-file label columns are blank.\n\n- **Don't split at random.** Documents in a thread or evidence chain answer each other.`benchmark/splits/evidence_packages_sample.csv` defines packages by family\nthat must stay together, currently 17 of them.\n- Task splits are `train` ,`test_id` (a skill seen in training, on an unseen family) and`test_ood` (held-out flagship compositions). The mapping is in the[benchmark README](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/README.md) .\n- **All of it is one matter.** A held-out split measures generalization within this matter,\nnot transfer to a new one. Building a second, independent matter for evaluation is still\nopen.\n- **Keep the model away from:** the load file's`TAG` and`PRIVILEGED` columns (withhold them; the emails themselves no longer\ncarry label headers),`CASE_BIBLE.md` ,`EVIDENCE_ARCS.md` , the manifest's`arc` /`event_id` /role columns, the\nledgers and`hidden_gold/` . The model may see the EML bodies, the Definitions files and\nordinary load-file metadata.\n\nThese come from `suggestions.md` and from `output/dataset_inconsistency_report.md`. The benchmark tasks\ndon't depend on them.\n\n- **Resolved (2026-09-28):** the § 7602 summons/subpoena framing, the two Kevin Tran identities,\nthe two Cascade addresses and the deadline chain. See`documentation/CONTRADICTIONS_RESOLVED.md` .\n- **Resolved (2026-09-29):** the standalone exhibits in`data/emails/Exhibits/` now agree with the\nexecuted contracts and seed emails (Clearwater option dates and installments, buyer lot amounts,\nOR/WA filing-receipt dates, status footers, placeholder numbers). Rebuilt by`tools/exhibit_kit/build_exhibits.py` ; the resolution log is at the end of the inconsistency report.\n- **Still open, by design or header-locked:**  - The seed parcel CSV and GreenAcre survey (EMAIL-037/039) use NW-1042…NW-1093 and different\nacreage pairs from the NW-01…NW-08 table in `CASE_BIBLE.md` §2. No crosswalk exists; don't\nmerge them by position.\n  - \"Batch N\" and \"buyer lot N\" in generated email subjects are templated labels that don't track the exhibits or the executed agreements. Join buyer records on buyer, date and amount, not on batch or lot number.\n  - EMAIL-793 (9/8/22) has a subject that cites the 9/15 notice; four generated \"extension chain\" notes predate the summons. Both are Subject-header anachronisms, kept because headers are locked.\n- The seed parcel CSV and GreenAcre survey (EMAIL-037/039) use NW-1042…NW-1093 and different\nacreage pairs from the NW-01…NW-08 table in \n- **The privilege log is out of step with the corpus.** Most of its 1,009 rows name people who\ndon't appear in the corpus.\n- **Workbook search queries and chronology are stale.** They cover only EMAIL-001–030.\n- **Emails are short and convenient.** Replies average about 32 new words, and some\nadmissions are too tidy (e.g. \"keep this between us\"). The`thread_kit` threads push back\non this: they have longer replies, a validator that flags tidy admissions, and hedged\nreadings that stay open.\n- **Generation isn't reproducible.** No generator or model versions or file hashes are\nrecorded.\n\nFor the single-turn, multi-model RLVR tasks (date extraction, evidence audit, and\ncollection reconciliation), see\n[`benchmark/rlvr/README.md`](https://github.com/decover-tech/frontier-legal/blob/main/benchmark/rlvr/README.md). It includes an offline\nself-test, OpenAI/Anthropic/Gemini/OpenRouter adapters, deterministic scoring, and saved\nJSON run records.\n\n1. Parse the `.eml` files with any MIME library. Use`DATESENT` , or the`Date:` header, for\nchronology. The DocID order follows collection order, not time.\n2. Withhold the load file's `TAG` /`PRIVILEGED` columns and`benchmark/GOLD_LABELS.csv` from model inputs. The`.eml` files are already clean: there are no`X-Decover-*` headers.\n3. Build examples by evidence package, not by document, and hold out whole packages.\n4. For any label beyond the seed tags, generate targets from `CASE_BIBLE.md` and the\nbenchmark gold, and have an expert review them. Treat`[PROPOSED]` facts as unknowns\nuntil the documents that support them exist.\n5. Score with `benchmark/eval/scoring_rubric.md` . It covers evidence recall,\njoin correctness, counter-evidence, restraint on uncertain points, citation accuracy and\nstopping efficiency.\n\n- All 1,400 planned documents are rendered. Thread expansion has added 54 more: THR-001\nClearwater scouting (EMAIL-1401–1412), THR-002 Bellhaven registration (1421–1434), THR-003\nQ4 push-through (1441–1454) and THR-004 the $2.4M estimate (1461–1474). Their specs and\nreports are in `Logs/threads/` .\n- The authorized density pass adds THR-005–007 (EMAIL-1475–1506), bringing the total to\n1,486 messages. All original emails remain byte-identical. Story specs, a metadata\nsynchronization/checking script and validation results are in\n`tools/thread_kit/expansions/` ;[the authoring register](https://github.com/decover-tech/frontier-legal/blob/main/documentation/DENSITY_EXPANSION.md) explains the additions and their limits.\n- The top fixes for training readiness, from `suggestions.md` :\n  1. Move labels out of the headers.\n  2. Write 150–200 examples reviewed by experts, with multi-label responsiveness, privilege basis, chronology and grounded QA, including examples where the right answer is to abstain.\n  3. Resolve the contradictions listed above.\n  4. Fix the summons/subpoena mislabel.\n  5. Add a second matter for evaluation, plus hashes.", "url": "https://wpnews.pro/news/building-an-rl-environment-and-dataset-for-legal", "canonical_source": "https://github.com/decover-tech/frontier-legal", "published_at": "2026-10-09 23:12:52+00:00", "updated_at": "2026-10-09 23:26:11.657338+00:00", "lang": "en", "topics": ["ai-agents", "ai-research", "large-language-models", "ai-safety"], "entities": ["DecoverAI", "USA v. Cascade Timber Holdings, Inc.", "Cascade Timber", "Alder Point Partners", "Bellhaven Advisory", "L&L Associates", "GreenAcre", "Moss & Lane"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-an-rl-environment-and-dataset-for-legal", "markdown": "https://wpnews.pro/news/building-an-rl-environment-and-dataset-for-legal.md", "text": "https://wpnews.pro/news/building-an-rl-environment-and-dataset-for-legal.txt", "jsonld": "https://wpnews.pro/news/building-an-rl-environment-and-dataset-for-legal.jsonld"}}