As Lyle Lanley might say, a law firm with AI is a little like the mule with a spinning wheel — no one knows how he got it and danged if he knows how to use it. Firms are falling all over themselves to be able to tell the world they have cutting edge AI capabilities, but ask a lawyer why its AI assistant runs on the most expensive frontier model on the menu and the answer lands somewhere between “it’s the best one” and a shrug. Ask what it should actually cost to get the right answer and you get the shrug without the preamble.
That’s a survivable state of affairs when everything gets priced out as a flat-rate seat license and the compute bill becomes somebody else’s problem. That won’t cut it forever and when consumption pricing finally takes over, every question a lawyer fires at a chatbot generates a line item, and nobody in the building can tell you whether that line item is a penny or 70 cents. As NetDocuments CEO Josh Baxter put it in announcing the company’s new benchmark product, the industry is currently oscillating between “spending without limits and cutting without strategy.”
NetDocuments published the Legal Context Engineering Benchmark ahead of ILTACON, and the proposition is pretty straigthforward. Three things determine whether a legal AI agent works: the model, the harness wrapped around it, and the context it can reach while it works. Existing benchmarks offer an incomplete picture. The goal of the NetDocuments product is to freeze the model and harness, run the same 300 questions across 10 real matters twice — once with only search and retrieval, once with the Legal Context Graph switched on — and measure what moves.
[
From ‘Vendor’ To ‘Partner’: How LexisNexis Is Deepening Law Firm Relationships ](https://abovethelaw.com/2026/08/white-glove-service-in-ai-era/)
The company is emphasizing ‘white glove service’ in the AI era. Here’s what the initiative is delivering for clients.
What this reveals is the key detail other benchmarks miss:
Accuracy alone would be the wrong scorecard, because accuracy can nearly always be bought with more spending: a bigger model, more reasoning, more retrieval.
NetDocuments, as a document management system, is focused on the role context plays. Their point with this exercise isn’t to prove that users can spend less without forfeiting accuracy, but that excellent, well-managed context makes AI run more efficiently — which can either be used to save money or reinvested into chasing more accuracy.
But along the way, the results lent support to the idea that while the market obsesses over models achieving accuracy-above-all, as a practical matter, if “human-in-the-loop” means anything at all, it means someone should be there to intervene to address a couple percent accuracy dip when it saves the user hundreds of thousands of dollars.
[
Beyond Recall And Precision: Strategic Oversight Of AI Managed Review ](https://abovethelaw.com/2026/08/beyond-recall-and-precision-strategic-oversight-of-ai-managed-review/)
AI can accelerate document review, but metrics alone aren't enough. Learn how Managing Attorney oversight turns AI-driven review into a more adaptive, strategic and defensible process.
Trading off accuracy gets some people around these legal tech conversations skittish. But think of it this way — assigning some tasks to a first-year rather than a mid-level trades accuracy too, but we do it because we know it’s cheaper and the sacrifice is something we can easily address on the back end. Why isn’t that how we think about AI?
The headline result is that the context layer cuts it roughly in half. They call the ratio the Context Value Ratio and it’s what shows how much the firm can accomplish with more robust context. According to the company, using their process to enrich context and prevent models from casually wandering around millions of tokens of unfocused context, token consumption per answer fell 52 percent.
NetDocuments ran the whole benchmark on three tiers of the same model family. On the economy tier, answering all 300 questions cost $4.35 and produced 191.9 fully accurate answers. On the frontier tier, the same 300 questions cost $151.10 and produced 221.8 correct answers. Thirty-five times the spend to get 30 more right answers out of three hundred.
The economy model availed of the NetDocuments context graph gets 195.3 questions right for $2.83. Comparing it to the 221.8 answers of the frontier model with the benefit of the context graph, that means 26 additional correct answers cost about $148 — roughly $5.60 apiece, against an average of a penny and a half.
On bet-the-company questions $5.60 is worth it. But most legal queries do not carry final bet-the-company stakes and wasting $5.60 on them is a decision that nobody is actually aware enough to be making. Firms are throwing money at the frontier tier the way they buy Herman Miller chairs, and the report all but says so, noting that the gains from context are front-loaded across tiers and that the mid tier “is where most production work will actually run.”
At a press briefing, NetDocuments projecting a firm could save around $1 million a year in token savings. That estimate rested on a 2,000-lawyer firm asking about four million questions a year, which works out to eight AI queries per professional per working day.
That seems reasonable — and it might even undercount usage once the agents takeover and start churning.
Joe Patrice is a senior editor at Above the Law and co-host of Thinking Like A Lawyer. Feel free to email any tips, questions, or comments. Follow him on Twitter or Bluesky if you’re interested in law, politics, and a healthy dose of college sports news.