cd /news/ai-agents/agent-token-spend-distribution-codin… · home › topics › ai-agents › article
[ARTICLE · art-141337] src=lexifina.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Agent token spend distribution: coding versus legal work

A token-spend breakdown comparing Cursor's coding agents with Lexifina's legal agents shows system and tool definitions consume 34.4% of Lexifina's spend versus 17% for Cursor, while Cursor allocates 20% to read-tool outputs and 18% to other tool outputs against Lexifina's 9.0% and 5.3%. Lexifina attributes its higher system-instruction share to more exhaustive user steering and cloud-based agents, and reports reasoning at 24.8% of spend versus Cursor's 10%. Lexifina says it keeps recurring instructions and tool descriptions in consistent order and uses cache boundaries to separate stable blocks from growing conversation to make context reusable.

by read2 min views1 publishedSep 28, 2026
Agent token spend distribution: coding versus legal work
Image: source

Coding agents and legal agents work on documents with a very different composition. Coding functions can span multiple pages, have adverse effects, and complex dependencies. So can legal clauses and sections, but often in a very different way. For these tasks, we present some aggregate token consumptions, approximated by various usage bins.

Where agent spending goes #

Source Cursor Lexifina
System & tool defs 17% 34.4%
Read tool outputs 20% 9.0%
Other tool outputs 18% 5.3%
Tool call arguments 21% 17.6%
Reasoning 10% 24.8%
Skills & plugins 8% 7.1%
Assistant text 4% 1.6%
User text 3% 0.37%

Some insights into this utilisation #

System instructions explain how the agent should work. Tool definitions describe the actions it can take and the information each action needs. Lexifina spends more on this, presumably to provide more exhaustive user steering, and because our agents are primarily cloud-based.

Lexifina loads some tool definitions on demand. Definitions available to the model still occupy space in its request, alongside system instructions. Caching can reduce the cost of reading that text again, but it also puts pressure on the active context width. Rather, many secondary and tertiary tool calls are "delegated", advertised and discoverable to the agent but not actively loaded into context.

Tool outputs include the information returned by an agent after an action. Cursor assigns 20% of spending to read-tool outputs and 18% to other tool outputs. Lexifina's corresponding shares here are 9.0% and 5.3%. The primary user action involves redlining (marking up a paragraph) rather than raw document generation, and coding also generally consumes more context outside of complex legal work.

Skills and plugins account for 7.1% of spending here. A skill that is hard to discover can cause retries, so large definitions help the agent find the right action.

For Lexifina, useful efficiency means lower total cost across drafting, review, research, and matter summaries while preserving source coverage and reducing the corrections a lawyer must make. Time spent, failed actions, retries, and all delegated work belong in that assessment. A shorter request that leads to repeated repairs can cost more overall. In terms of raw token optimisation, we strive to make context reusable by keeping recurring instructions and tool descriptions in a consistent order, and by using cache boundaries to separate stable blocks from a growing conversation.

── more in #ai-agents 4 stories · sorted by recency
── more on @cursor 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agent-token-spend-di…] indexed:0 read:2min 2026-09-28 · —