cd /news/large-language-models/i-counted-tokens-for-the-same-data-i… · home › topics › large-language-models › article
[ARTICLE · art-144025] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

I counted tokens for the same data in 7 formats. Pretty JSON costs 3 CSV

A developer measured token counts for the same 20-row, 5-field product table serialized in seven formats using the o200k_base tokenizer behind GPT-4o, finding pretty-printed JSON costs roughly 3× the tokens of CSV (884 vs 300) and XML 3.63× (1,088 tokens). The same data minified as JSON came to 525 tokens and YAML to 649, while TSV (296) and CSV (300) were cheapest; the older cl100k_base tokenizer produced nearly identical results within 2%. At 30,000 monthly requests with a 20-row table attached, the developer estimated CSV costs about $18/month versus $53 for pretty JSON at $2 per million input tokens, and released a browser-based token counter that runs the same tokenizer locally.

by read2 min views3 publishedOct 2, 2026

Most of us paste data into LLM prompts as JSON, often straight from JSON.stringify(data, null, 2). I had never checked what that costs in tokens, so I took one table, wrote it in seven formats and counted.

Short version: pretty-printed JSON uses about 3× the tokens of CSV.

A product table: 20 rows, 5 fields (id, name, price, in_stock, category). One row in CSV:

1001,Wireless Mouse,9.99,false,electronics

Same data, seven formats, counted with o200k_base, the tokenizer behind GPT-4o and later OpenAI models:

import { getEncoding } from "js-tiktoken";
const enc = getEncoding("o200k_base");
const count = (s) => enc.encode(s).length;

count(csv);                            // 300
count(JSON.stringify(rows));           // 525
count(JSON.stringify(rows, null, 2));  // 884
Format Tokens vs CSV
TSV 296 0.99×
CSV 300 1.00×
Markdown table 373 1.24×
JSON, minified 525 1.75×
YAML 649 2.16×
JSON, pretty (2-space) 884 2.95×
XML 1,088 3.63×

The older cl100k_base tokenizer gave nearly identical numbers, within 2% for every format.

"name":, "price": twenty times. CSV writes them once, in the header. YAML is the odd one: fewer characters than minified JSON, more tokens, because every field gets its own line and its own key.

Coding agents send a lot of code, so I tried a few things:

Test Result
16-line Python file: 4 spaces vs 2 spaces vs tabs 135 / 135 / 133 tokens, basically no difference
Same file without its one-line docstring 135 → 123 (−9%)
10-line JS function, minified 92 → 47 (−49%)
One UUID 18 tokens

subtotal and taxRate, which is exactly what helps the model understand the code. Dropping unused fields, null fields and extra decimal places helps too.

Say you attach a 20-row table to every request, 1,000 requests a day, 30,000 a month, at $2 per million input tokens:

Format Tokens / month Cost / month
CSV 9.0M $18
JSON, minified 15.8M $32
JSON, pretty 26.5M $53
XML 32.6M $65

Invisible per request, real for a product, and it scales with bigger tables, RAG results and long API responses.

I built a free token counter for this. It runs the same o200k tokenizer in your browser, so nothing you paste is uploaded. Paste your data in two formats and compare tokens and cost per model.

Full write-up with more detail: JSON vs YAML vs CSV: which format uses the fewest tokens?

── more in #large-language-models 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-counted-tokens-for…] indexed:0 read:2min 2026-10-02 · —