Most of us paste data into LLM prompts as JSON, often straight from JSON.stringify(data, null, 2). I had never checked what that costs in tokens, so I took one table, wrote it in seven formats and counted.
Short version: pretty-printed JSON uses about 3× the tokens of CSV.
A product table: 20 rows, 5 fields (id, name, price, in_stock, category). One row in CSV:
1001,Wireless Mouse,9.99,false,electronics
Same data, seven formats, counted with o200k_base, the tokenizer behind GPT-4o and later OpenAI models:
import { getEncoding } from "js-tiktoken";
const enc = getEncoding("o200k_base");
const count = (s) => enc.encode(s).length;
count(csv); // 300
count(JSON.stringify(rows)); // 525
count(JSON.stringify(rows, null, 2)); // 884
| Format | Tokens | vs CSV |
|---|---|---|
| TSV | 296 | 0.99× |
| CSV | 300 | 1.00× |
| Markdown table | 373 | 1.24× |
| JSON, minified | 525 | 1.75× |
| YAML | 649 | 2.16× |
| JSON, pretty (2-space) | 884 | 2.95× |
| XML | 1,088 | 3.63× |
The older cl100k_base tokenizer gave nearly identical numbers, within 2% for every format.
"name":, "price": twenty times. CSV writes them once, in the header.
YAML is the odd one: fewer characters than minified JSON, more tokens, because every field gets its own line and its own key.
Coding agents send a lot of code, so I tried a few things:
| Test | Result |
|---|---|
| 16-line Python file: 4 spaces vs 2 spaces vs tabs | 135 / 135 / 133 tokens, basically no difference |
| Same file without its one-line docstring | 135 → 123 (−9%) |
| 10-line JS function, minified | 92 → 47 (−49%) |
| One UUID | 18 tokens |
subtotal and taxRate, which is exactly what helps the model understand the code.
Dropping unused fields, null fields and extra decimal places helps too.
Say you attach a 20-row table to every request, 1,000 requests a day, 30,000 a month, at $2 per million input tokens:
| Format | Tokens / month | Cost / month |
|---|---|---|
| CSV | 9.0M | $18 |
| JSON, minified | 15.8M | $32 |
| JSON, pretty | 26.5M | $53 |
| XML | 32.6M | $65 |
Invisible per request, real for a product, and it scales with bigger tables, RAG results and long API responses.
I built a free token counter for this. It runs the same o200k tokenizer in your browser, so nothing you paste is uploaded. Paste your data in two formats and compare tokens and cost per model.
Full write-up with more detail: JSON vs YAML vs CSV: which format uses the fewest tokens?