# I counted tokens for the same data in 7 formats. Pretty JSON costs 3 CSV

> Source: <https://dev.to/jaehyun_cho_0dff271e0d2e5/i-counted-tokens-for-the-same-data-in-7-formats-pretty-json-costs-3x-csv-8a2>
> Published: 2026-10-02 17:03:19+00:00

Most of us paste data into LLM prompts as JSON, often straight from `JSON.stringify(data, null, 2)`. I had never checked what that costs in tokens, so I took one table, wrote it in seven formats and counted.

Short version: **pretty-printed JSON uses about 3× the tokens of CSV.**

A product table: 20 rows, 5 fields (id, name, price, in_stock, category). One row in CSV:

```
1001,Wireless Mouse,9.99,false,electronics
```

Same data, seven formats, counted with `o200k_base`, the tokenizer behind GPT-4o and later OpenAI models:

``` js
import { getEncoding } from "js-tiktoken";
const enc = getEncoding("o200k_base");
const count = (s) => enc.encode(s).length;

count(csv);                            // 300
count(JSON.stringify(rows));           // 525
count(JSON.stringify(rows, null, 2));  // 884
```

| Format | Tokens | vs CSV | 
|---|---|---|
| TSV | 296 | 0.99× | 
| CSV | 300 | 1.00× | 
| Markdown table | 373 | 1.24× | 
| JSON, minified | 525 | 1.75× | 
| YAML | 649 | 2.16× | 
| JSON, pretty (2-space) | 884 | **2.95×** | 
| XML | 1,088 | **3.63×** | 

The older `cl100k_base` tokenizer gave nearly identical numbers, within 2% for every format.

`"name":`, `"price":` twenty times. CSV writes them once, in the header.
YAML is the odd one: fewer characters than minified JSON, more tokens, because every field gets its own line and its own key.

Coding agents send a lot of code, so I tried a few things:

| Test | Result | 
|---|---|
| 16-line Python file: 4 spaces vs 2 spaces vs tabs | 135 / 135 / 133 tokens, **basically no difference** | 
| Same file without its one-line docstring | 135 → 123 (−9%) | 
| 10-line JS function, minified | 92 → 47 (−49%) | 
| One UUID | **18 tokens** | 

`subtotal` and `taxRate`, which is exactly what helps the model understand the code.
Dropping unused fields, `null` fields and extra decimal places helps too.

Say you attach a 20-row table to every request, 1,000 requests a day, 30,000 a month, at $2 per million input tokens:

| Format | Tokens / month | Cost / month | 
|---|---|---|
| CSV | 9.0M | $18 | 
| JSON, minified | 15.8M | $32 | 
| JSON, pretty | 26.5M | $53 | 
| XML | 32.6M | $65 | 

Invisible per request, real for a product, and it scales with bigger tables, RAG results and long API responses.

I built a free [token counter](https://tokensave.app/) for this. It runs the same `o200k` tokenizer in your browser, so nothing you paste is uploaded. Paste your data in two formats and compare tokens and cost per model.

Full write-up with more detail: [JSON vs YAML vs CSV: which format uses the fewest tokens?](https://tokensave.app/blog/json-vs-yaml-vs-csv-tokens)
