{"slug": "i-counted-tokens-for-the-same-data-in-7-formats-pretty-json-costs-3-csv", "title": "I counted tokens for the same data in 7 formats. Pretty JSON costs 3 CSV", "summary": "A developer measured token counts for the same 20-row, 5-field product table serialized in seven formats using the o200k_base tokenizer behind GPT-4o, finding pretty-printed JSON costs roughly 3× the tokens of CSV (884 vs 300) and XML 3.63× (1,088 tokens). The same data minified as JSON came to 525 tokens and YAML to 649, while TSV (296) and CSV (300) were cheapest; the older cl100k_base tokenizer produced nearly identical results within 2%. At 30,000 monthly requests with a 20-row table attached, the developer estimated CSV costs about $18/month versus $53 for pretty JSON at $2 per million input tokens, and released a browser-based token counter that runs the same tokenizer locally.", "body_md": "Most of us paste data into LLM prompts as JSON, often straight from `JSON.stringify(data, null, 2)`. I had never checked what that costs in tokens, so I took one table, wrote it in seven formats and counted.\n\nShort version: **pretty-printed JSON uses about 3× the tokens of CSV.**\n\nA product table: 20 rows, 5 fields (id, name, price, in_stock, category). One row in CSV:\n\n```\n1001,Wireless Mouse,9.99,false,electronics\n```\n\nSame data, seven formats, counted with `o200k_base`, the tokenizer behind GPT-4o and later OpenAI models:\n\n``` js\nimport { getEncoding } from \"js-tiktoken\";\nconst enc = getEncoding(\"o200k_base\");\nconst count = (s) => enc.encode(s).length;\n\ncount(csv);                            // 300\ncount(JSON.stringify(rows));           // 525\ncount(JSON.stringify(rows, null, 2));  // 884\n```\n\n| Format | Tokens | vs CSV | \n|---|---|---|\n| TSV | 296 | 0.99× | \n| CSV | 300 | 1.00× | \n| Markdown table | 373 | 1.24× | \n| JSON, minified | 525 | 1.75× | \n| YAML | 649 | 2.16× | \n| JSON, pretty (2-space) | 884 | **2.95×** | \n| XML | 1,088 | **3.63×** | \n\nThe older `cl100k_base` tokenizer gave nearly identical numbers, within 2% for every format.\n\n`\"name\":`, `\"price\":` twenty times. CSV writes them once, in the header.\nYAML is the odd one: fewer characters than minified JSON, more tokens, because every field gets its own line and its own key.\n\nCoding agents send a lot of code, so I tried a few things:\n\n| Test | Result | \n|---|---|\n| 16-line Python file: 4 spaces vs 2 spaces vs tabs | 135 / 135 / 133 tokens, **basically no difference** | \n| Same file without its one-line docstring | 135 → 123 (−9%) | \n| 10-line JS function, minified | 92 → 47 (−49%) | \n| One UUID | **18 tokens** | \n\n`subtotal` and `taxRate`, which is exactly what helps the model understand the code.\nDropping unused fields, `null` fields and extra decimal places helps too.\n\nSay you attach a 20-row table to every request, 1,000 requests a day, 30,000 a month, at $2 per million input tokens:\n\n| Format | Tokens / month | Cost / month | \n|---|---|---|\n| CSV | 9.0M | $18 | \n| JSON, minified | 15.8M | $32 | \n| JSON, pretty | 26.5M | $53 | \n| XML | 32.6M | $65 | \n\nInvisible per request, real for a product, and it scales with bigger tables, RAG results and long API responses.\n\nI built a free [token counter](https://tokensave.app/) for this. It runs the same `o200k` tokenizer in your browser, so nothing you paste is uploaded. Paste your data in two formats and compare tokens and cost per model.\n\nFull write-up with more detail: [JSON vs YAML vs CSV: which format uses the fewest tokens?](https://tokensave.app/blog/json-vs-yaml-vs-csv-tokens)", "url": "https://wpnews.pro/news/i-counted-tokens-for-the-same-data-in-7-formats-pretty-json-costs-3-csv", "canonical_source": "https://dev.to/jaehyun_cho_0dff271e0d2e5/i-counted-tokens-for-the-same-data-in-7-formats-pretty-json-costs-3x-csv-8a2", "published_at": "2026-10-02 17:03:19+00:00", "updated_at": "2026-10-02 17:07:31.857434+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "developer-tools", "mlops"], "entities": ["OpenAI", "GPT-4o", "tokensave.app", "js-tiktoken", "o200k_base", "cl100k_base"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-counted-tokens-for-the-same-data-in-7-formats-pretty-json-costs-3-csv", "markdown": "https://wpnews.pro/news/i-counted-tokens-for-the-same-data-in-7-formats-pretty-json-costs-3-csv.md", "text": "https://wpnews.pro/news/i-counted-tokens-for-the-same-data-in-7-formats-pretty-json-costs-3-csv.txt", "jsonld": "https://wpnews.pro/news/i-counted-tokens-for-the-same-data-in-7-formats-pretty-json-costs-3-csv.jsonld"}}