{"slug": "claude-haiku-5-5-is-cheap-until-100k-tokens-build-a-tiny-price-gate-in", "title": "Claude Haiku 5.5 Is Cheap Until 100k Tokens. Build a Tiny Price Gate in TypeScript.", "summary": "A developer published a TypeScript price gate for Anthropic's Claude Haiku 5.5, released October 7, 2026, that returns ALLOW, REVIEW, or REFUSE before a request is sent based on prompt size, task type, and effort level. The gate applies Haiku 5.5's short-tier rates ($0.10/$0.50 per million input/output tokens) only up to 100,000 prompt tokens, switching to long-tier rates ($0.50/$2.50) and a REVIEW decision above that line, and flags agentic-coding and computer-use tasks for review even on short prompts, citing Anthropic's Terminal-Bench 4.0 scores of 39.2% for Haiku 5.5 versus 70.6% for Sonnet 5.5. It also accounts for the new tokenizer's roughly 30% higher token count and the model's low/medium/high/xhigh/max effort dial, defaulting to medium when effort is omitted.", "body_md": "On October 7, 2026, Anthropic released Claude Haiku 5.5. The model id is `claude-haiku-5-5`. For prompts up to 100,000 tokens, input is $0.10 per million tokens and output is $0.50. That is 90% below Haiku 4.5's $1.00 and $5.00. Cross 100,000 tokens and the rates become $0.50 and $2.50. That is a 50% cut, not a 90% cut.\n\nAnthropic also says the new tokenizer uses about 30% more tokens than Haiku 4.5 for the same text. A cheaper token is not automatically a cheaper task. And Haiku 5.5 is the first Haiku with an effort dial. The API default is `medium`. `low`, `high`, `xhigh`, and `max` are real settings. They change how much the model thinks, which changes the bill. The list price does not tell you that.\n\nThe useful question is what your code does before the request leaves.\n\nThis toy does not call Claude. You hand it token counts you already have, a task name, and an effort. It answers ALLOW, REVIEW, or REFUSE, and it prints the list-price bill in USD.\n\nPrompt size is input tokens plus cache reads plus cache writes. At or under 100,000, the short rates apply. Over that line, the long rates apply and the decision is REVIEW.\n\nALLOW is only for five narrow jobs: `classify`, `extract`, `route`, `summarize`, and `compact`. Effort must be `low` or `medium`. If you omit effort, the gate stamps `medium`, because that is the Claude API default for this model.\n\n`agentic-coding` and `computer-use` are REVIEW even on a short prompt. Anthropic's own Terminal-Bench 4.0 numbers are 39.2% for Haiku 5.5 and 70.6% for Sonnet 5.5. A lower price is not the same job. `high`, `xhigh`, and `max` are also REVIEW. This program has no eval that says those levels pay for themselves.\n\nBad counts and unknown labels are REFUSE. A negative token count should not become a discount.\n\nThe full source is below. Save it as `src/gate.ts`. Node 22 or newer. No packages. No key.\n\n```\nnode --experimental-strip-types src/gate.ts\n/**\n * Price gate for Claude Haiku 5.5.\n * No network. No API key. Rates are list prices in microdollars per million tokens.\n *\n * Sources, read 2026-10-08:\n * - https://www.anthropic.com/claude-haiku-5-5 (October 7, 2026)\n * - https://platform.claude.com/docs/en/models/haiku-5-5/overview\n * - https://platform.claude.com/docs/en/build-with-claude/effort\n */\n\nconst MILLION = 1_000_000n;\nconst PROMPT_LIMIT = 100_000n;\nconst RECOUNT_NUM = 130n;\nconst RECOUNT_DEN = 100n;\n\nconst ALLOW_TASKS = new Set([\n  \"classify\",\n  \"extract\",\n  \"route\",\n  \"summarize\",\n  \"compact\",\n]);\n\nconst REVIEW_TASKS = new Set([\"agentic-coding\", \"computer-use\"]);\n\nconst ALLOW_EFFORT = new Set([\"low\", \"medium\"]);\nconst ALL_EFFORT = new Set([\"low\", \"medium\", \"high\", \"xhigh\", \"max\"]);\n\ntype Effort = \"low\" | \"medium\" | \"high\" | \"xhigh\" | \"max\";\ntype Ttl = \"5m\" | \"1h\";\ntype Tier = \"short\" | \"long\";\n\ntype Rates = {\n  input: bigint;\n  output: bigint;\n  cacheRead: bigint;\n  cacheWrite5m: bigint;\n  cacheWrite1h: bigint;\n};\n\nconst HAIKU_55: Record<Tier, Rates> = {\n  short: {\n    input: 100_000n,\n    output: 500_000n,\n    cacheRead: 10_000n,\n    cacheWrite5m: 125_000n,\n    cacheWrite1h: 200_000n,\n  },\n  long: {\n    input: 500_000n,\n    output: 2_500_000n,\n    cacheRead: 50_000n,\n    cacheWrite5m: 625_000n,\n    cacheWrite1h: 1_000_000n,\n  },\n};\n\nconst HAIKU_45_5M: Rates = {\n  input: 1_000_000n,\n  output: 5_000_000n,\n  cacheRead: 100_000n,\n  cacheWrite5m: 1_250_000n,\n  cacheWrite1h: 0n,\n};\n\nconst SONNET_55_CACHE_READ_NEW = 100_000n;\nconst SONNET_55_CACHE_READ_OLD = 200_000n;\n\nexport type Job = {\n  name: string;\n  task: string;\n  effort?: string;\n  inputTokens: number;\n  outputTokens: number;\n  cacheReadTokens: number;\n  cacheWriteTokens: number;\n  cacheWriteTtl?: Ttl;\n  recountFromHaiku45?: boolean;\n};\n\nexport type Decision = \"ALLOW\" | \"REVIEW\" | \"REFUSE\";\n\nexport type Quote = {\n  name: string;\n  decision: Decision;\n  reason: string;\n  model: \"claude-haiku-5-5\";\n  effort: Effort | null;\n  assumedDefaultEffort: boolean;\n  tier: Tier | null;\n  promptTokens: bigint | null;\n  haiku55Usd: string | null;\n  haiku45Usd: string | null;\n  savingsBps: number | null;\n  sonnetCacheReadNewUsd: string | null;\n  sonnetCacheReadOldUsd: string | null;\n};\n\nfunction isWhole(value: number): boolean {\n  return Number.isFinite(value) && Number.isInteger(value) && value >= 0;\n}\n\nfunction mulDiv(tokens: bigint, perMillion: bigint): bigint {\n  return (tokens * perMillion) / MILLION;\n}\n\nfunction usd(micro: bigint): string {\n  const whole = micro / MILLION;\n  const frac = (micro % MILLION).toString().padStart(6, \"0\");\n  return `${whole}.${frac}`;\n}\n\nfunction bill(tokens: {\n  input: bigint;\n  output: bigint;\n  cacheRead: bigint;\n  cacheWrite: bigint;\n}, rates: Rates, ttl: Ttl): bigint {\n  const writeRate = ttl === \"1h\" ? rates.cacheWrite1h : rates.cacheWrite5m;\n  return (\n    mulDiv(tokens.input, rates.input) +\n    mulDiv(tokens.output, rates.output) +\n    mulDiv(tokens.cacheRead, rates.cacheRead) +\n    mulDiv(tokens.cacheWrite, writeRate)\n  );\n}\n\nfunction savingsBps(before: bigint, after: bigint): number | null {\n  if (before <= 0n) return null;\n  return Number(((before - after) * 10_000n) / before);\n}\n\nexport function quoteJob(job: Job): Quote {\n  const base = {\n    name: job.name,\n    model: \"claude-haiku-5-5\" as const,\n    effort: null,\n    assumedDefaultEffort: false,\n    tier: null,\n    promptTokens: null,\n    haiku55Usd: null,\n    haiku45Usd: null,\n    savingsBps: null,\n    sonnetCacheReadNewUsd: null,\n    sonnetCacheReadOldUsd: null,\n  };\n\n  const counts = [\n    job.inputTokens,\n    job.outputTokens,\n    job.cacheReadTokens,\n    job.cacheWriteTokens,\n  ];\n  if (!counts.every(isWhole)) {\n    return {\n      ...base,\n      decision: \"REFUSE\",\n      reason: \"token counts must be non-negative integers\",\n    };\n  }\n\n  const ttl = job.cacheWriteTtl ?? \"5m\";\n  if (ttl !== \"5m\" && ttl !== \"1h\") {\n    return { ...base, decision: \"REFUSE\", reason: \"cache write ttl must be 5m or 1h\" };\n  }\n\n  let effort = job.effort;\n  let assumedDefaultEffort = false;\n  if (effort === undefined) {\n    effort = \"medium\";\n    assumedDefaultEffort = true;\n  }\n  if (!ALL_EFFORT.has(effort)) {\n    return { ...base, decision: \"REFUSE\", reason: `unknown effort ${job.effort}` };\n  }\n\n  const knownTask = ALLOW_TASKS.has(job.task) || REVIEW_TASKS.has(job.task);\n  if (!knownTask) {\n    return { ...base, decision: \"REFUSE\", reason: `unknown task ${job.task}` };\n  }\n\n  let input = BigInt(job.inputTokens);\n  let cacheRead = BigInt(job.cacheReadTokens);\n  let cacheWrite = BigInt(job.cacheWriteTokens);\n  const output = BigInt(job.outputTokens);\n\n  if (job.recountFromHaiku45) {\n    input = (input * RECOUNT_NUM) / RECOUNT_DEN;\n    cacheRead = (cacheRead * RECOUNT_NUM) / RECOUNT_DEN;\n    cacheWrite = (cacheWrite * RECOUNT_NUM) / RECOUNT_DEN;\n  }\n\n  const promptTokens = input + cacheRead + cacheWrite;\n  const tier: Tier = promptTokens <= PROMPT_LIMIT ? \"short\" : \"long\";\n  const haiku55 = bill(\n    { input, output, cacheRead, cacheWrite },\n    HAIKU_55[tier],\n    ttl,\n  );\n\n  let haiku45: bigint | null = null;\n  if (ttl === \"5m\") {\n    const oldInput = BigInt(job.inputTokens);\n    const oldRead = BigInt(job.cacheReadTokens);\n    const oldWrite = BigInt(job.cacheWriteTokens);\n    haiku45 = bill(\n      { input: oldInput, output, cacheRead: oldRead, cacheWrite: oldWrite },\n      HAIKU_45_5M,\n      \"5m\",\n    );\n  }\n\n  const reasons: string[] = [];\n  if (!ALLOW_TASKS.has(job.task)) {\n    reasons.push(`${job.task} stays off the Haiku allow-list`);\n  }\n  if (!ALLOW_EFFORT.has(effort)) {\n    reasons.push(`effort ${effort} needs an eval this toy does not have`);\n  }\n  if (tier === \"long\") {\n    reasons.push(\"prompt is over 100000 tokens, so the long-tier rates apply\");\n  }\n  if (job.recountFromHaiku45) {\n    reasons.push(\"prompt tokens were recounted at 130/100 before the Haiku 5.5 bill\");\n  }\n  if (assumedDefaultEffort) {\n    reasons.push(\"missing effort was stamped medium, the API default\");\n  }\n\n  const decision: Decision = reasons.some((reason) =>\n    reason.startsWith(\"prompt is over\") ||\n    reason.includes(\"allow-list\") ||\n    reason.includes(\"needs an eval\")\n  )\n    ? \"REVIEW\"\n    : \"ALLOW\";\n\n  if (decision === \"ALLOW\") {\n    reasons.unshift(\"short prompt, allow-listed task, effort is low or medium\");\n  }\n  if (ttl === \"1h\") {\n    reasons.push(\"one-hour cache writes are quoted for Haiku 5.5 only\");\n  }\n\n  const sonnetNew = mulDiv(cacheRead, SONNET_55_CACHE_READ_NEW);\n  const sonnetOld = mulDiv(BigInt(job.cacheReadTokens), SONNET_55_CACHE_READ_OLD);\n\n  return {\n    ...base,\n    decision,\n    reason: reasons.join(\"; \"),\n    effort: effort as Effort,\n    assumedDefaultEffort,\n    tier,\n    promptTokens,\n    haiku55Usd: usd(haiku55),\n    haiku45Usd: haiku45 === null ? null : usd(haiku45),\n    savingsBps: haiku45 === null ? null : savingsBps(haiku45, haiku55),\n    sonnetCacheReadNewUsd: usd(sonnetNew),\n    sonnetCacheReadOldUsd: usd(sonnetOld),\n  };\n}\n\nfunction line(quote: Quote): string {\n  const parts = [\n    quote.name,\n    quote.decision,\n    quote.model,\n    quote.effort === null ? \"effort=none\" : `effort=${quote.effort}`,\n    quote.tier === null ? \"tier=none\" : `tier=${quote.tier}`,\n    quote.promptTokens === null ? \"prompt=none\" : `prompt=${quote.promptTokens}`,\n    quote.haiku55Usd === null ? \"haiku55=none\" : `haiku55=${quote.haiku55Usd}`,\n    quote.haiku45Usd === null ? \"haiku45=none\" : `haiku45=${quote.haiku45Usd}`,\n    quote.savingsBps === null ? \"savings_bps=none\" : `savings_bps=${quote.savingsBps}`,\n    quote.reason,\n  ];\n  return parts.join(\" | \");\n}\n\nconst fixtures: Job[] = [\n  {\n    name: \"classify-short\",\n    task: \"classify\",\n    effort: \"medium\",\n    inputTokens: 2_000,\n    outputTokens: 200,\n    cacheReadTokens: 0,\n    cacheWriteTokens: 0,\n  },\n  {\n    name: \"compact-under-limit\",\n    task: \"compact\",\n    effort: \"low\",\n    inputTokens: 4_000,\n    outputTokens: 500,\n    cacheReadTokens: 80_000,\n    cacheWriteTokens: 0,\n  },\n  {\n    name: \"compact-over-limit\",\n    task: \"compact\",\n    effort: \"medium\",\n    inputTokens: 20_000,\n    outputTokens: 400,\n    cacheReadTokens: 90_000,\n    cacheWriteTokens: 0,\n  },\n  {\n    name: \"coding-short\",\n    task: \"agentic-coding\",\n    effort: \"medium\",\n    inputTokens: 8_000,\n    outputTokens: 1_200,\n    cacheReadTokens: 0,\n    cacheWriteTokens: 0,\n  },\n  {\n    name: \"classify-max-effort\",\n    task: \"classify\",\n    effort: \"max\",\n    inputTokens: 2_000,\n    outputTokens: 200,\n    cacheReadTokens: 0,\n    cacheWriteTokens: 0,\n  },\n  {\n    name: \"extract-default-effort\",\n    task: \"extract\",\n    inputTokens: 1_500,\n    outputTokens: 300,\n    cacheReadTokens: 0,\n    cacheWriteTokens: 0,\n  },\n  {\n    name: \"recount-30\",\n    task: \"summarize\",\n    effort: \"medium\",\n    inputTokens: 10_000,\n    outputTokens: 500,\n    cacheReadTokens: 0,\n    cacheWriteTokens: 0,\n    recountFromHaiku45: true,\n  },\n  {\n    name: \"one-hour-write\",\n    task: \"route\",\n    effort: \"low\",\n    inputTokens: 1_000,\n    outputTokens: 100,\n    cacheReadTokens: 0,\n    cacheWriteTokens: 4_000,\n    cacheWriteTtl: \"1h\",\n  },\n  {\n    name: \"bad-tokens\",\n    task: \"classify\",\n    effort: \"low\",\n    inputTokens: -1,\n    outputTokens: 10,\n    cacheReadTokens: 0,\n    cacheWriteTokens: 0,\n  },\n  {\n    name: \"unknown-effort\",\n    task: \"classify\",\n    effort: \"turbo\",\n    inputTokens: 100,\n    outputTokens: 10,\n    cacheReadTokens: 0,\n    cacheWriteTokens: 0,\n  },\n];\n\nconst expected: Record<string, { decision: Decision; haiku55Usd: string | null; savingsBps: number | null }> = {\n  \"classify-short\": { decision: \"ALLOW\", haiku55Usd: \"0.000300\", savingsBps: 9000 },\n  \"compact-under-limit\": { decision: \"ALLOW\", haiku55Usd: \"0.001450\", savingsBps: 9000 },\n  \"compact-over-limit\": { decision: \"REVIEW\", haiku55Usd: \"0.015500\", savingsBps: 5000 },\n  \"coding-short\": { decision: \"REVIEW\", haiku55Usd: \"0.001400\", savingsBps: 9000 },\n  \"classify-max-effort\": { decision: \"REVIEW\", haiku55Usd: \"0.000300\", savingsBps: 9000 },\n  \"extract-default-effort\": { decision: \"ALLOW\", haiku55Usd: \"0.000300\", savingsBps: 9000 },\n  \"recount-30\": { decision: \"ALLOW\", haiku55Usd: \"0.001550\", savingsBps: 8760 },\n  \"one-hour-write\": { decision: \"ALLOW\", haiku55Usd: \"0.000950\", savingsBps: null },\n  \"bad-tokens\": { decision: \"REFUSE\", haiku55Usd: null, savingsBps: null },\n  \"unknown-effort\": { decision: \"REFUSE\", haiku55Usd: null, savingsBps: null },\n};\n\nfunction main(): void {\n  let failed = 0;\n  for (const job of fixtures) {\n    const quote = quoteJob(job);\n    const want = expected[job.name];\n    const ok =\n      want.decision === quote.decision &&\n      want.haiku55Usd === quote.haiku55Usd &&\n      want.savingsBps === quote.savingsBps;\n    if (!ok) {\n      failed += 1;\n      console.error(`FAIL ${job.name}`);\n      console.error(quote);\n    }\n    console.log(line(quote));\n  }\n\n  const cacheRead = 100_000n;\n  const sonnetNew = usd(mulDiv(cacheRead, SONNET_55_CACHE_READ_NEW));\n  const sonnetOld = usd(mulDiv(cacheRead, SONNET_55_CACHE_READ_OLD));\n  console.log(\n    `sonnet-cache-read | QUOTE | claude-sonnet-5-5 | cache_read_tokens=100000 | new=${sonnetNew} | old=${sonnetOld} | October 7 cache-read cut, not a routing decision`,\n  );\n  if (sonnetNew !== \"0.010000\" || sonnetOld !== \"0.020000\") {\n    failed += 1;\n    console.error(\"FAIL sonnet-cache-read\");\n  }\n\n  if (failed > 0) {\n    console.error(`${failed} check(s) failed`);\n    process.exit(1);\n  }\n  console.log(`checks=11 passed`);\n}\n\nmain();\n```\n\nChecked on October 8, 2026, Node.js 25.6.0. Eleven checks passed:\n\n```\nclassify-short | ALLOW | claude-haiku-5-5 | effort=medium | tier=short | prompt=2000 | haiku55=0.000300 | haiku45=0.003000 | savings_bps=9000 | short prompt, allow-listed task, effort is low or medium\ncompact-under-limit | ALLOW | claude-haiku-5-5 | effort=low | tier=short | prompt=84000 | haiku55=0.001450 | haiku45=0.014500 | savings_bps=9000 | short prompt, allow-listed task, effort is low or medium\ncompact-over-limit | REVIEW | claude-haiku-5-5 | effort=medium | tier=long | prompt=110000 | haiku55=0.015500 | haiku45=0.031000 | savings_bps=5000 | prompt is over 100000 tokens, so the long-tier rates apply\ncoding-short | REVIEW | claude-haiku-5-5 | effort=medium | tier=short | prompt=8000 | haiku55=0.001400 | haiku45=0.014000 | savings_bps=9000 | agentic-coding stays off the Haiku allow-list\nclassify-max-effort | REVIEW | claude-haiku-5-5 | effort=max | tier=short | prompt=2000 | haiku55=0.000300 | haiku45=0.003000 | savings_bps=9000 | effort max needs an eval this toy does not have\nextract-default-effort | ALLOW | claude-haiku-5-5 | effort=medium | tier=short | prompt=1500 | haiku55=0.000300 | haiku45=0.003000 | savings_bps=9000 | short prompt, allow-listed task, effort is low or medium; missing effort was stamped medium, the API default\nrecount-30 | ALLOW | claude-haiku-5-5 | effort=medium | tier=short | prompt=13000 | haiku55=0.001550 | haiku45=0.012500 | savings_bps=8760 | short prompt, allow-listed task, effort is low or medium; prompt tokens were recounted at 130/100 before the Haiku 5.5 bill\none-hour-write | ALLOW | claude-haiku-5-5 | effort=low | tier=short | prompt=5000 | haiku55=0.000950 | haiku45=none | savings_bps=none | short prompt, allow-listed task, effort is low or medium; one-hour cache writes are quoted for Haiku 5.5 only\nbad-tokens | REFUSE | claude-haiku-5-5 | effort=none | tier=none | prompt=none | haiku55=none | haiku45=none | savings_bps=none | token counts must be non-negative integers\nunknown-effort | REFUSE | claude-haiku-5-5 | effort=none | tier=none | prompt=none | haiku55=none | haiku45=none | savings_bps=none | unknown effort turbo\nsonnet-cache-read | QUOTE | claude-sonnet-5-5 | cache_read_tokens=100000 | new=0.010000 | old=0.020000 | October 7 cache-read cut, not a routing decision\nchecks=11 passed\n```\n\n`savings_bps` is basis points against a Haiku 4.5 bill at the original token counts. 9000 is 90.00%. 5000 is 50.00%. 8760 is 87.60%.\n\nRead three of those lines closely.\n\n`classify-short` is the brochure case. 2,000 input tokens and 200 output tokens, effort `medium`. Haiku 5.5 is $0.000300. Haiku 4.5 is $0.003000. Same tokens, 90% less.\n\n`compact-over-limit` is the line you can miss. 20,000 fresh input tokens plus 90,000 cache-read tokens is a 110,000-token prompt. The job is still `compact`, which is on the allow-list, and effort is still `medium`. The decision is REVIEW. Haiku 5.5 costs $0.015500. Haiku 4.5 costs $0.031000. The saving is 50%, because the long-tier rates are five times the short-tier rates, and Haiku 4.5's flat rate is ten times the short-tier rate.\n\n`recount-30` starts from a Haiku 4.5 count of 10,000 input tokens and 500 output tokens. The gate multiplies the input by 130/100 before it prices Haiku 5.5, so the new bill uses 13,000 input tokens. Haiku 5.5 is $0.001550. Haiku 4.5 is $0.012500. That is 87.60% less, not 90%.\n\nThe last quote is not a route. On October 7, Anthropic cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens. 100,000 cache-read tokens are $0.010000 now and were $0.020000. The gate prints that. It does not send the job to Sonnet.\n\n`max` on a classify job still shows the short-tier list price, then REVIEW.`haiku45` empty.\nRates used, USD per million tokens:\n\n| Rate | Haiku 5.5, prompt ≤100k | Haiku 5.5, prompt >100k | Haiku 4.5 | \n|---|---|---|---|\n| Input | 0.10 | 0.50 | 1.00 | \n| Output | 0.50 | 2.50 | 5.00 | \n| Cache read | 0.01 | 0.05 | 0.10 | \n| Cache write, 5 min | 0.125 | 0.625 | 1.25 | \n| Cache write, 1 hour | 0.20 | 1.00 | not compared | \n\nI'm building Roster, AI employees that do real work. A smaller model still needs a gate that knows which price it is about to pay.", "url": "https://wpnews.pro/news/claude-haiku-5-5-is-cheap-until-100k-tokens-build-a-tiny-price-gate-in", "canonical_source": "https://dev.to/bobbyhalljr/claude-haiku-55-is-cheap-until-100k-tokens-build-a-tiny-price-gate-in-typescript-3k4", "published_at": "2026-10-08 21:40:03+00:00", "updated_at": "2026-10-08 21:48:37.749223+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-products", "developer-tools"], "entities": ["Anthropic", "Claude Haiku 5.5", "Claude Haiku 4.5", "Claude Sonnet 5.5", "Terminal-Bench 4.0", "TypeScript", "Node.js"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/claude-haiku-5-5-is-cheap-until-100k-tokens-build-a-tiny-price-gate-in", "markdown": "https://wpnews.pro/news/claude-haiku-5-5-is-cheap-until-100k-tokens-build-a-tiny-price-gate-in.md", "text": "https://wpnews.pro/news/claude-haiku-5-5-is-cheap-until-100k-tokens-build-a-tiny-price-gate-in.txt", "jsonld": "https://wpnews.pro/news/claude-haiku-5-5-is-cheap-until-100k-tokens-build-a-tiny-price-gate-in.jsonld"}}