cd /news/ai-tools/claude-haiku-5-5-is-cheap-until-100k… · home › topics › ai-tools › article
[ARTICLE · art-147873] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Claude Haiku 5.5 Is Cheap Until 100k Tokens. Build a Tiny Price Gate in TypeScript.

A developer published a TypeScript price gate for Anthropic's Claude Haiku 5.5, released October 7, 2026, that returns ALLOW, REVIEW, or REFUSE before a request is sent based on prompt size, task type, and effort level. The gate applies Haiku 5.5's short-tier rates ($0.10/$0.50 per million input/output tokens) only up to 100,000 prompt tokens, switching to long-tier rates ($0.50/$2.50) and a REVIEW decision above that line, and flags agentic-coding and computer-use tasks for review even on short prompts, citing Anthropic's Terminal-Bench 4.0 scores of 39.2% for Haiku 5.5 versus 70.6% for Sonnet 5.5. It also accounts for the new tokenizer's roughly 30% higher token count and the model's low/medium/high/xhigh/max effort dial, defaulting to medium when effort is omitted.

by read11 min views2 publishedOct 8, 2026

On October 7, 2026, Anthropic released Claude Haiku 5.5. The model id is claude-haiku-5-5. For prompts up to 100,000 tokens, input is $0.10 per million tokens and output is $0.50. That is 90% below Haiku 4.5's $1.00 and $5.00. Cross 100,000 tokens and the rates become $0.50 and $2.50. That is a 50% cut, not a 90% cut.

Anthropic also says the new tokenizer uses about 30% more tokens than Haiku 4.5 for the same text. A cheaper token is not automatically a cheaper task. And Haiku 5.5 is the first Haiku with an effort dial. The API default is medium. low, high, xhigh, and max are real settings. They change how much the model thinks, which changes the bill. The list price does not tell you that.

The useful question is what your code does before the request leaves.

This toy does not call Claude. You hand it token counts you already have, a task name, and an effort. It answers ALLOW, REVIEW, or REFUSE, and it prints the list-price bill in USD.

Prompt size is input tokens plus cache reads plus cache writes. At or under 100,000, the short rates apply. Over that line, the long rates apply and the decision is REVIEW.

ALLOW is only for five narrow jobs: classify, extract, route, summarize, and compact. Effort must be low or medium. If you omit effort, the gate stamps medium, because that is the Claude API default for this model.

agentic-coding and computer-use are REVIEW even on a short prompt. Anthropic's own Terminal-Bench 4.0 numbers are 39.2% for Haiku 5.5 and 70.6% for Sonnet 5.5. A lower price is not the same job. high, xhigh, and max are also REVIEW. This program has no eval that says those levels pay for themselves.

Bad counts and unknown labels are REFUSE. A negative token count should not become a discount.

The full source is below. Save it as src/gate.ts. Node 22 or newer. No packages. No key.

node --experimental-strip-types src/gate.ts
/**
 * Price gate for Claude Haiku 5.5.
 * No network. No API key. Rates are list prices in microdollars per million tokens.
 *
 * Sources, read 2026-10-08:
 * - https://www.anthropic.com/claude-haiku-5-5 (October 7, 2026)
 * - https://platform.claude.com/docs/en/models/haiku-5-5/overview
 * - https://platform.claude.com/docs/en/build-with-claude/effort
 */

const MILLION = 1_000_000n;
const PROMPT_LIMIT = 100_000n;
const RECOUNT_NUM = 130n;
const RECOUNT_DEN = 100n;

const ALLOW_TASKS = new Set([
  "classify",
  "extract",
  "route",
  "summarize",
  "compact",
]);

const REVIEW_TASKS = new Set(["agentic-coding", "computer-use"]);

const ALLOW_EFFORT = new Set(["low", "medium"]);
const ALL_EFFORT = new Set(["low", "medium", "high", "xhigh", "max"]);

type Effort = "low" | "medium" | "high" | "xhigh" | "max";
type Ttl = "5m" | "1h";
type Tier = "short" | "long";

type Rates = {
  input: bigint;
  output: bigint;
  cacheRead: bigint;
  cacheWrite5m: bigint;
  cacheWrite1h: bigint;
};

const HAIKU_55: Record<Tier, Rates> = {
  short: {
    input: 100_000n,
    output: 500_000n,
    cacheRead: 10_000n,
    cacheWrite5m: 125_000n,
    cacheWrite1h: 200_000n,
  },
  long: {
    input: 500_000n,
    output: 2_500_000n,
    cacheRead: 50_000n,
    cacheWrite5m: 625_000n,
    cacheWrite1h: 1_000_000n,
  },
};

const HAIKU_45_5M: Rates = {
  input: 1_000_000n,
  output: 5_000_000n,
  cacheRead: 100_000n,
  cacheWrite5m: 1_250_000n,
  cacheWrite1h: 0n,
};

const SONNET_55_CACHE_READ_NEW = 100_000n;
const SONNET_55_CACHE_READ_OLD = 200_000n;

export type Job = {
  name: string;
  task: string;
  effort?: string;
  inputTokens: number;
  outputTokens: number;
  cacheReadTokens: number;
  cacheWriteTokens: number;
  cacheWriteTtl?: Ttl;
  recountFromHaiku45?: boolean;
};

export type Decision = "ALLOW" | "REVIEW" | "REFUSE";

export type Quote = {
  name: string;
  decision: Decision;
  reason: string;
  model: "claude-haiku-5-5";
  effort: Effort | null;
  assumedDefaultEffort: boolean;
  tier: Tier | null;
  promptTokens: bigint | null;
  haiku55Usd: string | null;
  haiku45Usd: string | null;
  savingsBps: number | null;
  sonnetCacheReadNewUsd: string | null;
  sonnetCacheReadOldUsd: string | null;
};

function isWhole(value: number): boolean {
  return Number.isFinite(value) && Number.isInteger(value) && value >= 0;
}

function mulDiv(tokens: bigint, perMillion: bigint): bigint {
  return (tokens * perMillion) / MILLION;
}

function usd(micro: bigint): string {
  const whole = micro / MILLION;
  const frac = (micro % MILLION).toString().padStart(6, "0");
  return `${whole}.${frac}`;
}

function bill(tokens: {
  input: bigint;
  output: bigint;
  cacheRead: bigint;
  cacheWrite: bigint;
}, rates: Rates, ttl: Ttl): bigint {
  const writeRate = ttl === "1h" ? rates.cacheWrite1h : rates.cacheWrite5m;
  return (
    mulDiv(tokens.input, rates.input) +
    mulDiv(tokens.output, rates.output) +
    mulDiv(tokens.cacheRead, rates.cacheRead) +
    mulDiv(tokens.cacheWrite, writeRate)
  );
}

function savingsBps(before: bigint, after: bigint): number | null {
  if (before <= 0n) return null;
  return Number(((before - after) * 10_000n) / before);
}

export function quoteJob(job: Job): Quote {
  const base = {
    name: job.name,
    model: "claude-haiku-5-5" as const,
    effort: null,
    assumedDefaultEffort: false,
    tier: null,
    promptTokens: null,
    haiku55Usd: null,
    haiku45Usd: null,
    savingsBps: null,
    sonnetCacheReadNewUsd: null,
    sonnetCacheReadOldUsd: null,
  };

  const counts = [
    job.inputTokens,
    job.outputTokens,
    job.cacheReadTokens,
    job.cacheWriteTokens,
  ];
  if (!counts.every(isWhole)) {
    return {
      ...base,
      decision: "REFUSE",
      reason: "token counts must be non-negative integers",
    };
  }

  const ttl = job.cacheWriteTtl ?? "5m";
  if (ttl !== "5m" && ttl !== "1h") {
    return { ...base, decision: "REFUSE", reason: "cache write ttl must be 5m or 1h" };
  }

  let effort = job.effort;
  let assumedDefaultEffort = false;
  if (effort === undefined) {
    effort = "medium";
    assumedDefaultEffort = true;
  }
  if (!ALL_EFFORT.has(effort)) {
    return { ...base, decision: "REFUSE", reason: `unknown effort ${job.effort}` };
  }

  const knownTask = ALLOW_TASKS.has(job.task) || REVIEW_TASKS.has(job.task);
  if (!knownTask) {
    return { ...base, decision: "REFUSE", reason: `unknown task ${job.task}` };
  }

  let input = BigInt(job.inputTokens);
  let cacheRead = BigInt(job.cacheReadTokens);
  let cacheWrite = BigInt(job.cacheWriteTokens);
  const output = BigInt(job.outputTokens);

  if (job.recountFromHaiku45) {
    input = (input * RECOUNT_NUM) / RECOUNT_DEN;
    cacheRead = (cacheRead * RECOUNT_NUM) / RECOUNT_DEN;
    cacheWrite = (cacheWrite * RECOUNT_NUM) / RECOUNT_DEN;
  }

  const promptTokens = input + cacheRead + cacheWrite;
  const tier: Tier = promptTokens <= PROMPT_LIMIT ? "short" : "long";
  const haiku55 = bill(
    { input, output, cacheRead, cacheWrite },
    HAIKU_55[tier],
    ttl,
  );

  let haiku45: bigint | null = null;
  if (ttl === "5m") {
    const oldInput = BigInt(job.inputTokens);
    const oldRead = BigInt(job.cacheReadTokens);
    const oldWrite = BigInt(job.cacheWriteTokens);
    haiku45 = bill(
      { input: oldInput, output, cacheRead: oldRead, cacheWrite: oldWrite },
      HAIKU_45_5M,
      "5m",
    );
  }

  const reasons: string[] = [];
  if (!ALLOW_TASKS.has(job.task)) {
    reasons.push(`${job.task} stays off the Haiku allow-list`);
  }
  if (!ALLOW_EFFORT.has(effort)) {
    reasons.push(`effort ${effort} needs an eval this toy does not have`);
  }
  if (tier === "long") {
    reasons.push("prompt is over 100000 tokens, so the long-tier rates apply");
  }
  if (job.recountFromHaiku45) {
    reasons.push("prompt tokens were recounted at 130/100 before the Haiku 5.5 bill");
  }
  if (assumedDefaultEffort) {
    reasons.push("missing effort was stamped medium, the API default");
  }

  const decision: Decision = reasons.some((reason) =>
    reason.startsWith("prompt is over") ||
    reason.includes("allow-list") ||
    reason.includes("needs an eval")
  )
    ? "REVIEW"
    : "ALLOW";

  if (decision === "ALLOW") {
    reasons.unshift("short prompt, allow-listed task, effort is low or medium");
  }
  if (ttl === "1h") {
    reasons.push("one-hour cache writes are quoted for Haiku 5.5 only");
  }

  const sonnetNew = mulDiv(cacheRead, SONNET_55_CACHE_READ_NEW);
  const sonnetOld = mulDiv(BigInt(job.cacheReadTokens), SONNET_55_CACHE_READ_OLD);

  return {
    ...base,
    decision,
    reason: reasons.join("; "),
    effort: effort as Effort,
    assumedDefaultEffort,
    tier,
    promptTokens,
    haiku55Usd: usd(haiku55),
    haiku45Usd: haiku45 === null ? null : usd(haiku45),
    savingsBps: haiku45 === null ? null : savingsBps(haiku45, haiku55),
    sonnetCacheReadNewUsd: usd(sonnetNew),
    sonnetCacheReadOldUsd: usd(sonnetOld),
  };
}

function line(quote: Quote): string {
  const parts = [
    quote.name,
    quote.decision,
    quote.model,
    quote.effort === null ? "effort=none" : `effort=${quote.effort}`,
    quote.tier === null ? "tier=none" : `tier=${quote.tier}`,
    quote.promptTokens === null ? "prompt=none" : `prompt=${quote.promptTokens}`,
    quote.haiku55Usd === null ? "haiku55=none" : `haiku55=${quote.haiku55Usd}`,
    quote.haiku45Usd === null ? "haiku45=none" : `haiku45=${quote.haiku45Usd}`,
    quote.savingsBps === null ? "savings_bps=none" : `savings_bps=${quote.savingsBps}`,
    quote.reason,
  ];
  return parts.join(" | ");
}

const fixtures: Job[] = [
  {
    name: "classify-short",
    task: "classify",
    effort: "medium",
    inputTokens: 2_000,
    outputTokens: 200,
    cacheReadTokens: 0,
    cacheWriteTokens: 0,
  },
  {
    name: "compact-under-limit",
    task: "compact",
    effort: "low",
    inputTokens: 4_000,
    outputTokens: 500,
    cacheReadTokens: 80_000,
    cacheWriteTokens: 0,
  },
  {
    name: "compact-over-limit",
    task: "compact",
    effort: "medium",
    inputTokens: 20_000,
    outputTokens: 400,
    cacheReadTokens: 90_000,
    cacheWriteTokens: 0,
  },
  {
    name: "coding-short",
    task: "agentic-coding",
    effort: "medium",
    inputTokens: 8_000,
    outputTokens: 1_200,
    cacheReadTokens: 0,
    cacheWriteTokens: 0,
  },
  {
    name: "classify-max-effort",
    task: "classify",
    effort: "max",
    inputTokens: 2_000,
    outputTokens: 200,
    cacheReadTokens: 0,
    cacheWriteTokens: 0,
  },
  {
    name: "extract-default-effort",
    task: "extract",
    inputTokens: 1_500,
    outputTokens: 300,
    cacheReadTokens: 0,
    cacheWriteTokens: 0,
  },
  {
    name: "recount-30",
    task: "summarize",
    effort: "medium",
    inputTokens: 10_000,
    outputTokens: 500,
    cacheReadTokens: 0,
    cacheWriteTokens: 0,
    recountFromHaiku45: true,
  },
  {
    name: "one-hour-write",
    task: "route",
    effort: "low",
    inputTokens: 1_000,
    outputTokens: 100,
    cacheReadTokens: 0,
    cacheWriteTokens: 4_000,
    cacheWriteTtl: "1h",
  },
  {
    name: "bad-tokens",
    task: "classify",
    effort: "low",
    inputTokens: -1,
    outputTokens: 10,
    cacheReadTokens: 0,
    cacheWriteTokens: 0,
  },
  {
    name: "unknown-effort",
    task: "classify",
    effort: "turbo",
    inputTokens: 100,
    outputTokens: 10,
    cacheReadTokens: 0,
    cacheWriteTokens: 0,
  },
];

const expected: Record<string, { decision: Decision; haiku55Usd: string | null; savingsBps: number | null }> = {
  "classify-short": { decision: "ALLOW", haiku55Usd: "0.000300", savingsBps: 9000 },
  "compact-under-limit": { decision: "ALLOW", haiku55Usd: "0.001450", savingsBps: 9000 },
  "compact-over-limit": { decision: "REVIEW", haiku55Usd: "0.015500", savingsBps: 5000 },
  "coding-short": { decision: "REVIEW", haiku55Usd: "0.001400", savingsBps: 9000 },
  "classify-max-effort": { decision: "REVIEW", haiku55Usd: "0.000300", savingsBps: 9000 },
  "extract-default-effort": { decision: "ALLOW", haiku55Usd: "0.000300", savingsBps: 9000 },
  "recount-30": { decision: "ALLOW", haiku55Usd: "0.001550", savingsBps: 8760 },
  "one-hour-write": { decision: "ALLOW", haiku55Usd: "0.000950", savingsBps: null },
  "bad-tokens": { decision: "REFUSE", haiku55Usd: null, savingsBps: null },
  "unknown-effort": { decision: "REFUSE", haiku55Usd: null, savingsBps: null },
};

function main(): void {
  let failed = 0;
  for (const job of fixtures) {
    const quote = quoteJob(job);
    const want = expected[job.name];
    const ok =
      want.decision === quote.decision &&
      want.haiku55Usd === quote.haiku55Usd &&
      want.savingsBps === quote.savingsBps;
    if (!ok) {
      failed += 1;
      console.error(`FAIL ${job.name}`);
      console.error(quote);
    }
    console.log(line(quote));
  }

  const cacheRead = 100_000n;
  const sonnetNew = usd(mulDiv(cacheRead, SONNET_55_CACHE_READ_NEW));
  const sonnetOld = usd(mulDiv(cacheRead, SONNET_55_CACHE_READ_OLD));
  console.log(
    `sonnet-cache-read | QUOTE | claude-sonnet-5-5 | cache_read_tokens=100000 | new=${sonnetNew} | old=${sonnetOld} | October 7 cache-read cut, not a routing decision`,
  );
  if (sonnetNew !== "0.010000" || sonnetOld !== "0.020000") {
    failed += 1;
    console.error("FAIL sonnet-cache-read");
  }

  if (failed > 0) {
    console.error(`${failed} check(s) failed`);
    process.exit(1);
  }
  console.log(`checks=11 passed`);
}

main();

Checked on October 8, 2026, Node.js 25.6.0. Eleven checks passed:

classify-short | ALLOW | claude-haiku-5-5 | effort=medium | tier=short | prompt=2000 | haiku55=0.000300 | haiku45=0.003000 | savings_bps=9000 | short prompt, allow-listed task, effort is low or medium
compact-under-limit | ALLOW | claude-haiku-5-5 | effort=low | tier=short | prompt=84000 | haiku55=0.001450 | haiku45=0.014500 | savings_bps=9000 | short prompt, allow-listed task, effort is low or medium
compact-over-limit | REVIEW | claude-haiku-5-5 | effort=medium | tier=long | prompt=110000 | haiku55=0.015500 | haiku45=0.031000 | savings_bps=5000 | prompt is over 100000 tokens, so the long-tier rates apply
coding-short | REVIEW | claude-haiku-5-5 | effort=medium | tier=short | prompt=8000 | haiku55=0.001400 | haiku45=0.014000 | savings_bps=9000 | agentic-coding stays off the Haiku allow-list
classify-max-effort | REVIEW | claude-haiku-5-5 | effort=max | tier=short | prompt=2000 | haiku55=0.000300 | haiku45=0.003000 | savings_bps=9000 | effort max needs an eval this toy does not have
extract-default-effort | ALLOW | claude-haiku-5-5 | effort=medium | tier=short | prompt=1500 | haiku55=0.000300 | haiku45=0.003000 | savings_bps=9000 | short prompt, allow-listed task, effort is low or medium; missing effort was stamped medium, the API default
recount-30 | ALLOW | claude-haiku-5-5 | effort=medium | tier=short | prompt=13000 | haiku55=0.001550 | haiku45=0.012500 | savings_bps=8760 | short prompt, allow-listed task, effort is low or medium; prompt tokens were recounted at 130/100 before the Haiku 5.5 bill
one-hour-write | ALLOW | claude-haiku-5-5 | effort=low | tier=short | prompt=5000 | haiku55=0.000950 | haiku45=none | savings_bps=none | short prompt, allow-listed task, effort is low or medium; one-hour cache writes are quoted for Haiku 5.5 only
bad-tokens | REFUSE | claude-haiku-5-5 | effort=none | tier=none | prompt=none | haiku55=none | haiku45=none | savings_bps=none | token counts must be non-negative integers
unknown-effort | REFUSE | claude-haiku-5-5 | effort=none | tier=none | prompt=none | haiku55=none | haiku45=none | savings_bps=none | unknown effort turbo
sonnet-cache-read | QUOTE | claude-sonnet-5-5 | cache_read_tokens=100000 | new=0.010000 | old=0.020000 | October 7 cache-read cut, not a routing decision
checks=11 passed

savings_bps is basis points against a Haiku 4.5 bill at the original token counts. 9000 is 90.00%. 5000 is 50.00%. 8760 is 87.60%.

Read three of those lines closely.

classify-short is the brochure case. 2,000 input tokens and 200 output tokens, effort medium. Haiku 5.5 is $0.000300. Haiku 4.5 is $0.003000. Same tokens, 90% less.

compact-over-limit is the line you can miss. 20,000 fresh input tokens plus 90,000 cache-read tokens is a 110,000-token prompt. The job is still compact, which is on the allow-list, and effort is still medium. The decision is REVIEW. Haiku 5.5 costs $0.015500. Haiku 4.5 costs $0.031000. The saving is 50%, because the long-tier rates are five times the short-tier rates, and Haiku 4.5's flat rate is ten times the short-tier rate.

recount-30 starts from a Haiku 4.5 count of 10,000 input tokens and 500 output tokens. The gate multiplies the input by 130/100 before it prices Haiku 5.5, so the new bill uses 13,000 input tokens. Haiku 5.5 is $0.001550. Haiku 4.5 is $0.012500. That is 87.60% less, not 90%.

The last quote is not a route. On October 7, Anthropic cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens. 100,000 cache-read tokens are $0.010000 now and were $0.020000. The gate prints that. It does not send the job to Sonnet.

max on a classify job still shows the short-tier list price, then REVIEW.haiku45 empty. Rates used, USD per million tokens:

Rate Haiku 5.5, prompt ≤100k Haiku 5.5, prompt >100k Haiku 4.5
Input 0.10 0.50 1.00
Output 0.50 2.50 5.00
Cache read 0.01 0.05 0.10
Cache write, 5 min 0.125 0.625 1.25
Cache write, 1 hour 0.20 1.00 not compared

I'm building Roster, AI employees that do real work. A smaller model still needs a gate that knows which price it is about to pay.

── more in #ai-tools 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-haiku-5-5-is-…] indexed:0 read:11min 2026-10-08 · —