cd /news/ai-agents/claude-haiku-5-5-work-queues-design-… · home › topics › ai-agents › article
[ARTICLE · art-148888] src=pub.towardsai.net ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Claude Haiku 5.5 Work Queues: Design High-Volume AI Pipelines That Know When to Stop

A practical guide published alongside Anthropic's Claude Haiku 5.5 outlines a work-queue pattern that treats the fast, low-cost model as a bounded worker rather than the owner of an AI workflow, using immutable task packets with explicit stop rules, retry limits, and escalation paths. The guide recommends routing only narrow, verifiable jobs — such as ticket triage, invoice extraction, and document checks — to Claude Haiku 5.5, while keeping decisions on eligibility, policy exceptions, production data changes, regulated advice, and irreversible actions out of the fast lane. Anthropic positions larger models as the better option for complex agentic coding, and the pattern applies to any fast model placed in front of a more capable model or a human reviewer.

by read11 min views1 publishedOct 10, 2026

Fast AI workers are valuable only when their boundaries are clear. Here is a practical queue pattern for letting a small model handle the volume without letting it quietly own the hard calls.

A faster, cheaper model can create a strange new failure mode: your AI feature stops feeling expensive enough to supervise. A team that once ran a careful review loop for a few requests may suddenly send thousands of classifications, summaries, support drafts, or subagent jobs through the same system. The work moves. The queue clears. Then a bad answer is accepted because nobody decided what “good enough” meant.

That is why a high-volume AI pipeline needs a stop rule before it needs a clever router.

Claude Haiku 5.5 is designed for quick, repetitive work such as classification, compaction, extraction, routing, and subagent tasks. Anthropic also positions larger models as the better option for complex agentic coding. Those two ideas fit together neatly: use the small model as a bounded worker, not as the owner of the workflow.

This guide shows how to build that boundary. The pattern works for Claude Haiku 5.5, but it is useful for any fast model placed in front of a more capable model or a human reviewer.

A model router asks, “Which model should receive this prompt?” A work queue asks a more useful series of questions:

This changes the design from a vague intelligence contest into an operations system. Haiku can be excellent at pulling fields from an invoice, assigning a ticket type, finding a source passage, checking a narrow rule, or drafting a response from approved facts. It should not silently decide whether a disputed payment is valid, invent a policy exception, or approve an irreversible action.

The difference is not the number of tokens in the prompt. It is the consequence of being wrong and whether code can verify the output.

Do not push a free-form user message straight to a queue. First turn it into an immutable task packet. The packet gives the worker only the facts, tools, and authority needed for one small job. It also gives the next worker enough context to continue if the first one stops.

For example, a support system might queue “classify the issue and extract account identifiers,” not “solve this customer’s problem.” The second wording invites the model to improvise. The first creates a narrow, testable deliverable.

type TaskPacket = {  id: string;  type: "ticket_triage" | "invoice_extract" | "doc_check";  policyVersion: string;  input: {    text: string;    allowedDocumentIds: string[];    locale: string;  };  expected: {    schema: "triage.v1";    allowedLabels: string[];    requiresEvidence: boolean;  };  limits: {    maxAttempts: number;    deadlineMs: number;    consequence: "low" | "review_required";  };  lineage: {    parentJobId?: string;    createdAt: string;  };};

Notice what is missing: a broad instruction to “use your judgment,” access to unrelated customer records, and permission to mutate state. Those belong outside the worker. The queue owns retries and deadlines. A policy layer owns privileges. A deterministic service owns side effects.

A packet should also be idempotent. If a job is delivered twice, the system must not send two messages, create two refunds, or write two copies of the same record. Give the task an ID and make downstream writes conditional on it.

“The model sounded confident” is not an acceptance check. It is a reason to look for one.

Put work in the Haiku lane when you can state three things plainly. First, the result has a small output contract. Second, an external check can reject obvious bad results. Third, a failed task can be escalated without losing the original input or creating harm.

Good fast-lane examples include:

Keep a task out of the fast lane when it decides eligibility, interprets a policy exception, changes production data, gives regulated advice, chooses an irreversible action, or requires a deep synthesis that cannot be checked cheaply. A stronger model may help with those jobs, but it should still hand its recommendation to an application-controlled gate.

The worker may produce structured output. That is useful, but structured output alone does not prove the content is right. An acceptance gate should test the shape, the grounding, and the business rule separately.

For a ticket-triage task, the gate could require a label from the allowed set, evidence spans that occur in the source text, and a rule that says high-consequence labels always move to review. A model is allowed to propose; code decides whether the proposal can travel forward.

function decide(result: TriageResult, task: TaskPacket): Decision {  if (!matchesSchema(result, task.expected.schema)) {    return { kind: "escalate", reason: "invalid_schema" };  }
if (!task.expected.allowedLabels.includes(result.label)) {    return { kind: "escalate", reason: "unknown_label" };  }
if (!evidenceOccursInInput(result.evidence, task.input.text)) {    return { kind: "escalate", reason: "ungrounded_evidence" };  }
if (result.label === "account_access" || result.riskFlags.length > 0) {    return { kind: "review", reason: "sensitive_route" };  }
return { kind: "accept", reason: "checks_passed" };}

This code is intentionally ordinary. Your valuable logic should be ordinary enough to test. A prompt can tell Haiku to abstain when it lacks evidence, but the gate should enforce the same rule even if the prompt is ignored, shortened, or changed later.

Most queues do not need an elaborate model tournament. They need a short, explainable ladder.

Accept only when the output passes every deterministic check and the task’s consequence class is low. Store the worker’s model ID, prompt or policy version, task type, accepted output, and reason code. You will need these records when quality changes.

Escalate when the schema is valid but evidence is weak, a required field is absent, tool results conflict, the task exceeds its time or attempt budget, or the packet is outside the worker’s declared scope. Send a stronger model the original packet, the fast worker’s attempted result, and the specific failed checks. Do not ask the next model to start over with a vague “please solve this.”

Send a person the cases where no model should be final: an irreversible action, a policy exception, a high-risk customer issue, repeated disagreement, or a downstream system that cannot safely roll back. The reviewer should see the compact evidence packet, not a 200-message transcript.

Do not make a model’s self-reported confidence the sole trigger for any step. Use it as one signal, then compare it with missing fields, evidence checks, rule hits, disagreement, and task age. Recent routing research makes the same warning: a promising escalation signal can merely be tracking how difficult the question looks. Test it against simpler baselines before it earns a place in production.

Escalation is not an error path to erase. It is evidence. If Haiku classified an item as “billing” but failed to attach a source span, the stronger model or reviewer should know that. It may reveal an unclear taxonomy, poor source material, or a changed customer pattern.

Attach these fields to each handoff:

This helps you debug the system without asking an operator to reconstruct what happened from logs. It also lets you study escalation patterns later. If one label drives half of escalations, improve that label’s examples or make it a review-only route. If a connector often returns partial data, fix the connector rather than blaming the model.

Fast models tempt teams to do everything in a synchronous request handler. Resist that for jobs that can fan out, call tools, or wait on a human. A queue gives you backpressure, leases, retry discipline, and a place to record decisions.

async function runWorker(task: TaskPacket) {  const lease = await queue.claim(task.id, task.limits.deadlineMs);  if (!lease) return;
js
  try {    const answer = await claude.messages.create({      model: "claude-haiku-5-5",      max_tokens: 700,      // Set effort only after comparing your own fixtures.      thinking: { type: "adaptive", effort: "low" },      system: workerInstructions(task),      messages: [{ role: "user", content: JSON.stringify(task) }]    });
js
    const result = parseWorkerOutput(answer);    const decision = decide(result, task);
await events.append({ taskId: task.id, result, decision });    await dispatch(decision, task, result);  } catch (error) {    await retryOrEscalate(task, error);  } finally {    await queue.release(lease);  }}

The SDK details will evolve, so confirm current parameters against the Claude Platform documentation. The architecture should not: claim a job, do bounded work, validate outside the model, save a decision record, then dispatch the next controlled step.

The first mistake is using another model call as the only validator. A second model can be a helpful reviewer, especially for nuanced language, but it is not a substitute for a fact that code can check. If the task requires an account ID from supplied text, verify the ID against the supplied text. If it requires a permitted action, check the allow-list. Reserve model-to-model review for cases where the output still needs interpretation after deterministic checks have done their work.

The second mistake is retrying a failed task with the same packet and prompt until it happens to pass. Repeated retries can turn an ambiguous input into a fabricated answer and waste more money than a deliberate escalation. Give each task type a small attempt budget. Change something meaningful on retry: refresh an allowed source, shrink an overlarge packet, ask for a missing field, or send the task to the escalation lane. Record why the prior attempt failed so the next worker does not repeat the same dead end.

A useful rule is simple: retries repair transient systems; escalations repair uncertainty. Keep those paths separate in both code and reporting.

Build a small fixture set from real work before changing production traffic. Include ordinary cases, malformed inputs, ambiguous language, conflicting source records, missing attachments, prompt-injection attempts in source text, sensitive cases, and known historical failures.

For each fixture, define the expected terminal state: accept, escalate to a stronger model, or route to a person. Then compare three configurations: the current workflow, Haiku at a conservative effort setting, and Haiku with your acceptance gate. This reveals whether your savings come from real automation or from quietly accepting more mistakes.

Test the negative path as hard as the happy path. Delete a required field. Change an allowed label. Give the worker a plausible but unsupported claim. Delay a tool result. Send the same task twice. The queue should stop, preserve evidence, and route correctly every time.

Cost per million tokens is useful for planning. It is not the business metric. A cheap worker that creates customer follow-up, rework, or silent data errors is expensive.

Start with a small dashboard that tracks:

Do not aim for the lowest escalation rate. A sudden drop can mean the gate became too permissive. Aim for an escalation profile that matches the task’s consequence. Low-stakes extraction can tolerate a wider fast lane than account recovery or a production code change.

Begin in shadow mode. Let Haiku produce a result and decision record, but continue using your existing route. Compare the two on the same task stream. Next, enable acceptance only for a narrow task type with a strong deterministic check. Cap concurrency, set a budget, and create an immediate kill switch that sends new work to the previous path.

Review the first batch by failure reason, not by averages. “Ungrounded evidence” may mean your source packaging is poor. “Unknown label” may mean product vocabulary changed. “Timed out” may mean the task needs a smaller packet, not a more powerful model.

Only widen the route when the fixture suite and live outcomes agree. Repeat the comparison whenever you update the model snapshot, effort setting, prompt template, tool contract, or taxonomy. Fast-model economics can change quickly; your acceptance policy should be stable enough to absorb that change.

Claude Haiku 5.5 makes a large class of useful work inexpensive enough to run at volume. That is an opportunity to design better systems, not merely cheaper prompts. Give the worker a clear packet, small authority, a result it can prove, and a graceful exit when it cannot.

When the queue knows how to stop, you can safely let it move faster.

It is a strong fit for high-volume, latency-sensitive work with a narrow output contract: classification, extraction, summaries, routing, document checks, and clearly scoped subagent tasks. Confirm the current capability and pricing details in Anthropic’s official documentation before you ship.

No. Route based on consequence and verifiability, not only apparent simplicity. A short request that can change money, access, health, legal status, or production data may still require a stronger route or human review.

Escalate when the output fails schema, evidence, or business-rule checks; when required data is missing or conflicts; when the work exceeds its attempt budget; or when the task is outside the fast worker’s authority. Treat self-reported confidence as supplementary evidence, not a final gate.

Structured output makes parsing safer. It does not make the content true. Pair the schema with evidence checks, allow-lists, deterministic rules, idempotent writes, and an escalation path.

Track verified completion, false acceptance, escalation reason, queue latency, cost per verified completion, and outcome changes after model or policy updates. These show whether the fast lane is earning trust instead of merely lowering a token bill.

For a single low-risk synchronous classification, perhaps not. For work that fans out, calls tools, retries, waits on people, or can create side effects, a durable queue gives you backpressure, leases, traceability, and safe recovery.

Further reading: Anthropic’s Claude Haiku 5.5 announcement, Claude Haiku 5.5 model overview, and Anthropic’s ticket-routing guide.

Claude Haiku 5.5 Work Queues: Design High-Volume AI Pipelines That Know When to Stop was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #ai-agents 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-haiku-5-5-wor…] indexed:0 read:11min 2026-10-10 · —