Haiku 5.5 Slashes Costs; New Agents Reshape Docs Anthropic's Claude Haiku 5.5 cuts per-token pricing by 90% for requests under 100k tokens, but a new tokenizer inflates token counts roughly 30% per request and an adjustable effort dial can swing actual bills, so the writeup recommends validating jobs against the new pricing tiers before migrating production Haiku 4.5 workloads. The same roundup notes StepFun's Step 5 Preview (1M-token context) and xAI's text-to-video model are now routable through Vercel's AI Gateway, and that Prisma Compute reached general availability with per-branch preview databases and deploy-time schema migrations. This week's tooling news splits neatly into two themes: pricing traps disguised as savings, and a quiet consolidation of agentic infrastructure under Vercel's AI Gateway umbrella. Both deserve more scrutiny than the headlines suggest. Claude Haiku 5.5 drops per-token pricing by 90% for requests under 100k tokens. That number will pull people into a migration without reading the footnotes—and the footnotes matter. The new tokenizer inflates token counts roughly 30% per request, and the effort dial think: compute intensity per task can swing your actual bill in either direction. List price is a floor, not a ceiling. Why it matters now: If you're running Haiku 4.5 workloads in production, you're probably doing rough napkin math on whether to migrate. That math is wrong until you recount tokens against Haiku 5.5's tokenizer and stress-test the effort dial against your actual task distribution. Cost modeling on the old tokenizer will underestimate spend. Verdict: Ship—but gate it first. The included TypeScript utility validates jobs against Haiku 5.5's pricing tiers before the API call hits. It forces you to declare task type, effort level, and token count upfront, replacing ad hoc mental math. Requires Node 22+, zero dependencies. Run it against your existing Haiku 4.5 workloads before touching production config. StepFun's Step 5 Preview is now routable through Vercel's AI Gateway, bringing a 1M-token context window into the same unified API layer as Claude Code and standard chat completions. The integration follows the standard gateway pattern: swap model names, keep your client code. Why it matters now: A 1M-token context window is a qualitatively different tool than a 200k window. Entire codebases, full document stacks, long conversation histories—these fit in a single request, which eliminates a class of context-chunking logic that's currently living in your application layer. If you're writing retrieval or chunking code to work around context limits, this is worth benchmarking. Verdict: Evaluate. If you're already inside the Vercel ecosystem, adding Step 5 Preview is a config change. If you're not, weigh gateway latency and cost overhead against calling StepFun directly. Don't adopt the gateway abstraction just for this model unless the consolidation story actually applies to your stack. xAI's text-to-video model lands on AI Gateway with native audio support, 480p–1080p resolution output, and seven aspect ratios. It's reachable via the standard AI SDK or CLI with no new provider authentication if you're already using BYOK. Why it matters now: The interesting part isn't the video model—it's the pipeline implication. You can now chain image and video generation in a single SDK call without stitching together separate provider integrations or managing divergent auth flows. For teams building content generation tooling, that's a meaningful reduction in orchestration code. Verdict: Ship. Production-grade, playground available for testing, no new auth overhead if you're already on the gateway. If video generation is in scope for your project, there's no reason to wait. Prisma Compute is now GA. The core mechanic: declare your TypeScript app and Prisma Postgres instance as a single unit; get automatic preview environments per Git branch with isolated databases and schema migrations applied at deploy time. Environment-specific connection strings are generated at runtime from declarative config, not copy-pasted across dashboards. Why it matters now: The preview database story is the real unlock. Schema changes tested against isolated branch databases before they touch production is not a new idea, but operationalizing it typically requires custom deploy scripts and careful environment variable management. Prisma Compute collapses that into a GitHub Actions PR merge. For teams doing rapid schema iteration, that's a meaningful reduction in foot-gun surface area. Verdict: Evaluate. If your stack is TypeScript/Node/Bun/Next.js and you're already using Prisma ORM, this is worth trying now—the migration cost is low and the branch preview workflow pays for itself quickly. If you're not using Prisma ORM, factor in the schema refactor cost. Requires Bun runtime, which may be a blocker depending on your deployment environment. Glyph Cluster is a reasoning model aimed at code review, multi-step analysis, and debugging—available via AI SDK, OpenAI Chat Completions API, and coding agents. It supports function calling, which means it can chain reasoning steps with custom tools. It's free during stealth. Why it matters now: The function calling support is what elevates this above a standard model swap. Agents that need to reason through a problem and then execute tool calls can wire Glyph Cluster in without a separate orchestration layer. The setup is minimal: change the model name to stealth/glyph-cluster and authenticate through the gateway. Verdict: Evaluate. Free during stealth makes the experimentation cost zero, but confirm one thing before you do: Zero Data Retention is not available, which matters if your code review workloads touch proprietary or sensitive code. Evaluate training data implications against your use case before routing production workloads through it. v0 now auto-loads Resend, Clerk, MongoDB, Algolia, and OpenSearch directly in the chat interface. It handles environment variable injection and loads provider-specific agent skills without leaving the v0 workflow. The integration replaces manual marketplace setup in the Vercel dashboard. Why it matters now: The friction reduction is real but scoped. If you're already using v0 to scaffold features, eliminating the credential setup and config boilerplate steps matters at the margin. If v0 isn't already in your workflow, this doesn't change the calculus much—the five supported providers are common but not universal. Verdict: Ship if you're already in v0. The workflow improvement is immediate and requires no setup. If you're not using v0 for code generation, the provider integration story alone isn't a reason to start. If you're building with AI tooling and want analysis like this in your inbox every week, subscribe at thedevsignal.com https://thedevsignal.com —it's written for engineers who want signal, not summaries. Each issue covers what to ship, what to wait on, and why the difference matters.