cd /news/ai-products/kimi-k3-and-kimi-k3-fast-with-zdr-an… · home topics ai-products article
[ARTICLE · art-75573] src=vercel.com ↗ pub= topic=ai-products verified=true sentiment=· neutral

Kimi K3 and Kimi K3 Fast with ZDR and US-based providers now on AI Gateway

Moonshot AI's Kimi K3 and Kimi K3 Fast models are now available on Vercel's AI Gateway from US-based providers including Baseten and Fireworks, with Zero Data Retention (ZDR) support. The models run on US infrastructure for data residency compliance, and AI Gateway automatically routes across providers for failover and higher uptime. Kimi K3 Fast costs ~50% more per token for lower latency, and regional pricing for US-only inference is ~10% higher.

read2 min views1 publishedJul 27, 2026

Kimi K3 from Moonshot AI and its faster serving path, Kimi K3 Fast, are now available from US-based providers on AI Gateway, including Baseten and Fireworks. Zero Data Retention (ZDR) is also supported for both models.

As of 8:20 AM PDT, Fireworks is available. Baseten and other providers will be online shortly.

Running Kimi K3 on US-based providers lets teams with data residency and compliance requirements use the model on US infrastructure. Because AI Gateway serves the models from multiple providers, it automatically routes across them for failover, higher uptime, and more available throughput than any single provider offers. You call the same moonshotai/kimi-k3

model ID, and the gateway handles provider selection and fallback.

Kimi K3 Fast trades a higher per-token cost for lower latency. Request it with the speed

option on the base model, which stays on moonshotai/kimi-k3

and falls back to standard speed when the fast tier is unavailable. Alternatively, use moonshotai/kimi-k3-fast

. The fast variant costs ~50% more than the base model.

To use Kimi K3, set model

to moonshotai/kimi-k3

in the AI SDK: To route Kimi K3 requests to use only US data centers for inference, set inferenceRegion

. Regional pricing is ~10% more than the regular variant.

Zero Data Retention for Kimi K3 is also available. Turn on Zero Data Retention for every request from the AI Gateway dashboard settings, or set it per request with zeroDataRetention

:

To see every provider serving Kimi K3, along with per-provider pricing, supported parameters, uptime, throughput, and latency, call the model endpoints API:

Model prices vary by provider and variant type.

Run vercel ai-gateway coding-agents setup

and select Kimi K3. This will detect the agents on your machine, provision an AI Gateway key, and write their config. See how to set it up via the Vercel CLI.

Try Kimi K3 in the model playground.

── more in #ai-products 4 stories · sorted by recency
── more on @moonshot ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/kimi-k3-and-kimi-k3-…] indexed:0 read:2min 2026-07-27 ·