{"slug": "prism-inference", "title": "Prism Inference", "summary": "Prism Inference launched an OpenAI- and Anthropic-compatible inference API serving open-weight frontier models including DeepSeek-V4.1-Flash, DeepSeek-V4-Flash, GLM-5.3, Kimi K3, Qwen3.8, and Qwen, with the company claiming up to 5.8× faster throughput on DeepSeek V4.1 at 550 tok/s versus 95–247 tok/s on major providers. Prism, backed by Y Combinator, prices usage per token with no per-seat fee and claims up to 50% less than major cloud providers, offering a 99.99% uptime SLA and sub-50ms P50 latency across global points of presence. The company states it applies strict zero data retention to API inputs and outputs and never trains on them, and offers dedicated private deployments in its cloud or the customer's.", "body_md": "[Backed byCombinator](https://www.ycombinator.com/companies/prism)\n\n# Lightning fast inferenceoptimized for you\n\nAccess to DeepSeek-V4.1-Flash, GLM-5.3, Kimi K3, and Qwen3.8 through OpenAI-compatible and Anthropic-compatible APIs.\n\n02 / Lightning fast inference\n\nDeepSeek V4.1\n\n550tok/s\n\nup to 5.8× faster\n\n95–247tok/s\n\n## Prism Engine\n\n$ prism chat --model deepseek-v4.1\n\n›\n\n## Major Providers\n\n$ chat --model deepseek-v4.1\n\n›\n\n03 / Serverless Model APIs\n\n## Run open source models through a single API.No infrastructure to manage.\n\nOne platform for the full AI inference stack\n\nModel APIs, GPU infrastructure, and agent runtimes. All in one platform.\n\nModel APIs\n\nGPU infrastructure\n\nAgent runtimes\n\nTop Open Source models\n\nDeepSeek-V4.1-Flash, DeepSeek-V4-Flash, GLM-5.3, Kimi K3, Qwen3.8, and Qwen on one key.\n\nBetter price-performance\n\nUp to 50% less than major cloud providers. Not because we cut corners, because we built the infrastructure.\n\nScale with your workload\n\nStart small and scale seamlessly from APIs to dedicated clusters.\n\nServerless API\n\nDedicated endpoints\n\nGPU control\n\nBuilt for production reliability\n\nStable infrastructure with low latency, high throughput, and reliable uptime at scale.\n\n99.99%\n\nUptime SLA\n\n<50ms\n\nP50 latency\n\nGlobal\n\nPoPs\n\n04 / Questions\n\n## What is Prism?\n\nPrism is an OpenAI- and Anthropic-compatible inference API for coding agents. It serves open-weight frontier models so you can run the primary agent loop without standing up your own GPU stack.\n\n## Which models can I call?\n\nThe public lineup is [DeepSeek-V4.1-Flash](https://prisminference.com/models?model=deepseek-v4.1-flash), [DeepSeek-V4-Flash](https://prisminference.com/models?model=deepseek-v4-flash), [GLM-5.3](https://prisminference.com/models?model=glm-5.3), [Kimi K3](https://prisminference.com/models?model=kimi-k3), [Qwen3.8](https://prisminference.com/models?model=qwen-3.8), and [Qwen](https://prisminference.com/models?model=qwen). Currently available model ids are `deepseek-v4.1-flash`, `deepseek-v4-flash`, `glm-5.3`, `kimi-k3`, `qwen-3.8`, and `qwen`.\n\n## Which API formats are supported?\n\nPrism supports OpenAI Chat Completions at `https://api.prisminference.com/v1` and Anthropic Messages at `https://api.prisminference.com`. Keep your existing client and change the base URL, API key, and model id.\n\n## How do I get an API key?\n\n[Create a Prism account](https://prisminference.com/signup), then create an API key from your account settings. Store it as `PRISM_API_KEY` and send it with your requests.\n\n## How is inference priced?\n\nUsage is billed per token. There is no per-seat fee. See [Pricing](https://prisminference.com/pricing) for serverless, elastic, dedicated, and batch options.\n\n## Which coding agents work with Prism?\n\nAny client or gateway that can send OpenAI Chat Completions or Anthropic Messages, including Cursor, Claude Code, and configurable custom loops. Clients that require the OpenAI Responses API need a compatibility gateway.\n\n## Does Prism retain prompts or train on them?\n\nNo. Prism applies strict zero data retention to API inputs and outputs and never uses them for training. See the [Privacy Policy](https://prisminference.com/privacy) for details.\n\n## Can I get a private deployment?\n\nYes. [Book a call](https://cal.com/team/prismai/demo) for dedicated capacity in our cloud or yours.\n\n05 / Get a key", "url": "https://wpnews.pro/news/prism-inference", "canonical_source": "https://prisminference.com", "published_at": "2026-09-25 06:57:29+00:00", "updated_at": "2026-09-25 07:30:08.860315+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-products", "large-language-models", "ai-agents", "developer-tools"], "entities": ["Prism Inference", "Y Combinator", "DeepSeek-V4.1-Flash", "DeepSeek-V4-Flash", "GLM-5.3", "Kimi K3", "Qwen3.8", "Claude Code"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/prism-inference", "markdown": "https://wpnews.pro/news/prism-inference.md", "text": "https://wpnews.pro/news/prism-inference.txt", "jsonld": "https://wpnews.pro/news/prism-inference.jsonld"}}