SiliconFlow API Review 2026: Setup, Models and Real Pricing A developer published a hands-on review of SiliconFlow's OpenAI-compatible LLM API, finding that the platform's chat completions reference lists deprecated models such as zai-org/GLM-4.7 and GLM-4.6 while omitting live ones like deepseek-ai/DeepSeek-V4.1-Flash, and recommending developers query the /models endpoint instead. The review also documents automatic traffic rerouting — on June 11, 2026, calls to zai-org/GLM-5 and moonshotai/Kimi-K2.5 were silently redirected to GLM-5.1 and Kimi-K2.6 at the same price — warning that the model answering a request can change without any code change. Verdict up front: SiliconFlow is a solid, OpenAI-compatible way to call open-weight LLMs DeepSeek, Qwen, GLM, Kimi, MiniMax from one key, with an Anthropic-compatible endpoint as a bonus. It is not automatically the cheapest host for every model. Check the specific model you need before you commit. Prices below were read off each provider's public pricing page on October 6, 2026. I went looking for a current, practical write-up on SiliconFlow and mostly found launch posts. SiliconFlow's own dev.to account announces models as they arrive, which is useful, but its two newest posts there cover GLM-4.6V and GLM-4.7, and both have since been deprecated on the platform. So this is the guide I wanted: how to connect, how to find the model IDs that actually exist today, what it costs compared with the alternatives, and where it falls short. One limit to know early: SiliconFlow is mostly a text and LLM platform. Its pricing page lists a handful of FLUX image models, two Wan 2.2 video models and two TTS models. For image, video and speech work I use synexa https://synexa.ai/?utm source=devto&utm medium=ugc&utm campaign=gretavolkov&utm content=siliconflow-intro instead, which runs 100+ media models behind a single predictions endpoint. The rest of this post is about the LLM side, where SiliconFlow is strongest. Setup is the standard OpenAI-compatible routine. Create an account, open API Keys in the console, generate a key, and point any OpenAI SDK at https://api.siliconflow.com/v1 . New accounts get $1 in free credits, and billing after that is pay-as-you-go with no minimum commitment. The first thing I do with any new provider is list the models through the API instead of trusting the marketing page: export SILICONFLOW API KEY="sk-..." Chat models only; other sub type values: embedding, reranker, text-to-image, image-to-image, speech-to-text, text-to-video curl -s "https://api.siliconflow.com/v1/models?sub type=chat" \ -H "Authorization: Bearer $SILICONFLOW API KEY" | jq -r '.data .id' Then a minimal Python call with the official openai package: python import os from openai import OpenAI client = OpenAI api key=os.environ "SILICONFLOW API KEY" , base url="https://api.siliconflow.com/v1", resp = client.chat.completions.create model="deepseek-ai/DeepSeek-V4.1-Flash", messages= {"role": "user", "content": "Summarise this changelog in 3 bullets: ..."} , max tokens=512, print resp.choices 0 .message.content print resp.usage log this; it's what you're billed on Model IDs use the org/Model form deepseek-ai/... , zai-org/... , moonshotai/... , Qwen/... , so copy them exactly from the /models output. /models beats the API reference This is the most useful thing I found. The chat completions reference documents the model parameter with an enum of allowed values, and that list is out of date. It still includes zai-org/GLM-4.7 , zai-org/GLM-4.6 and nex-agi/DeepSeek-V3.1-Nex-N1 , which the release notes say were deprecated in May and June 2026. It doesn't include deepseek-ai/DeepSeek-V4.1-Flash , which has a live model page and price. So treat the reference as a guide to parameters, not to models. Two parameters worth knowing about: enable thinking switches thinking mode on or off, but only for the models listed in the reference. Since it isn't a standard OpenAI field, pass it via extra body in the Python SDK. thinking budget caps chain-of-thought tokens for reasoning models minimum 128, maximum 32,768, default 4,096 . SiliconFlow retires models on a regular schedule and posts notices in its release notes. Sometimes it also reroutes traffic. On June 11, 2026, calls to zai-org/GLM-5 were automatically routed to zai-org/GLM-5.1 , and moonshotai/Kimi-K2.5 to Kimi-K2.6 , at the same price. That's convenient, but it means the model answering your request can change while your code stays the same. If you run evals or care about reproducible output, pin the successor name yourself and watch the release notes. Prices are per 1M tokens, shown as input / cached input / output. A dash means I couldn't find the model on that provider's pricing page. | Model | SiliconFlow | DeepSeek direct | Together AI | DeepInfra | |---|---|---|---|---| | DeepSeek V4.1 Flash | $0.15 / $0.003 / $0.60 | $0.15 / $0.003 / $0.60 off-peak; $0.30 / $0.006 / $1.20 peak | $0.30 / $0.006 / $1.20 | — | | DeepSeek V4 Pro 0813 | $1.32 / $0.044 / $3.96 | $0.66 / $0.022 / $1.98 off-peak; $1.32 / $0.044 / $3.96 peak | $1.32 / $0.13 / $3.96 | — | | DeepSeek V4 Flash 0731 | $0.22 / $0.014 / $0.66 | retired, served by V4.1 Flash | $0.14 / $0.03 / $0.28 | $0.06 / $0.015 / $0.18 | | GLM-5.3 | $1.40 / $0.26 / $4.40 | — | $1.40 / $0.26 / $4.40 | — | | Kimi K3 | $2.70 / $0.27 / $13.50 | — | $2.70 / $0.27 / $13.50 | $2.85 / $0.285 / $14.25 | | Kimi K2.6 | $0.77 / $0.14 / $3.40 | — | — | $0.75 / $0.15 / $3.50 | | MiniMax M3 | $0.30 / $0.06 / $1.20 | — | $0.30 / $0.06 / $1.20 | — | What the table actually says: If you'd rather keep several providers behind one key, OpenRouter says it passes provider prices through without markup and charges a 5.5% fee $0.80 minimum when you buy credits. Limits are per account, not per key, and they apply per model. Hitting the limit on one model doesn't throttle the others. Your tier depends on monthly spend the higher of last month and month-to-date and upgrades automatically: | Tier | RPM | TPM | |---|---|---| | L0 new accounts | 1,000 | 40,000 | | L2 | 2,000 | 80,000 | | L4 | 8,000 | 500,000 | | L5 | 10,000 | 2,000,000 | The catch: many current models advertise context windows of around 1M tokens, but an L0 account gets 40,000 tokens per minute. If you plan to send whole repositories or long documents, budget for spend that moves you up a tier, or test early whether you'll hit 429s. SiliconFlow also exposes an Anthropic-style Messages endpoint, and its docs include a Claude Code integration. The manual version is three environment variables: export ANTHROPIC BASE URL="https://api.siliconflow.com/" export ANTHROPIC MODEL="your-preferred-model" a model ID from /models export ANTHROPIC API KEY="sk-your-siliconflow-key" The docs also offer a one-line setup script piped from a URL. Read it before you run it; the docs say so too. Good fit: you want many open-weight LLMs behind one OpenAI-compatible key, your traffic peaks during DeepSeek's peak hours, or you want Claude Code-style tooling on top of open models. Look elsewhere or compare per model : you're cost-sensitive on older model revisions, you can batch DeepSeek Pro jobs off-peak, or your workload is mainly image and video generation. Is SiliconFlow OpenAI-compatible? Yes. Use base URL https://api.siliconflow.com/v1 with the OpenAI SDK. It supports most OpenAI parameters, plus its own fields like enable thinking and thinking budget . Does SiliconFlow have a free tier? New accounts get $1 in free credits. After that it's pay-as-you-go, with no minimum commitment. What are SiliconFlow's rate limits? They're tiered by monthly spend, starting at 1,000 RPM and 40,000 TPM for new accounts and going up to 10,000 RPM and 2,000,000 TPM at L5. Limits apply per account and per model. Is SiliconFlow cheaper than calling DeepSeek directly? For DeepSeek V4.1 Flash during DeepSeek's peak hours, yes. Off-peak it costs the same. For V4 Pro 0813, going direct off-peak costs half as much. SiliconFlow earns its place as a single endpoint for current open-weight LLMs, and the official quickstart https://docs.siliconflow.com/en/userguide/quickstart gets you to a first call in minutes. Just don't assume it's cheapest on every model: list models through the API, pin model names, and compare each model's price before you commit. For the media side of my projects images, video, speech I keep using synexa https://synexa.ai/?utm source=devto&utm medium=ugc&utm campaign=gretavolkov&utm content=siliconflow-outro , where the same /v1/predictions request shape works across its model gallery.