cd /news/large-language-models/siliconflow-api-review-2026-setup-mo… Β· home β€Ί topics β€Ί large-language-models β€Ί article
[ARTICLE Β· art-146500] src=dev.to β†— pub= topic=large-language-models verified=true sentiment=Β· neutral

SiliconFlow API Review 2026: Setup, Models and Real Pricing

A developer published a hands-on review of SiliconFlow's OpenAI-compatible LLM API, finding that the platform's chat completions reference lists deprecated models such as zai-org/GLM-4.7 and GLM-4.6 while omitting live ones like deepseek-ai/DeepSeek-V4.1-Flash, and recommending developers query the /models endpoint instead. The review also documents automatic traffic rerouting β€” on June 11, 2026, calls to zai-org/GLM-5 and moonshotai/Kimi-K2.5 were silently redirected to GLM-5.1 and Kimi-K2.6 at the same price β€” warning that the model answering a request can change without any code change.

by read6 min views3 publishedOct 7, 2026

Verdict up front: SiliconFlow is a solid, OpenAI-compatible way to call open-weight LLMs (DeepSeek, Qwen, GLM, Kimi, MiniMax) from one key, with an Anthropic-compatible endpoint as a bonus. It is not automatically the cheapest host for every model. Check the specific model you need before you commit. Prices below were read off each provider's public pricing page on October 6, 2026.

I went looking for a current, practical write-up on SiliconFlow and mostly found launch posts. SiliconFlow's own dev.to account announces models as they arrive, which is useful, but its two newest posts there cover GLM-4.6V and GLM-4.7, and both have since been deprecated on the platform. So this is the guide I wanted: how to connect, how to find the model IDs that actually exist today, what it costs compared with the alternatives, and where it falls short.

One limit to know early: SiliconFlow is mostly a text and LLM platform. Its pricing page lists a handful of FLUX image models, two Wan 2.2 video models and two TTS models. For image, video and speech work I use synexa instead, which runs 100+ media models behind a single predictions endpoint. The rest of this post is about the LLM side, where SiliconFlow is strongest.

Setup is the standard OpenAI-compatible routine. Create an account, open API Keys in the console, generate a key, and point any OpenAI SDK at https://api.siliconflow.com/v1. New accounts get $1 in free credits, and billing after that is pay-as-you-go with no minimum commitment.

The first thing I do with any new provider is list the models through the API instead of trusting the marketing page:

export SILICONFLOW_API_KEY="sk-..."

curl -s "https://api.siliconflow.com/v1/models?sub_type=chat" \
  -H "Authorization: Bearer $SILICONFLOW_API_KEY" | jq -r '.data[].id'

Then a minimal Python call with the official openai package:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["SILICONFLOW_API_KEY"],
    base_url="https://api.siliconflow.com/v1",
)

resp = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4.1-Flash",
    messages=[{"role": "user", "content": "Summarise this changelog in 3 bullets: ..."}],
    max_tokens=512,
)
print(resp.choices[0].message.content)
print(resp.usage)  # log this; it's what you're billed on

Model IDs use the org/Model form (deepseek-ai/..., zai-org/..., moonshotai/..., Qwen/...), so copy them exactly from the /models output.

/models beats the API reference This is the most useful thing I found. The chat completions reference documents the model parameter with an enum of allowed values, and that list is out of date. It still includes zai-org/GLM-4.7, zai-org/GLM-4.6 and nex-agi/DeepSeek-V3.1-Nex-N1, which the release notes say were deprecated in May and June 2026. It doesn't include deepseek-ai/DeepSeek-V4.1-Flash, which has a live model page and price.

So treat the reference as a guide to parameters, not to models. Two parameters worth knowing about:

enable_thinking switches thinking mode on or off, but only for the models listed in the reference. Since it isn't a standard OpenAI field, pass it via extra_body in the Python SDK.thinking_budget caps chain-of-thought tokens for reasoning models (minimum 128, maximum 32,768, default 4,096). SiliconFlow retires models on a regular schedule and posts notices in its release notes. Sometimes it also reroutes traffic. On June 11, 2026, calls to zai-org/GLM-5 were automatically routed to zai-org/GLM-5.1, and moonshotai/Kimi-K2.5 to Kimi-K2.6, at the same price.

That's convenient, but it means the model answering your request can change while your code stays the same. If you run evals or care about reproducible output, pin the successor name yourself and watch the release notes.

Prices are per 1M tokens, shown as input / cached input / output. A dash means I couldn't find the model on that provider's pricing page.

Model SiliconFlow DeepSeek direct Together AI DeepInfra
DeepSeek V4.1 Flash $0.15 / $0.003 / $0.60 $0.15 / $0.003 / $0.60 off-peak; $0.30 / $0.006 / $1.20 peak $0.30 / $0.006 / $1.20 β€”
DeepSeek V4 Pro 0813 $1.32 / $0.044 / $3.96 $0.66 / $0.022 / $1.98 off-peak; $1.32 / $0.044 / $3.96 peak $1.32 / $0.13 / $3.96 β€”
DeepSeek V4 Flash 0731 $0.22 / $0.014 / $0.66 retired, served by V4.1 Flash $0.14 / $0.03 / $0.28 $0.06 / $0.015 / $0.18
GLM-5.3 $1.40 / $0.26 / $4.40 β€” $1.40 / $0.26 / $4.40 β€”
Kimi K3 $2.70 / $0.27 / $13.50 β€” $2.70 / $0.27 / $13.50 $2.85 / $0.285 / $14.25
Kimi K2.6 $0.77 / $0.14 / $3.40 β€” β€” $0.75 / $0.15 / $3.50
MiniMax M3 $0.30 / $0.06 / $1.20 β€” $0.30 / $0.06 / $1.20 β€”

What the table actually says:

If you'd rather keep several providers behind one key, OpenRouter says it passes provider prices through without markup and charges a 5.5% fee ($0.80 minimum) when you buy credits.

Limits are per account, not per key, and they apply per model. Hitting the limit on one model doesn't throttle the others. Your tier depends on monthly spend (the higher of last month and month-to-date) and upgrades automatically:

Tier RPM TPM
L0 (new accounts) 1,000 40,000
L2 2,000 80,000
L4 8,000 500,000
L5 10,000 2,000,000

The catch: many current models advertise context windows of around 1M tokens, but an L0 account gets 40,000 tokens per minute. If you plan to send whole repositories or long documents, budget for spend that moves you up a tier, or test early whether you'll hit 429s.

SiliconFlow also exposes an Anthropic-style Messages endpoint, and its docs include a Claude Code integration. The manual version is three environment variables:

export ANTHROPIC_BASE_URL="https://api.siliconflow.com/"
export ANTHROPIC_MODEL="your-preferred-model"   # a model ID from /models
export ANTHROPIC_API_KEY="sk-your-siliconflow-key"

The docs also offer a one-line setup script piped from a URL. Read it before you run it; the docs say so too.

Good fit: you want many open-weight LLMs behind one OpenAI-compatible key, your traffic peaks during DeepSeek's peak hours, or you want Claude Code-style tooling on top of open models.

Look elsewhere (or compare per model): you're cost-sensitive on older model revisions, you can batch DeepSeek Pro jobs off-peak, or your workload is mainly image and video generation.

Is SiliconFlow OpenAI-compatible?

Yes. Use base URL https://api.siliconflow.com/v1 with the OpenAI SDK. It supports most OpenAI parameters, plus its own fields like enable_thinking and thinking_budget.

Does SiliconFlow have a free tier?

New accounts get $1 in free credits. After that it's pay-as-you-go, with no minimum commitment.

What are SiliconFlow's rate limits?

They're tiered by monthly spend, starting at 1,000 RPM and 40,000 TPM for new accounts and going up to 10,000 RPM and 2,000,000 TPM at L5. Limits apply per account and per model.

Is SiliconFlow cheaper than calling DeepSeek directly?

For DeepSeek V4.1 Flash during DeepSeek's peak hours, yes. Off-peak it costs the same. For V4 Pro 0813, going direct off-peak costs half as much.

SiliconFlow earns its place as a single endpoint for current open-weight LLMs, and the official quickstart gets you to a first call in minutes. Just don't assume it's cheapest on every model: list models through the API, pin model names, and compare each model's price before you commit. For the media side of my projects (images, video, speech) I keep using synexa, where the same /v1/predictions request shape works across its model gallery.

── more in #large-language-models 4 stories Β· sorted by recency
── more on @siliconflow 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/siliconflow-api-revi…] indexed:0 read:6min 2026-10-07 Β· β€”