Use one compatible API, switch models without rewriting your integration, and pay for the tokens you actually use. /api/v1/chat/completions
- Prompt Caching
- Spend Tracking
- Never Downgraded
glm-5.2
$0.0420
USD per 1M input tokens
gpt-5.6-luna
$0.0500
USD per 1M input tokens
glm-5.3-flash
$0.0555
USD per 1M input tokens
deepseek-v4-flash
$0.0660
USD per 1M input tokens
deepseek-v4-flash-vision-exp
$0.0660
USD per 1M input tokens
minimax-m3
$0.0900
USD per 1M input tokens
gpt-5.6-terra
$0.1200
USD per 1M input tokens
glm-5.3
$0.1260
USD per 1M input tokens
gemini-3.7-flash
$0.1350
USD per 1M input tokens
qwen3.8-2.4t-a95b
$0.1400
USD per 1M input tokens
kimi-k3
$0.1500
USD per 1M input tokens
deepseek-v4-pro
$0.1980
USD per 1M input tokens
gpt-5.6-sol
$0.2000
USD per 1M input tokens
sonnet-5
$0.3000
USD per 1M input tokens
grok-4.6
$0.4000
USD per 1M input tokens
opus-5
$0.4500
USD per 1M input tokens
fable-5
$1.5000
USD per 1M input tokens
Compare the numbers. #
| Relative | |||||
|---|---|---|---|---|---|
| GLM 5.2 | 0.0420 | 0.0078 | 0.1320 | 0.05106 | |
| GPT 5.6 Luna | 0.0500 | 0.0050 | 0.3000 | 0.0935 | |
| GLM 5.3 Flash | 0.0555 | 0.0111 | 0.1850 | 0.07067 | |
| DeepSeek V4 Flash 0731 | 0.0660 | 0.0066 | 0.1980 | 0.07392 | |
| DeepSeek V4 Flash Vision Exp | 0.0660 | 0.0066 | 0.1980 | 0.07392 |
Blended assumes 70% of input served from cache and output equal to 25% of input volume — a reference mix for comparison only, not a billed rate. The gateway records usage after a successful response; rates shown are the currently configured rates.
Change one base URL. #
Keep the request shape your application already understands. Change the model ID when the workload changes.
1endpoint
one base URL
Chat CompletionsPOST /chat/completions
ResponsesPOST /responses
MessagesPOST /messages
Cache decides the bill. #
Message 12 of a conversation costs 3.7× less than it would without cache — because every message resends everything before it, and we charge full price only for the part that is new.
GLM 5.2 rates, USD per 1M tokens — a cache hit is billed at 5× less than a miss. Illustration: a conversation of 12 messages, each one request that adds 10,000 input and 1,000 output tokens and resends everything before it, priced per 1,000 requests. Token counts are illustrative; the rates are published. Cache ratio differs per model.
Spend less on every token. #
Pay only for what you use. Input starts at $0.0420 per 1M tokens — lower cache rates are applied automatically, per model.
Usage-based pricing · 1,000 credits = $1