If your code already uses the OpenAI SDK, a gateway that speaks the OpenAI wire format requires exactly two changes:
from openai import OpenAI
client = OpenAI(
base_url="https://api.liurun.click/v1", # was "https://api.openai.com/v1"
api_key="sk-your-liurun-key", # was your OpenAI key
)
resp = client.chat.completions.create(
model="claude-sonnet-5", # any of 109 models, same call shape
messages=[{"role": "user", "content": "Explain WAL in Postgres"}],
)
print(resp.choices[0].message.content)
That model string is now the only knob between vendors. A fallback chain becomes a list, not a architecture project:
MODELS = ["claude-sonnet-5", "gpt-5.5", "deepseek-chat"] # try in order
curl https://api.liurun.click/v1/chat/completions \
-H "Authorization: Bearer sk-your-liurun-key" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-2.5-flash",
"messages": [{"role": "user", "content": "Say hi in 3 words"}],
"stream": true
}'
Streaming is standard SSE — data: lines, data: [DONE] sentinel, nothing exotic.
The gateway also accepts the Claude-format /v1/messages endpoint. So the Anthropic SDK works with a base URL swap:
import anthropic
client = anthropic.Anthropic(
base_url="https://api.liurun.click",
api_key="sk-your-liurun-key",
)
msg = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude"}],
)
Every request gets logged with its exact token counts and its exact cost. Not "credits" — dollars and cents, computed from a per-model price table that's published in full on a public pricing page:
| Model | Input $/1M | Output $/1M |
|---|---|---|
| Claude Sonnet 5 | 1.45 | 7.27 |
| Claude Opus 4.6 | 3.63 | 18.16 |
| GPT-5.5 | 2.72 | 16.35 |
| Gemini 2.5 Flash | 0.16 | 1.36 |
| DeepSeek Chat | 0.36 | 1.45 |
| Grok 4.5 | 1.09 | 3.27 |
| GPT-6 Luna (budget) | 0.05 | 0.27 |
Prompt-cache reads on Claude models bill at 10% of the input rate — cache-friendly agents stop being a budget gamble. Image models bill per image (GPT-Image-2 ≈ $0.051/image).
Billing is prepaid: you top up a USD balance from $1 and usage draws it down. No subscription, no seats, no minimum. If you stop liking the service, your remaining balance is the only thing at risk — and there's no contract to cancel.
The pattern I use in my own projects now:
def complete(messages, *, tier="smart", **kw):
models = {
"smart": ["claude-sonnet-5", "gpt-5.5"],
"cheap": ["gemini-2.5-flash", "deepseek-chat"],
"budget": ["gpt-6-luna"],
}[tier]
last = None
for m in models:
try:
return client.chat.completions.create(
model=m, messages=messages, **kw)
except Exception as e:
last = e
raise last
Same key, same endpoint, model choice becomes a cost/quality dial instead of an integration decision.
The gateway runs on New API, which is open source — if you'd rather keep everything in-house, the two-line migration above works against your own deployment too. I self-host the hosted version on a small AWS instance in Tokyo behind CloudFront; a t4g.small handles it comfortably.
Links: Sign up · Pricing · User Agreement
(Disclosure: I built and operate LiuRun API. The migration pattern above applies to any OpenAI-compatible gateway, self-hosted included.)