# Stop managing six AI vendor accounts: one key for 109 models

> Source: <https://dev.to/2048lr/stop-managing-six-ai-vendor-accounts-one-key-for-109-models-194a>
> Published: 2026-10-06 06:12:09+00:00

If your code already uses the OpenAI SDK, a gateway that speaks the OpenAI wire format requires exactly two changes:

``` python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.liurun.click/v1",   # was "https://api.openai.com/v1"
    api_key="sk-your-liurun-key",             # was your OpenAI key
)

resp = client.chat.completions.create(
    model="claude-sonnet-5",                  # any of 109 models, same call shape
    messages=[{"role": "user", "content": "Explain WAL in Postgres"}],
)
print(resp.choices[0].message.content)
```

That `model` string is now the only knob between vendors. A fallback chain becomes a list, not a architecture project:

```
MODELS = ["claude-sonnet-5", "gpt-5.5", "deepseek-chat"]  # try in order
curl https://api.liurun.click/v1/chat/completions \
  -H "Authorization: Bearer sk-your-liurun-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-2.5-flash",
    "messages": [{"role": "user", "content": "Say hi in 3 words"}],
    "stream": true
  }'
```

Streaming is standard SSE — `data:` lines, `data: [DONE]` sentinel, nothing exotic.

The gateway also accepts the Claude-format `/v1/messages` endpoint. So the Anthropic SDK works with a base URL swap:

``` python
import anthropic

client = anthropic.Anthropic(
    base_url="https://api.liurun.click",
    api_key="sk-your-liurun-key",
)

msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello, Claude"}],
)
```

Every request gets logged with its exact token counts and its exact cost. Not "credits" — dollars and cents, computed from a per-model price table that's published in full on a public pricing page:

| Model | Input $/1M | Output $/1M | 
|---|---|---|
| Claude Sonnet 5 | 1.45 | 7.27 | 
| Claude Opus 4.6 | 3.63 | 18.16 | 
| GPT-5.5 | 2.72 | 16.35 | 
| Gemini 2.5 Flash | 0.16 | 1.36 | 
| DeepSeek Chat | 0.36 | 1.45 | 
| Grok 4.5 | 1.09 | 3.27 | 
| GPT-6 Luna (budget) | 0.05 | 0.27 | 

Prompt-cache reads on Claude models bill at 10% of the input rate — cache-friendly agents stop being a budget gamble. Image models bill per image (GPT-Image-2 ≈ $0.051/image).

Billing is prepaid: you top up a USD balance from $1 and usage draws it down. No subscription, no seats, no minimum. If you stop liking the service, your remaining balance is the only thing at risk — and there's no contract to cancel.

The pattern I use in my own projects now:

``` python
def complete(messages, *, tier="smart", **kw):
    models = {
        "smart":  ["claude-sonnet-5", "gpt-5.5"],
        "cheap":  ["gemini-2.5-flash", "deepseek-chat"],
        "budget": ["gpt-6-luna"],
    }[tier]
    last = None
    for m in models:
        try:
            return client.chat.completions.create(
                model=m, messages=messages, **kw)
        except Exception as e:
            last = e
    raise last
```

Same key, same endpoint, model choice becomes a cost/quality dial instead of an integration decision.

The gateway runs on [New API](https://github.com/QuantumNous/new-api), which is open source — if you'd rather keep everything in-house, the two-line migration above works against your own deployment too. I self-host the hosted version on a small AWS instance in Tokyo behind CloudFront; a t4g.small handles it comfortably.

**Links:** [Sign up](https://api.liurun.click/register) · [Pricing](https://api.liurun.click/pricing) · [User Agreement](https://api.liurun.click/user-agreement)

*(Disclosure: I built and operate LiuRun API. The migration pattern above applies to any OpenAI-compatible gateway, self-hosted included.)*
