As developers, we’ve all been there: you build an AI-powered feature, deploy it to production, and suddenly get slapped with a massive OpenAI or Anthropic API bill because a user ran a recursive loop, or because your team shared raw API keys across staging environments. 💸
Sharing master API keys is a major security risk, and tracking individual user costs across multiple models (GPT-4o, Claude 3.5 Sonnet, Llama-3) is a management nightmare.
There is a better way. By self-hosting LiteLLM, you can spin up a unified, private LLM Gateway that acts as a proxy between your applications and your AI providers.
With LiteLLM, you get:
In this guide, we'll deploy a production-ready LiteLLM Instance with a beautiful UI Admin Dashboard on a high-speed Vultr High-Performance Cloud Instance in under 10 minutes.
When proxying API requests, latency is everything. Adding a proxy layer shouldn't introduce lag. Vultr's High-Performance Cloud offers high-frequency CPU cores and NVMe storage in over 32 global locations. This means your proxy can be co-located right next to your target audience (or your main application servers) to ensure sub-millisecond routing overhead.
Once your instance is ready, copy the IP address and connect to it via your terminal:
ssh root@YOUR_VULTR_IP
Update your system packages and install Docker to manage our containers easily:
sudo apt update && sudo apt upgrade -y
sudo apt install -y docker.io docker-compose
Verify the installation:
docker --version && docker-compose --version
Create a dedicated directory for LiteLLM:
mkdir litellm-gateway && cd litellm-gateway
Now, create a litellm_config.yaml file. This is where you configure the models you want to proxy. For this demo, we'll configure OpenAI's GPT-4o and Anthropic's Claude 3.5 Sonnet:
model_list:
- model_name: gpt-4o
litellm_params:
model: openai/gpt-4o
api_key: "os.environ/OPENAI_API_KEY"
- model_name: claude-3-5-sonnet
litellm_params:
model: anthropic/claude-3-5-sonnet-20240620
api_key: "os.environ/ANTHROPIC_API_KEY"
router_settings:
routing_strategy: latency-based-routing
general_settings:
master_key: "os.environ/LITELLM_MASTER_KEY"
We will set up three components:
Create a docker-compose.yml file in the same directory:
version: '3.8'
services:
db:
image: postgres:16-alpine
container_name: litellm-db
environment:
POSTGRES_DB: litellm
POSTGRES_USER: litellm_user
POSTGRES_PASSWORD: super_secret_db_password
volumes:
- pgdata:/var/lib/postgresql/data
ports:
- "5432:5432"
restart: always
litellm:
image: ghcr.io/berriai/litellm:main-latest
container_name: litellm-proxy
ports:
- "4000:4000"
volumes:
- ./litellm_config.yaml:/app/config.yaml
environment:
- DATABASE_URL=postgresql://litellm_user:super_secret_db_password@db:5432/litellm
- LITELLM_MASTER_KEY=sk-your-super-secure-master-admin-key-12345
- OPENAI_API_KEY=your_actual_openai_api_key_here
- ANTHROPIC_API_KEY=your_actual_anthropic_api_key_here
depends_on:
- db
command: ["--config", "/app/config.yaml", "--detailed_debug"]
restart: always
volumes:
pgdata:
💡 Note: Replace your_actual_openai_api_key_here and your_actual_anthropic_api_key_here with your real API keys, and change the LITELLM_MASTER_KEY to a secure, random string.
Run the Docker stack in detached mode:
docker-compose up -d
Check if everything is running correctly:
docker-compose ps
Your gateway is now active!
http://YOUR_VULTR_IP:4000/ui. LITELLM_MASTER_KEY you configured in your docker-compose.yml (e.g., sk-your-super-secure-master-admin-key-12345).
From this dashboard, you can:
To consume your self-hosted gateway, you simply point your existing OpenAI SDK to your Vultr server. No code changes required!
from openai import OpenAI
client = OpenAI(
api_key="sk-the-virtual-key-you-created-in-dashboard",
base_url="http://YOUR_VULTR_IP:4000"
)
response = client.chat.completions.create(
model="claude-3-5-sonnet",
messages=[{"role": "user", "content": "Explain quantum computing in 2 sentences."}]
)
print(response.choices[0].message.content)
Before pointing production applications to your gateway, ensure you secure it with SSL. You can easily install Nginx and Certbot on your Vultr instance to get a free Let's Encrypt SSL certificate:
sudo apt install -y nginx certbot python3-certbot-nginx
Configure Nginx to reverse proxy traffic from port 80/ 443 to http://localhost:4000 and run sudo certbot --nginx to enable HTTPS.
Ready to get your team's AI costs under lock and key? Build your high-performance, private LLM gateway today on Vultr High-Performance Cloud!
Liked this resource? Join our daily Telegram channel for more developer tools and cloud insights: @Libretech2026