cd /news/ai-infrastructure/stop-bleeding-ai-costs-self-host-you… · home › topics › ai-infrastructure › article
[ARTICLE · art-144460] src=dev.to ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Stop Bleeding AI Costs: Self-Host Your Own Private LLM Gateway (LiteLLM) in Under 10 Minutes

A developer published a step-by-step guide for self-hosting LiteLLM as a private LLM gateway on a Vultr high-performance cloud instance, using Docker Compose to run the LiteLLM proxy alongside a PostgreSQL database. The setup proxies requests to OpenAI's GPT-4o and Anthropic's Claude 3.5 Sonnet, centralizes API key management, and provides an admin dashboard for tracking per-user costs and issuing virtual keys.

by read3 min views1 publishedOct 3, 2026

As developers, we’ve all been there: you build an AI-powered feature, deploy it to production, and suddenly get slapped with a massive OpenAI or Anthropic API bill because a user ran a recursive loop, or because your team shared raw API keys across staging environments. 💸

Sharing master API keys is a major security risk, and tracking individual user costs across multiple models (GPT-4o, Claude 3.5 Sonnet, Llama-3) is a management nightmare.

There is a better way. By self-hosting LiteLLM, you can spin up a unified, private LLM Gateway that acts as a proxy between your applications and your AI providers.

With LiteLLM, you get:

In this guide, we'll deploy a production-ready LiteLLM Instance with a beautiful UI Admin Dashboard on a high-speed Vultr High-Performance Cloud Instance in under 10 minutes.

When proxying API requests, latency is everything. Adding a proxy layer shouldn't introduce lag. Vultr's High-Performance Cloud offers high-frequency CPU cores and NVMe storage in over 32 global locations. This means your proxy can be co-located right next to your target audience (or your main application servers) to ensure sub-millisecond routing overhead.

Once your instance is ready, copy the IP address and connect to it via your terminal:

ssh root@YOUR_VULTR_IP

Update your system packages and install Docker to manage our containers easily:

sudo apt update && sudo apt upgrade -y
sudo apt install -y docker.io docker-compose

Verify the installation:

docker --version && docker-compose --version

Create a dedicated directory for LiteLLM:

mkdir litellm-gateway && cd litellm-gateway

Now, create a litellm_config.yaml file. This is where you configure the models you want to proxy. For this demo, we'll configure OpenAI's GPT-4o and Anthropic's Claude 3.5 Sonnet:

model_list:
  - model_name: gpt-4o
    litellm_params:
      model: openai/gpt-4o
      api_key: "os.environ/OPENAI_API_KEY"
  - model_name: claude-3-5-sonnet
    litellm_params:
      model: anthropic/claude-3-5-sonnet-20240620
      api_key: "os.environ/ANTHROPIC_API_KEY"

router_settings:
  routing_strategy: latency-based-routing

general_settings:
  master_key: "os.environ/LITELLM_MASTER_KEY"

We will set up three components:

Create a docker-compose.yml file in the same directory:

version: '3.8'

services:
  db:
    image: postgres:16-alpine
    container_name: litellm-db
    environment:
      POSTGRES_DB: litellm
      POSTGRES_USER: litellm_user
      POSTGRES_PASSWORD: super_secret_db_password
    volumes:
      - pgdata:/var/lib/postgresql/data
    ports:
      - "5432:5432"
    restart: always

  litellm:
    image: ghcr.io/berriai/litellm:main-latest
    container_name: litellm-proxy
    ports:
      - "4000:4000"
    volumes:
      - ./litellm_config.yaml:/app/config.yaml
    environment:
      - DATABASE_URL=postgresql://litellm_user:super_secret_db_password@db:5432/litellm
      - LITELLM_MASTER_KEY=sk-your-super-secure-master-admin-key-12345
      - OPENAI_API_KEY=your_actual_openai_api_key_here
      - ANTHROPIC_API_KEY=your_actual_anthropic_api_key_here
    depends_on:
      - db
    command: ["--config", "/app/config.yaml", "--detailed_debug"]
    restart: always

volumes:
  pgdata:

💡 Note: Replace your_actual_openai_api_key_here and your_actual_anthropic_api_key_here with your real API keys, and change the LITELLM_MASTER_KEY to a secure, random string.

Run the Docker stack in detached mode:

docker-compose up -d

Check if everything is running correctly:

docker-compose ps

Your gateway is now active!

http://YOUR_VULTR_IP:4000/ui. LITELLM_MASTER_KEY you configured in your docker-compose.yml (e.g., sk-your-super-secure-master-admin-key-12345). From this dashboard, you can:

To consume your self-hosted gateway, you simply point your existing OpenAI SDK to your Vultr server. No code changes required!

from openai import OpenAI

client = OpenAI(
    api_key="sk-the-virtual-key-you-created-in-dashboard",
    base_url="http://YOUR_VULTR_IP:4000"
)

response = client.chat.completions.create(
    model="claude-3-5-sonnet",
    messages=[{"role": "user", "content": "Explain quantum computing in 2 sentences."}]
)

print(response.choices[0].message.content)

Before pointing production applications to your gateway, ensure you secure it with SSL. You can easily install Nginx and Certbot on your Vultr instance to get a free Let's Encrypt SSL certificate:

sudo apt install -y nginx certbot python3-certbot-nginx

Configure Nginx to reverse proxy traffic from port 80/ 443 to http://localhost:4000 and run sudo certbot --nginx to enable HTTPS.

Ready to get your team's AI costs under lock and key? Build your high-performance, private LLM gateway today on Vultr High-Performance Cloud!

Liked this resource? Join our daily Telegram channel for more developer tools and cloud insights: @Libretech2026

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @litellm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stop-bleeding-ai-cos…] indexed:0 read:3min 2026-10-03 · —