# Stop Bleeding AI Costs: Self-Host Your Own Private LLM Gateway (LiteLLM) in Under 10 Minutes

> Source: <https://dev.to/fejuno/stop-bleeding-ai-costs-self-host-your-own-private-llm-gateway-litellm-in-under-10-minutes-4h04>
> Published: 2026-10-03 13:34:17+00:00

As developers, we’ve all been there: you build an AI-powered feature, deploy it to production, and suddenly get slapped with a massive OpenAI or Anthropic API bill because a user ran a recursive loop, or because your team shared raw API keys across staging environments. 💸

Sharing master API keys is a major security risk, and tracking individual user costs across multiple models (GPT-4o, Claude 3.5 Sonnet, Llama-3) is a management nightmare.

There is a better way. By self-hosting **LiteLLM**, you can spin up a unified, private **LLM Gateway** that acts as a proxy between your applications and your AI providers. 

With LiteLLM, you get:

In this guide, we'll deploy a production-ready LiteLLM Instance with a beautiful UI Admin Dashboard on a high-speed [Vultr High-Performance Cloud Instance](https://www.vultr.com/?ref=9925661) in under 10 minutes.

When proxying API requests, **latency is everything**. Adding a proxy layer shouldn't introduce lag. [Vultr's High-Performance Cloud](https://www.vultr.com/?ref=9925661) offers high-frequency CPU cores and NVMe storage in over 32 global locations. This means your proxy can be co-located right next to your target audience (or your main application servers) to ensure sub-millisecond routing overhead.

Once your instance is ready, copy the IP address and connect to it via your terminal:

```
ssh root@YOUR_VULTR_IP
```

Update your system packages and install Docker to manage our containers easily:

```
sudo apt update && sudo apt upgrade -y
sudo apt install -y docker.io docker-compose
```

Verify the installation:

```
docker --version && docker-compose --version
```

Create a dedicated directory for LiteLLM:

```
mkdir litellm-gateway && cd litellm-gateway
```

Now, create a `litellm_config.yaml` file. This is where you configure the models you want to proxy. For this demo, we'll configure OpenAI's GPT-4o and Anthropic's Claude 3.5 Sonnet:

```
model_list:
  - model_name: gpt-4o
    litellm_params:
      model: openai/gpt-4o
      api_key: "os.environ/OPENAI_API_KEY"
  - model_name: claude-3-5-sonnet
    litellm_params:
      model: anthropic/claude-3-5-sonnet-20240620
      api_key: "os.environ/ANTHROPIC_API_KEY"

# Configure global database for key management and logs
router_settings:
  routing_strategy: latency-based-routing

general_settings:
  master_key: "os.environ/LITELLM_MASTER_KEY"
```

We will set up three components:

Create a `docker-compose.yml` file in the same directory:

```
version: '3.8'

services:
  db:
    image: postgres:16-alpine
    container_name: litellm-db
    environment:
      POSTGRES_DB: litellm
      POSTGRES_USER: litellm_user
      POSTGRES_PASSWORD: super_secret_db_password
    volumes:
      - pgdata:/var/lib/postgresql/data
    ports:
      - "5432:5432"
    restart: always

  litellm:
    image: ghcr.io/berriai/litellm:main-latest
    container_name: litellm-proxy
    ports:
      - "4000:4000"
    volumes:
      - ./litellm_config.yaml:/app/config.yaml
    environment:
      - DATABASE_URL=postgresql://litellm_user:super_secret_db_password@db:5432/litellm
      - LITELLM_MASTER_KEY=sk-your-super-secure-master-admin-key-12345
      - OPENAI_API_KEY=your_actual_openai_api_key_here
      - ANTHROPIC_API_KEY=your_actual_anthropic_api_key_here
    depends_on:
      - db
    command: ["--config", "/app/config.yaml", "--detailed_debug"]
    restart: always

volumes:
  pgdata:
```

💡 *Note: Replace `your_actual_openai_api_key_here` and `your_actual_anthropic_api_key_here` with your real API keys, and change the `LITELLM_MASTER_KEY` to a secure, random string.*

Run the Docker stack in detached mode:

```
docker-compose up -d
```

Check if everything is running correctly:

```
docker-compose ps
```

Your gateway is now active!

`http://YOUR_VULTR_IP:4000/ui`.` LITELLM_MASTER_KEY` you configured in your `docker-compose.yml` (e.g., `sk-your-super-secure-master-admin-key-12345`).
From this dashboard, you can:

To consume your self-hosted gateway, you simply point your existing OpenAI SDK to your Vultr server. No code changes required!

``` python
from openai import OpenAI

client = OpenAI(
    api_key="sk-the-virtual-key-you-created-in-dashboard",
    base_url="http://YOUR_VULTR_IP:4000"
)

# Call any model you defined in your LiteLLM config!
response = client.chat.completions.create(
    model="claude-3-5-sonnet",
    messages=[{"role": "user", "content": "Explain quantum computing in 2 sentences."}]
)

print(response.choices[0].message.content)
```

Before pointing production applications to your gateway, ensure you secure it with SSL. You can easily install **Nginx** and **Certbot** on your Vultr instance to get a free Let's Encrypt SSL certificate:

```
sudo apt install -y nginx certbot python3-certbot-nginx
```

Configure Nginx to reverse proxy traffic from port `80`/` 443` to `http://localhost:4000` and run `sudo certbot --nginx` to enable HTTPS.

Ready to get your team's AI costs under lock and key? Build your high-performance, private LLM gateway today on [Vultr High-Performance Cloud](https://www.vultr.com/?ref=9925661)!

*Liked this resource? Join our daily Telegram channel for more developer tools and cloud insights: @Libretech2026*
