cd /news/ai-infrastructure/litellm-rust-gateway-benchmarked-fas… · home › topics › ai-infrastructure › article
[ARTICLE · art-143552] src=dev.to ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

LiteLLM Rust Gateway Benchmarked: Fast and Tiny, but Not Yet a Python Proxy Replacement

A developer benchmarked the LiteLLM Rust gateway beta against the LiteLLM Python v1 proxy, Portkey OSS, and Bifrost 1.6.4, finding the Rust path added roughly 0.7 ms at p99, used 21.8 MB peak memory, and sustained about 2,814 requests per second against a deterministic local mock, versus 257.7 ms p99 and 329.5 MB for the Python proxy. The test disabled logging callbacks, persistence, and spend tracking and forwarded an Anthropic Messages body to a local Rust mock, so the author treats it as a runtime-isolation benchmark rather than a production architecture comparison. The conclusion is that the Rust gateway is fast and tiny but does not yet cover enough of a real gateway workload to replace a feature-complete Python deployment.

by read11 min views1 publishedOct 2, 2026

We did not bring LiteLLM into the lab because another millisecond matters on a 12-second reasoning request. We brought it in because gateway overhead becomes operationally expensive when traffic consists of embeddings, classifiers, guardrail calls, short agent turns, and other fast requests issued at high concurrency.

p99 added latency by gatewayLiteLLM Rust beta 0.7 msPortkey OSS 2.3 msBifrost 1.6.4 4.5 msLiteLLM Python v1 257.7 ms

Rust had the lowest p99 added latency in the local-mock test, but the test disabled logging, persistence, and spend tracking and is not a full-feature comparison.

The existing LiteLLM Python proxy offers a broad control plane: provider normalization, authentication, routing, budgets, callbacks, spend tracking, persistence, and an OpenAI-compatible API. The Rust work targets the forwarding hot path underneath that surface.

There are currently two materially different deployment modes:

That distinction matters. The impressive benchmark numbers belong to a narrow forwarding path, not automatically to a feature-complete replacement for an existing Python deployment.

In the July 22, 2026 benchmark artifacts we reviewed, the standalone Rust path added approximately 0.7 ms at p99, consumed 21.8 MB peak memory, and sustained roughly 2,814 requests per second against a deterministic local mock. The same test recorded the LiteLLM Python v1 proxy at 257.7 ms p99 added latency and 329.5 MB peak memory.

Those are large differences, but the test intentionally disabled logging callbacks, persistence, and spend tracking. It forwarded an Anthropic Messages body to a local Rust mock. We therefore treated it as a runtime-isolation benchmark, not a production architecture benchmark.

The earlier migration harness showed a different traffic shape and a different generation of the implementation: approximately 0.05 ms Rust overhead versus 7.5 ms for Python, 6,782 versus 453 requests per second, and 31.7 MB versus 358.9 MB peak memory. We did not merge those values into one result because the workloads and harnesses were not identical.

The practical question for us was not whether Rust can forward JSON faster than Python. It can. The real question was whether the current LiteLLM Rust surface covers enough of an actual gateway workload to justify migration.

Our executed local experiment tested the benchmark's response-presence guard using Requests response fixtures. It made no network calls and measured no latency.

For a follow-up gateway-overhead test, we would use three layers:

Using a mock would keep provider queueing, internet variance, rate limits, and model generation time out of the measurement. The following configuration illustrates how we would route /chat/completions traffic to that mock:

model_list:
  - model_name: fake-openai-endpoint
    litellm_params:
      model: openai/any
      api_base: http://mock-openai:8080/v1
      api_key: test

litellm_settings:
  callbacks: []
  num_retries: 0
  request_timeout: 30

general_settings:
  master_key: sk-local-benchmark-key

For that follow-up test, we would use a minimal mock rather than a live provider:

from fastapi import FastAPI
from pydantic import BaseModel
import time

app = FastAPI()

class ChatRequest(BaseModel):
    model: str
    messages: list
    max_tokens: int | None = None

@app.post("/v1/chat/completions")
async def chat_completions(request: ChatRequest):
    return {
        "id": "chatcmpl-local-fixed",
        "object": "chat.completion",
        "created": int(time.time()),
        "model": request.model,
        "choices": [{
            "index": 0,
            "message": {
                "role": "assistant",
                "content": "deterministic mock response"
            },
            "finish_reason": "stop"
        }],
        "usage": {
            "prompt_tokens": 8,
            "completion_tokens": 3,
            "total_tokens": 11
        }
    }

This proposed Docker Compose setup would keep the mock and proxy on the same bridge network:

services:
  mock-openai:
    image: python:3.12-slim
    working_dir: /app
    volumes:
      - ./mock.py:/app/mock.py:ro
    command: >
      sh -c "pip install --no-cache-dir fastapi==0.115.0
      uvicorn==0.30.6 pydantic==2.9.2 &&
      uvicorn mock:app --host 0.0.0.0 --port 8080"
    ports:
      - "8080:8080"

  litellm:
    image: ghcr.io/berriai/litellm:main-stable
    depends_on:
      - mock-openai
    volumes:
      - ./config.yaml:/app/config.yaml:ro
    command:
      - "--config"
      - "/app/config.yaml"
      - "--port"
      - "4000"
      - "--num_workers"
      - "4"
    ports:
      - "4000:4000"

To run the follow-up routing smoke test, we would start the stack and send a request through the proxy:

docker compose up -d

curl -i http://localhost:4000/chat/completions \
  -H 'Authorization: Bearer sk-local-benchmark-key' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "fake-openai-endpoint",
    "messages": [{"role": "user", "content": "ping"}],
    "max_tokens": 16
  }'

The following is an illustrative response shape, not captured smoke-test output. The 2.41 ms overhead value is an example, not a measured result; our executed local fixture test made no network calls and measured no durations:

HTTP/1.1 200 OK
content-type: application/json
x-litellm-overhead-duration-ms: 2.41

{
  "id": "chatcmpl-local-fixed",
  "object": "chat.completion",
  "model": "any",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "deterministic mock response"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 8,
    "completion_tokens": 3,
    "total_tokens": 11
  }
}

A real response from this setup would provide a routing and instrumentation check, not a latency benchmark. Our executed fixture test did not validate proxy routing.

For a follow-up load test, we would use a constant arrival rate rather than an unconstrained closed loop. This distinction is easy to miss. One thousand clients immediately submitting another request after every response can produce approximately the same throughput as a controlled test while holding far more work in flight, making p95 latency look dramatically worse.

In our Python proxy benchmark analysis, we examined a setup with 1,000 Locust users, 0.5 to 1 second of think time, roughly 1,170 RPS, and approximately 130 requests in flight. Under the four-instance configuration, the recorded total /chat/completions p95 was 150 ms, while the gateway's own x-litellm-overhead-duration-ms p95 was 8 ms. The often-repeated “8 ms p95 at 1k RPS” figure refers to measured proxy overhead, not complete client-visible request latency.

We also checked the benchmark listener itself. The sample listener uses:

if response and hasattr(response, "headers"):
    overhead = response.headers.get("x-litellm-overhead-duration-ms")

With requests==2.32.3, we constructed 200, 429, and 500 response fixtures carrying the same header. The tested response-presence guard reached the header on the 200 fixture but skipped the 429 and 500 fixtures because those requests.Response objects were falsey.

The safer guard is:

if response is not None and hasattr(response, "headers"):
    overhead = response.headers.get("x-litellm-overhead-duration-ms")

That does not prove LiteLLM always attaches the overhead header to error responses. Our fixture test shows that this guard can exclude header-bearing 429 and 500 responses. If real error responses carry a parseable overhead header, that exclusion could bias the collected distribution toward successful responses; we did not measure a distribution or test LiteLLM's error responses.

Our first roadblock was straightforward: there was no published prebuilt Docker image for the standalone Axum gateway. The available deployment path required building the server from the litellm-rust workspace with its server feature or obtaining early-beta access.

That means a normal docker compose pull workflow can reproduce the Python proxy test, but not the complete standalone Rust benchmark. We would not put an internally compiled beta binary into a production base image without pinning a commit, generating an SBOM, scanning dependencies, and owning the release pipeline.

The lower-risk hybrid mode was easier to reason about. We enabled Rust per model with rust: true and checked for the x-litellm-rust: true response header. If the header was absent, the request had stayed on or fallen back to Python.

The parity boundary was narrower than the phrase “Rust /chat/completions support” suggests. The Rust route accepted non-streaming text conversations for the supported Anthropic and Bedrock paths. It rejected or routed around requests containing:

stream: true`` response_format or JSON moden greater than onetop_k The route decision happens before calling the provider. Unsupported requests can therefore fall back to Python without issuing two billable calls. Once Rust has called the provider, however, it returns any downstream failure directly rather than retrying through Python. That behavior avoids duplicate provider charges but means “automatic fallback” is not a universal retry mechanism.

We also encountered a version-boundary problem. The per-model Rust flag for Anthropic /v1/messages arrived in v1.94.0, initially in v1.94.0-rc.1. We could not identify the exact release boundary for Rust /chat/completions support on Anthropic and Bedrock, so we treated that coverage as version-dependent; the standalone server remained an early beta. We would pin an exact version and run contract tests for every parameter shape used by production clients rather than trusting a broad endpoint label.

The benchmark isolation is both a strength and a weakness. A local deterministic mock removes provider variance, which is exactly what we want when measuring gateway overhead. It also excludes the network, rate limits, provider-side queueing, and multi-second generation time that dominate real chat requests.

The Rust comparison used single-host, per-scenario runs without repeated-trial error bars. We consider the order-of-magnitude memory difference credible enough to investigate, but not sufficient for capacity planning.

Finally, the auto-routing result exposed a different class of risk: cost optimization can silently become quality degradation. On the three-model Gemini 3.x ladder, the semantic router at a 0.2 threshold saved 58.2% and achieved a 46.8% win rate against the flagship, with 82 of 90 exact-match answers. A fitted rule-based complexity router saved 58.3% but dropped to a 41.5% win rate and 71 of 90 exact matches.

The apparent middle tier was particularly poor. It cost 14.1 times as much as the cheap model while producing worse measured results. Its 981,308 thinking tokens also reduced a nominal fourfold output-price advantage to only a 2.7-fold measured advantage.

We could not verify the stated 51% production cost savings from the available benchmark evidence. We found controlled savings of 47.5%, 58.2%, and 58.3% on one model ladder and one evaluation set, but no production dataset establishing 51%. We would not place that number in a business case without billing exports, production-quality labels, and the exact routing policy that generated it.

The cleanest cross-gateway comparison came from the identical local-mock forwarding test:

Gateway p99 added latency Peak memory Estimated gateway cost per 1M requests Sustained throughput
LiteLLM Rust beta 0.7 ms 21.8 MB $0.000175 About 2,814 RPS
Portkey OSS 2.3 ms 90.4 MB $0.001042 Not reported in the same summary
Bifrost 1.6.4 4.5 ms 199.1 MB $0.001008 About 2,744 RPS
LiteLLM Python v1 257.7 ms 329.5 MB $0.015354 Not reported in the same summary

We distinguish measured p99 added latency, peak memory, and sustained throughput from the estimated gateway-cost column. The cost estimates use measured CPU, peak RSS, and sustained throughput on a 4 vCPU / 16 GB instance and exclude token charges. The forwarding tests used a deterministic local upstream with logging callbacks, spend tracking, and persistence disabled.

That last sentence matters more than the ranking. Rust used roughly one-fifteenth of Python's measured memory in this specific run, but the estimated gateway cost difference was only about $0.01518 per million requests. Even at one billion requests per month, the direct estimate implies approximately $15.18 in gateway-footprint savings. That number is too small to fund a migration by itself.

The economic case comes from secondary effects:

For ordinary chat traffic, token spend dwarfs gateway infrastructure. Auto-routing can therefore be much more valuable than shaving gateway CPU, but only if quality survives.

Using the controlled routing benchmark as an illustration, flagship-only responses cost $9.320 for the evaluation set. The semantic router reduced that to $4.893 at a 0.3 threshold or $3.894 at a 0.2 threshold. The latter preserved 82 of 90 exact-match answers, but its 95% confidence interval for win rate was 44.0% to 49.5%. It did not conclusively satisfy both a 45% quality floor and a 50% savings target.

For a hypothetical $100,000 monthly model bill, a genuine 51% reduction would save $51,000 per month. If migration and evaluation cost $10,000 upfront, a one-month payback would require roughly $19,608 in monthly baseline spend, assuming the 51% reduction and excluding ongoing costs:

Break-even monthly spend = migration cost / savings rate
                          = $10,000 / 0.51
                          = $19,607.84

That arithmetic is useful, but the 51% input is not verified production evidence. Our actual approval gate would use:

Net savings =
  model savings
  - router inference cost
  - evaluation cost
  - quality-related rework
  - additional observability
  - migration and maintenance cost

Teams evaluating adjacent gateways can compare our other infrastructure reviews in the Effloow tools collection. For workload-specific capacity and routing analysis, our AI infrastructure services cover the parts a synthetic benchmark cannot settle.

Deploy or trial LiteLLM's Rust path if:

x-litellm-rust: true. Hold off or avoid the standalone gateway if:

Treat auto-routing as a separate production project if:

Our final assessment is positive but narrow. The Rust implementation demonstrates a real systems advantage: dramatically lower forwarding overhead and memory in controlled tests. It is especially compelling as an opt-in acceleration path inside an existing LiteLLM deployment.

It is not yet a drop-in replacement for the full Python proxy. The missing standalone image, incomplete endpoint surface, automatic Python fallback, and benchmark configuration without production features prevent us from recommending a wholesale migration.

We would deploy hybrid mode behind telemetry, measure the Rust hit rate by request shape, and keep Python as the compatibility path. We would move to standalone Rust only after our production contract suite reached full parity and container distribution met our release requirements.

The benchmark methodology is available in LiteLLM's Rust gateway benchmark, the Python proxy baseline is described in its gateway benchmark guide, and the routing tradeoffs are detailed in the cost-ladder benchmark. If you want a second set of eyes on your own numbers before committing engineering time, contact our team.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @litellm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/litellm-rust-gateway…] indexed:0 read:11min 2026-10-02 · —