{"slug": "litellm-rust-gateway-benchmarked-fast-and-tiny-but-not-yet-a-python-proxy", "title": "LiteLLM Rust Gateway Benchmarked: Fast and Tiny, but Not Yet a Python Proxy Replacement", "summary": "A developer benchmarked the LiteLLM Rust gateway beta against the LiteLLM Python v1 proxy, Portkey OSS, and Bifrost 1.6.4, finding the Rust path added roughly 0.7 ms at p99, used 21.8 MB peak memory, and sustained about 2,814 requests per second against a deterministic local mock, versus 257.7 ms p99 and 329.5 MB for the Python proxy. The test disabled logging callbacks, persistence, and spend tracking and forwarded an Anthropic Messages body to a local Rust mock, so the author treats it as a runtime-isolation benchmark rather than a production architecture comparison. The conclusion is that the Rust gateway is fast and tiny but does not yet cover enough of a real gateway workload to replace a feature-complete Python deployment.", "body_md": "We did not bring LiteLLM into the lab because another millisecond matters on a 12-second reasoning request. We brought it in because gateway overhead becomes operationally expensive when traffic consists of embeddings, classifiers, guardrail calls, short agent turns, and other fast requests issued at high concurrency.\n\np99 added latency by gatewayLiteLLM Rust beta 0.7 msPortkey OSS 2.3 msBifrost 1.6.4 4.5 msLiteLLM Python v1 257.7 ms\n\nRust had the lowest p99 added latency in the local-mock test, but the test disabled logging, persistence, and spend tracking and is not a full-feature comparison.\n\nThe existing LiteLLM Python proxy offers a broad control plane: provider normalization, authentication, routing, budgets, callbacks, spend tracking, persistence, and an OpenAI-compatible API. The Rust work targets the forwarding hot path underneath that surface.\n\nThere are currently two materially different deployment modes:\n\nThat distinction matters. The impressive benchmark numbers belong to a narrow forwarding path, not automatically to a feature-complete replacement for an existing Python deployment.\n\nIn the July 22, 2026 benchmark artifacts we reviewed, the standalone Rust path added approximately **0.7 ms at p99**, consumed **21.8 MB peak memory**, and sustained roughly **2,814 requests per second** against a deterministic local mock. The same test recorded the LiteLLM Python v1 proxy at **257.7 ms p99 added latency** and **329.5 MB peak memory**.\n\nThose are large differences, but the test intentionally disabled logging callbacks, persistence, and spend tracking. It forwarded an Anthropic Messages body to a local Rust mock. We therefore treated it as a runtime-isolation benchmark, not a production architecture benchmark.\n\nThe earlier migration harness showed a different traffic shape and a different generation of the implementation: approximately **0.05 ms Rust overhead versus 7.5 ms for Python**, **6,782 versus 453 requests per second**, and **31.7 MB versus 358.9 MB peak memory**. We did not merge those values into one result because the workloads and harnesses were not identical.\n\nThe practical question for us was not whether Rust can forward JSON faster than Python. It can. The real question was whether the current LiteLLM Rust surface covers enough of an actual gateway workload to justify migration.\n\nOur executed local experiment tested the benchmark's response-presence guard using Requests response fixtures. It made no network calls and measured no latency.\n\nFor a follow-up gateway-overhead test, we would use three layers:\n\nUsing a mock would keep provider queueing, internet variance, rate limits, and model generation time out of the measurement. The following configuration illustrates how we would route `/chat/completions` traffic to that mock:\n\n```\n# config.yaml\nmodel_list:\n  - model_name: fake-openai-endpoint\n    litellm_params:\n      model: openai/any\n      api_base: http://mock-openai:8080/v1\n      api_key: test\n\nlitellm_settings:\n  callbacks: []\n  num_retries: 0\n  request_timeout: 30\n\ngeneral_settings:\n  master_key: sk-local-benchmark-key\n```\n\nFor that follow-up test, we would use a minimal mock rather than a live provider:\n\n``` python\n# mock.py\nfrom fastapi import FastAPI\nfrom pydantic import BaseModel\nimport time\n\napp = FastAPI()\n\nclass ChatRequest(BaseModel):\n    model: str\n    messages: list\n    max_tokens: int | None = None\n\n@app.post(\"/v1/chat/completions\")\nasync def chat_completions(request: ChatRequest):\n    return {\n        \"id\": \"chatcmpl-local-fixed\",\n        \"object\": \"chat.completion\",\n        \"created\": int(time.time()),\n        \"model\": request.model,\n        \"choices\": [{\n            \"index\": 0,\n            \"message\": {\n                \"role\": \"assistant\",\n                \"content\": \"deterministic mock response\"\n            },\n            \"finish_reason\": \"stop\"\n        }],\n        \"usage\": {\n            \"prompt_tokens\": 8,\n            \"completion_tokens\": 3,\n            \"total_tokens\": 11\n        }\n    }\n```\n\nThis proposed Docker Compose setup would keep the mock and proxy on the same bridge network:\n\n```\n# compose.yaml\nservices:\n  mock-openai:\n    image: python:3.12-slim\n    working_dir: /app\n    volumes:\n      - ./mock.py:/app/mock.py:ro\n    command: >\n      sh -c \"pip install --no-cache-dir fastapi==0.115.0\n      uvicorn==0.30.6 pydantic==2.9.2 &&\n      uvicorn mock:app --host 0.0.0.0 --port 8080\"\n    ports:\n      - \"8080:8080\"\n\n  litellm:\n    image: ghcr.io/berriai/litellm:main-stable\n    depends_on:\n      - mock-openai\n    volumes:\n      - ./config.yaml:/app/config.yaml:ro\n    command:\n      - \"--config\"\n      - \"/app/config.yaml\"\n      - \"--port\"\n      - \"4000\"\n      - \"--num_workers\"\n      - \"4\"\n    ports:\n      - \"4000:4000\"\n```\n\nTo run the follow-up routing smoke test, we would start the stack and send a request through the proxy:\n\n```\ndocker compose up -d\n\ncurl -i http://localhost:4000/chat/completions \\\n  -H 'Authorization: Bearer sk-local-benchmark-key' \\\n  -H 'Content-Type: application/json' \\\n  -d '{\n    \"model\": \"fake-openai-endpoint\",\n    \"messages\": [{\"role\": \"user\", \"content\": \"ping\"}],\n    \"max_tokens\": 16\n  }'\n```\n\nThe following is an illustrative response shape, not captured smoke-test output. The 2.41 ms overhead value is an example, not a measured result; our executed local fixture test made no network calls and measured no durations:\n\n```\nHTTP/1.1 200 OK\ncontent-type: application/json\nx-litellm-overhead-duration-ms: 2.41\n\n{\n  \"id\": \"chatcmpl-local-fixed\",\n  \"object\": \"chat.completion\",\n  \"model\": \"any\",\n  \"choices\": [\n    {\n      \"index\": 0,\n      \"message\": {\n        \"role\": \"assistant\",\n        \"content\": \"deterministic mock response\"\n      },\n      \"finish_reason\": \"stop\"\n    }\n  ],\n  \"usage\": {\n    \"prompt_tokens\": 8,\n    \"completion_tokens\": 3,\n    \"total_tokens\": 11\n  }\n}\n```\n\nA real response from this setup would provide a routing and instrumentation check, not a latency benchmark. Our executed fixture test did not validate proxy routing.\n\nFor a follow-up load test, we would use a constant arrival rate rather than an unconstrained closed loop. This distinction is easy to miss. One thousand clients immediately submitting another request after every response can produce approximately the same throughput as a controlled test while holding far more work in flight, making p95 latency look dramatically worse.\n\nIn our Python proxy benchmark analysis, we examined a setup with 1,000 Locust users, 0.5 to 1 second of think time, roughly 1,170 RPS, and approximately 130 requests in flight. Under the four-instance configuration, the recorded total `/chat/completions` p95 was **150 ms**, while the gateway's own `x-litellm-overhead-duration-ms` p95 was **8 ms**. The often-repeated “8 ms p95 at 1k RPS” figure refers to measured proxy overhead, not complete client-visible request latency.\n\nWe also checked the benchmark listener itself. The sample listener uses:\n\n```\nif response and hasattr(response, \"headers\"):\n    overhead = response.headers.get(\"x-litellm-overhead-duration-ms\")\n```\n\nWith `requests==2.32.3`, we constructed 200, 429, and 500 response fixtures carrying the same header. The tested response-presence guard reached the header on the 200 fixture but skipped the 429 and 500 fixtures because those `requests.Response` objects were falsey.\n\nThe safer guard is:\n\n```\nif response is not None and hasattr(response, \"headers\"):\n    overhead = response.headers.get(\"x-litellm-overhead-duration-ms\")\n```\n\nThat does not prove LiteLLM always attaches the overhead header to error responses. Our fixture test shows that this guard can exclude header-bearing 429 and 500 responses. If real error responses carry a parseable overhead header, that exclusion could bias the collected distribution toward successful responses; we did not measure a distribution or test LiteLLM's error responses.\n\nOur first roadblock was straightforward: there was no published prebuilt Docker image for the standalone Axum gateway. The available deployment path required building the server from the `litellm-rust` workspace with its server feature or obtaining early-beta access.\n\nThat means a normal `docker compose pull` workflow can reproduce the Python proxy test, but not the complete standalone Rust benchmark. We would not put an internally compiled beta binary into a production base image without pinning a commit, generating an SBOM, scanning dependencies, and owning the release pipeline.\n\nThe lower-risk hybrid mode was easier to reason about. We enabled Rust per model with `rust: true` and checked for the `x-litellm-rust: true` response header. If the header was absent, the request had stayed on or fallen back to Python.\n\nThe parity boundary was narrower than the phrase “Rust `/chat/completions` support” suggests. The Rust route accepted non-streaming text conversations for the supported Anthropic and Bedrock paths. It rejected or routed around requests containing:\n\n`stream: true`` response_format` or JSON mode`n` greater than one`top_k`\nThe route decision happens before calling the provider. Unsupported requests can therefore fall back to Python without issuing two billable calls. Once Rust has called the provider, however, it returns any downstream failure directly rather than retrying through Python. That behavior avoids duplicate provider charges but means “automatic fallback” is not a universal retry mechanism.\n\nWe also encountered a version-boundary problem. The per-model Rust flag for Anthropic `/v1/messages` arrived in v1.94.0, initially in v1.94.0-rc.1. We could not identify the exact release boundary for Rust `/chat/completions` support on Anthropic and Bedrock, so we treated that coverage as version-dependent; the standalone server remained an early beta. We would pin an exact version and run contract tests for every parameter shape used by production clients rather than trusting a broad endpoint label.\n\nThe benchmark isolation is both a strength and a weakness. A local deterministic mock removes provider variance, which is exactly what we want when measuring gateway overhead. It also excludes the network, rate limits, provider-side queueing, and multi-second generation time that dominate real chat requests.\n\nThe Rust comparison used single-host, per-scenario runs without repeated-trial error bars. We consider the order-of-magnitude memory difference credible enough to investigate, but not sufficient for capacity planning.\n\nFinally, the auto-routing result exposed a different class of risk: cost optimization can silently become quality degradation. On the three-model Gemini 3.x ladder, the semantic router at a 0.2 threshold saved **58.2%** and achieved a **46.8% win rate against the flagship**, with **82 of 90 exact-match answers**. A fitted rule-based complexity router saved **58.3%** but dropped to a **41.5% win rate** and **71 of 90 exact matches**.\n\nThe apparent middle tier was particularly poor. It cost **14.1 times** as much as the cheap model while producing worse measured results. Its 981,308 thinking tokens also reduced a nominal fourfold output-price advantage to only a 2.7-fold measured advantage.\n\nWe could not verify the stated **51% production cost savings** from the available benchmark evidence. We found controlled savings of 47.5%, 58.2%, and 58.3% on one model ladder and one evaluation set, but no production dataset establishing 51%. We would not place that number in a business case without billing exports, production-quality labels, and the exact routing policy that generated it.\n\nThe cleanest cross-gateway comparison came from the identical local-mock forwarding test:\n\n| Gateway | p99 added latency | Peak memory | Estimated gateway cost per 1M requests | Sustained throughput | \n|---|---|---|---|---|\n| LiteLLM Rust beta | 0.7 ms | 21.8 MB | $0.000175 | About 2,814 RPS | \n| Portkey OSS | 2.3 ms | 90.4 MB | $0.001042 | Not reported in the same summary | \n| Bifrost 1.6.4 | 4.5 ms | 199.1 MB | $0.001008 | About 2,744 RPS | \n| LiteLLM Python v1 | 257.7 ms | 329.5 MB | $0.015354 | Not reported in the same summary | \n\nWe distinguish measured p99 added latency, peak memory, and sustained throughput from the estimated gateway-cost column. The cost estimates use measured CPU, peak RSS, and sustained throughput on a 4 vCPU / 16 GB instance and exclude token charges. The forwarding tests used a deterministic local upstream with logging callbacks, spend tracking, and persistence disabled.\n\nThat last sentence matters more than the ranking. Rust used roughly one-fifteenth of Python's measured memory in this specific run, but the estimated gateway cost difference was only about **$0.01518 per million requests**. Even at one billion requests per month, the direct estimate implies approximately **$15.18** in gateway-footprint savings. That number is too small to fund a migration by itself.\n\nThe economic case comes from secondary effects:\n\nFor ordinary chat traffic, token spend dwarfs gateway infrastructure. Auto-routing can therefore be much more valuable than shaving gateway CPU, but only if quality survives.\n\nUsing the controlled routing benchmark as an illustration, flagship-only responses cost $9.320 for the evaluation set. The semantic router reduced that to $4.893 at a 0.3 threshold or $3.894 at a 0.2 threshold. The latter preserved 82 of 90 exact-match answers, but its 95% confidence interval for win rate was 44.0% to 49.5%. It did not conclusively satisfy both a 45% quality floor and a 50% savings target.\n\nFor a hypothetical $100,000 monthly model bill, a genuine 51% reduction would save $51,000 per month. If migration and evaluation cost $10,000 upfront, a one-month payback would require roughly $19,608 in monthly baseline spend, assuming the 51% reduction and excluding ongoing costs:\n\n```\nBreak-even monthly spend = migration cost / savings rate\n                          = $10,000 / 0.51\n                          = $19,607.84\n```\n\nThat arithmetic is useful, but the 51% input is not verified production evidence. Our actual approval gate would use:\n\n```\nNet savings =\n  model savings\n  - router inference cost\n  - evaluation cost\n  - quality-related rework\n  - additional observability\n  - migration and maintenance cost\n```\n\nTeams evaluating adjacent gateways can compare our other infrastructure reviews in the [Effloow tools collection](https://dev.to/tools). For workload-specific capacity and routing analysis, our [AI infrastructure services](https://dev.to/services) cover the parts a synthetic benchmark cannot settle.\n\n**Deploy or trial LiteLLM's Rust path if:**\n\n`x-litellm-rust: true`.\n**Hold off or avoid the standalone gateway if:**\n\n**Treat auto-routing as a separate production project if:**\n\nOur final assessment is positive but narrow. The Rust implementation demonstrates a real systems advantage: dramatically lower forwarding overhead and memory in controlled tests. It is especially compelling as an opt-in acceleration path inside an existing LiteLLM deployment.\n\nIt is not yet a drop-in replacement for the full Python proxy. The missing standalone image, incomplete endpoint surface, automatic Python fallback, and benchmark configuration without production features prevent us from recommending a wholesale migration.\n\nWe would deploy hybrid mode behind telemetry, measure the Rust hit rate by request shape, and keep Python as the compatibility path. We would move to standalone Rust only after our production contract suite reached full parity and container distribution met our release requirements.\n\nThe benchmark methodology is available in LiteLLM's [Rust gateway benchmark](https://docs.litellm.ai/blog/rust-ai-gateway-benchmarks), the Python proxy baseline is described in its [gateway benchmark guide](https://docs.litellm.ai/docs/benchmarks), and the routing tradeoffs are detailed in the [cost-ladder benchmark](https://docs.litellm.ai/docs/proxy/auto_routing_benchmark). If you want a second set of eyes on your own numbers before committing engineering time, [contact our team](https://dev.to/contact).", "url": "https://wpnews.pro/news/litellm-rust-gateway-benchmarked-fast-and-tiny-but-not-yet-a-python-proxy", "canonical_source": "https://dev.to/jangwook_kim_e31e7291ad98/litellm-rust-gateway-benchmarked-fast-and-tiny-but-not-yet-a-python-proxy-replacement-4hhl", "published_at": "2026-10-02 00:40:09+00:00", "updated_at": "2026-10-02 00:44:32.833097+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-tools", "mlops"], "entities": ["LiteLLM", "Portkey", "Bifrost", "Anthropic", "OpenAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/litellm-rust-gateway-benchmarked-fast-and-tiny-but-not-yet-a-python-proxy", "markdown": "https://wpnews.pro/news/litellm-rust-gateway-benchmarked-fast-and-tiny-but-not-yet-a-python-proxy.md", "text": "https://wpnews.pro/news/litellm-rust-gateway-benchmarked-fast-and-tiny-but-not-yet-a-python-proxy.txt", "jsonld": "https://wpnews.pro/news/litellm-rust-gateway-benchmarked-fast-and-tiny-but-not-yet-a-python-proxy.jsonld"}}