Shepherd Model Gateway – Model-routing gateway for LLM deployments LightSeek Org released Shepherd Model Gateway (SMG), an engine-agnostic, high-performance model-routing gateway for large-scale LLM deployments that centralizes worker lifecycle management and balances traffic across HTTP, gRPC, and OpenAI-compatible backends. The gateway supports cache-aware routing for vLLM, TensorRT-LLM, TokenSpeed, and SGLang, offers multi-tenant rate limiting with OIDC and WebAssembly plugins, and provides 40+ Prometheus metrics with OpenTelemetry tracing. SMG is available via Docker, pip, and cargo, and can be run with a single command to point at inference workers. Engine-agnostic, high-performance model-routing gateway for large-scale LLM deployments. Centralizes worker lifecycle management, balances traffic across HTTP/gRPC/OpenAI-compatible backends, and provides enterprise-ready control over history storage, MCP tooling, and privacy-sensitive workflows. 🚀 Maximize GPU Utilization | Cache-aware routing understands your inference engine's KV cache state—whether vLLM, TensorRT-LLM, TokenSpeed, or SGLang—to reuse prefixes and reduce redundant computation. | 🔌 One API, Any Backend | Route to self-hosted models vLLM, TensorRT-LLM, TokenSpeed, SGLang or cloud providers OpenAI, Anthropic, Gemini, Bedrock, and more through a single unified endpoint. | ⚡ Built for Speed | Native Rust with gRPC pipelines, sub-millisecond routing decisions, and zero-copy tokenization. Circuit breakers and automatic failover keep things running. | 🔒 Enterprise Control | Multi-tenant rate limiting with OIDC, WebAssembly plugins for custom logic, and a privacy boundary that keeps conversation history within your infrastructure. | 📊 Full Observability | 40+ Prometheus metrics, OpenTelemetry tracing, and structured JSON logs with request correlation—know exactly what's happening at every layer. | API Coverage: OpenAI Chat/Completions/Embeddings, Responses API for agents, Anthropic Messages, and MCP tool execution. Install — pick your preferred method: Docker docker pull lightseekorg/smg:latest Python pip install smg Rust cargo install smg Run — point SMG at your inference workers: Single worker smg launch --worker-urls http://localhost:8000 Multiple workers with cache-aware routing smg launch --worker-urls http://gpu1:8000 http://gpu2:8000 --policy cache aware With high availability mesh smg launch --worker-urls http://gpu1:8000 --enable-mesh \ --mesh-advertise-host 10.0.0.1 --mesh-peer-urls 10.0.0.2:39527 Use — send requests to the gateway: curl http://localhost:30000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model": "llama3", "messages": {"role": "user", "content": "Hello "} }' That's it. SMG is now load-balancing requests across your workers. | Self-Hosted | Cloud Providers | |---|---| | vLLM | OpenAI | | TensorRT-LLM | Anthropic | | TokenSpeed | Google Gemini | | SGLang | AWS Bedrock | | Ollama | Azure OpenAI | | Any OpenAI-compatible server | Any OpenAI-compatible provider | | Feature | Description | |---|---| | cache aware, round robin, power of two, consistent hashing, prefix hash, manual, random, bucket | | Native gRPC with streaming, reasoning extraction, and tool call parsing | | Connect external tool servers via Model Context Protocol | | Mesh networking with SWIM protocol for multi-node deployments | | Pluggable storage: PostgreSQL, Oracle, Redis, or in-memory | | Extend with custom WebAssembly logic | | Circuit breakers, retries with backoff, rate limiting | | Architecture /lightseekorg/smg/blob/main/docs/concepts/architecture/overview.md Configuration /lightseekorg/smg/blob/main/docs/reference/configuration.md API Reference /lightseekorg/smg/blob/main/docs/reference/api/openai.md Kubernetes Setup /lightseekorg/smg/blob/main/docs/getting-started/service-discovery.md We welcome contributions See Contributing Guide /lightseekorg/smg/blob/main/docs/contributing/index.md for details.