Hi everyone!
Managing multiple LLM providers, rate limits, and API downtime while keeping latency and costs low became a hassle, so I built llmproxy.
What it does:
⚡ Response Caching: Reduces duplicate requests and cuts API costs.
🔄 Automatic Failover: Seamlessly fallbacks to backup providers or local models if your primary API is rate-limited or down.
🏡 Local + Cloud Routing: Route light prompts to local models (Ollama/vLLM) and heavy ones to cloud APIs (OpenAI/Anthropic).
🔌 OpenAI Compatible: Drop-in replacement—just point your base_url to llmproxy.
It’s completely open-source and easy to spin up with Docker.
📂 GitHub: https://github.com/lordraw77/llmproxy I’d love to get your feedback, feature requests, or suggestions on how to improve it!