{"slug": "routing-is-not-failover-a-practical-multi-provider-llm-pattern", "title": "Routing Is Not Failover: A Practical Multi-Provider LLM Pattern", "summary": "DVARA Open Source 1.8.0 introduces explicit multi-provider LLM routing that separates provider selection, retry and failover into distinct stages, with governance and capability checks applied before any upstream is chosen. The release supports static weighted, model-prefix, round-robin and canary routing configured in gateway.yaml, while latency-aware, cost-aware and geo-aware strategies remain Enterprise-only. The project warns that a text request succeeding against a fallback does not prove structured output, tool calling, streaming or vision will work, so each request shape must be tested and failure made explicit when no compatible provider remains.", "body_md": "Adding a second LLM provider does not automatically make an application resilient.\n\nThe common mistake is treating provider routing, retry and failover as one decision. They answer different questions at different points in a request. When those boundaries are unclear, teams can send traffic to an incompatible provider, bypass controls or struggle to explain why a request went where it did.\n\nA safer multi-provider design keeps the stages explicit:\n\n```\nApplication\n  → governance checks\n  → capability filter\n  → routing strategy\n  → selected provider\n\nAfter an eligible failure:\n  selected provider\n  → retry same provider\n  → compatible failover\n```\n\nProvider selection should not happen before policy enforcement. The request should first be associated with its workspace and credentials, then pass through the required policy, PII, guardrail and rate-limit checks.\n\nCapability filtering comes next. A provider may be healthy but still unable to satisfy a request involving structured output, tool calling, vision or another required feature. Filtering before routing ensures that the strategy operates only on eligible providers.\n\nOnly then should the gateway select an upstream.\n\nStatic weighted routing is useful when two compatible providers should receive a controlled traffic split. In [DVARA Open Source 1.8.0](https://github.com/dvarahq/dvara), a route can be expressed in `gateway.yaml`:\n\n```\nproviders:\n  - type: openai\n    api_key: ${OPENAI_API_KEY}\n  - type: azure-openai\n    api_key: ${AZURE_OPENAI_API_KEY}\n    base_url: ${AZURE_OPENAI_BASE_URL}\n\nroutes:\n  - id: weighted-gpt\n    model: \"gpt*\"\n    strategy: weighted\n    providers:\n      - provider: openai\n        weight: 80\n      - provider: azure-openai\n        weight: 20\n```\n\nBoth providers must accept the model identifier sent by the application. Routing selects a provider; it does not translate one provider's model name into another provider's deployment name.\n\nOpen Source reads this file at startup, so a routing change requires a Gateway restart. After restarting, verify the result over a meaningful sample of requests. A few calls are not enough to judge a probabilistic distribution. The response trace and structured access log's `provider` field show which upstream handled each request.\n\nDVARA Open Source 1.8.0 also supports model-prefix, round-robin and canary routing. Latency-aware, cost-aware, geo-aware and intelligent routing are Enterprise capabilities, not Open Source strategies.\n\nRouting asks: **Which eligible provider should receive this request first?**\n\nRetry asks: **Should the same provider receive another attempt after a retryable failure?**\n\nFailover asks: **Can a compatible alternative safely take over after eligible retries or provider unavailability?**\n\nThat last word—compatible—is important. A text request succeeding against a fallback does not prove that structured output, tool calling, streaming or vision will also work. Test each request shape and make failure explicit when no compatible provider remains.\n\nKeeping recovery separate also makes the evidence easier to interpret. A trace should show the original routing decision, the provider that received the request, and whether retry or fallback occurred afterward.\n\nThe practical goal is not merely to distribute traffic. It is to preserve the same governance and evidence path regardless of which provider is selected.\n\nStart with one model family and one operational goal: an even split, a controlled canary or a static weight. Confirm that policy and capability checks still apply, simulate an unavailable provider, and inspect the resulting trace before expanding the route.\n\nFor the complete configuration, verification steps and strategy boundaries, read the **[practical guide to multi-provider LLM routing](https://dvarahq.com/blog/multi-provider-llm-routing?utm_source=devto&utm_medium=referral&utm_campaign=seo_refresh_q4_2026&utm_content=multi_provider_routing_excerpt)**.", "url": "https://wpnews.pro/news/routing-is-not-failover-a-practical-multi-provider-llm-pattern", "canonical_source": "https://dev.to/dvarahq/routing-is-not-failover-a-practical-multi-provider-llm-pattern-2e1p", "published_at": "2026-09-28 13:08:30+00:00", "updated_at": "2026-09-28 13:20:06.655972+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-agents", "developer-tools", "mlops"], "entities": ["DVARA Open Source", "DVARA", "OpenAI", "Azure OpenAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/routing-is-not-failover-a-practical-multi-provider-llm-pattern", "markdown": "https://wpnews.pro/news/routing-is-not-failover-a-practical-multi-provider-llm-pattern.md", "text": "https://wpnews.pro/news/routing-is-not-failover-a-practical-multi-provider-llm-pattern.txt", "jsonld": "https://wpnews.pro/news/routing-is-not-failover-a-practical-multi-provider-llm-pattern.jsonld"}}