Adding a second LLM provider does not automatically make an application resilient.
The common mistake is treating provider routing, retry and failover as one decision. They answer different questions at different points in a request. When those boundaries are unclear, teams can send traffic to an incompatible provider, bypass controls or struggle to explain why a request went where it did.
A safer multi-provider design keeps the stages explicit:
Application
→ governance checks
→ capability filter
→ routing strategy
→ selected provider
After an eligible failure:
selected provider
→ retry same provider
→ compatible failover
Provider selection should not happen before policy enforcement. The request should first be associated with its workspace and credentials, then pass through the required policy, PII, guardrail and rate-limit checks.
Capability filtering comes next. A provider may be healthy but still unable to satisfy a request involving structured output, tool calling, vision or another required feature. Filtering before routing ensures that the strategy operates only on eligible providers.
Only then should the gateway select an upstream.
Static weighted routing is useful when two compatible providers should receive a controlled traffic split. In DVARA Open Source 1.8.0, a route can be expressed in gateway.yaml:
providers:
- type: openai
api_key: ${OPENAI_API_KEY}
- type: azure-openai
api_key: ${AZURE_OPENAI_API_KEY}
base_url: ${AZURE_OPENAI_BASE_URL}
routes:
- id: weighted-gpt
model: "gpt*"
strategy: weighted
providers:
- provider: openai
weight: 80
- provider: azure-openai
weight: 20
Both providers must accept the model identifier sent by the application. Routing selects a provider; it does not translate one provider's model name into another provider's deployment name.
Open Source reads this file at startup, so a routing change requires a Gateway restart. After restarting, verify the result over a meaningful sample of requests. A few calls are not enough to judge a probabilistic distribution. The response trace and structured access log's provider field show which upstream handled each request.
DVARA Open Source 1.8.0 also supports model-prefix, round-robin and canary routing. Latency-aware, cost-aware, geo-aware and intelligent routing are Enterprise capabilities, not Open Source strategies.
Routing asks: Which eligible provider should receive this request first?
Retry asks: Should the same provider receive another attempt after a retryable failure?
Failover asks: Can a compatible alternative safely take over after eligible retries or provider unavailability?
That last word—compatible—is important. A text request succeeding against a fallback does not prove that structured output, tool calling, streaming or vision will also work. Test each request shape and make failure explicit when no compatible provider remains.
Keeping recovery separate also makes the evidence easier to interpret. A trace should show the original routing decision, the provider that received the request, and whether retry or fallback occurred afterward.
The practical goal is not merely to distribute traffic. It is to preserve the same governance and evidence path regardless of which provider is selected.
Start with one model family and one operational goal: an even split, a controlled canary or a static weight. Confirm that policy and capability checks still apply, simulate an unavailable provider, and inspect the resulting trace before expanding the route.
For the complete configuration, verification steps and strategy boundaries, read the practical guide to multi-provider LLM routing.