cd /news/large-language-models/routing-is-not-failover-a-practical-… · home › topics › large-language-models › article
[ARTICLE · art-140999] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Routing Is Not Failover: A Practical Multi-Provider LLM Pattern

DVARA Open Source 1.8.0 introduces explicit multi-provider LLM routing that separates provider selection, retry and failover into distinct stages, with governance and capability checks applied before any upstream is chosen. The release supports static weighted, model-prefix, round-robin and canary routing configured in gateway.yaml, while latency-aware, cost-aware and geo-aware strategies remain Enterprise-only. The project warns that a text request succeeding against a fallback does not prove structured output, tool calling, streaming or vision will work, so each request shape must be tested and failure made explicit when no compatible provider remains.

by read3 min views1 publishedSep 28, 2026

Adding a second LLM provider does not automatically make an application resilient.

The common mistake is treating provider routing, retry and failover as one decision. They answer different questions at different points in a request. When those boundaries are unclear, teams can send traffic to an incompatible provider, bypass controls or struggle to explain why a request went where it did.

A safer multi-provider design keeps the stages explicit:

Application
  → governance checks
  → capability filter
  → routing strategy
  → selected provider

After an eligible failure:
  selected provider
  → retry same provider
  → compatible failover

Provider selection should not happen before policy enforcement. The request should first be associated with its workspace and credentials, then pass through the required policy, PII, guardrail and rate-limit checks.

Capability filtering comes next. A provider may be healthy but still unable to satisfy a request involving structured output, tool calling, vision or another required feature. Filtering before routing ensures that the strategy operates only on eligible providers.

Only then should the gateway select an upstream.

Static weighted routing is useful when two compatible providers should receive a controlled traffic split. In DVARA Open Source 1.8.0, a route can be expressed in gateway.yaml:

providers:
  - type: openai
    api_key: ${OPENAI_API_KEY}
  - type: azure-openai
    api_key: ${AZURE_OPENAI_API_KEY}
    base_url: ${AZURE_OPENAI_BASE_URL}

routes:
  - id: weighted-gpt
    model: "gpt*"
    strategy: weighted
    providers:
      - provider: openai
        weight: 80
      - provider: azure-openai
        weight: 20

Both providers must accept the model identifier sent by the application. Routing selects a provider; it does not translate one provider's model name into another provider's deployment name.

Open Source reads this file at startup, so a routing change requires a Gateway restart. After restarting, verify the result over a meaningful sample of requests. A few calls are not enough to judge a probabilistic distribution. The response trace and structured access log's provider field show which upstream handled each request.

DVARA Open Source 1.8.0 also supports model-prefix, round-robin and canary routing. Latency-aware, cost-aware, geo-aware and intelligent routing are Enterprise capabilities, not Open Source strategies.

Routing asks: Which eligible provider should receive this request first?

Retry asks: Should the same provider receive another attempt after a retryable failure?

Failover asks: Can a compatible alternative safely take over after eligible retries or provider unavailability?

That last word—compatible—is important. A text request succeeding against a fallback does not prove that structured output, tool calling, streaming or vision will also work. Test each request shape and make failure explicit when no compatible provider remains.

Keeping recovery separate also makes the evidence easier to interpret. A trace should show the original routing decision, the provider that received the request, and whether retry or fallback occurred afterward.

The practical goal is not merely to distribute traffic. It is to preserve the same governance and evidence path regardless of which provider is selected.

Start with one model family and one operational goal: an even split, a controlled canary or a static weight. Confirm that policy and capability checks still apply, simulate an unavailable provider, and inspect the resulting trace before expanding the route.

For the complete configuration, verification steps and strategy boundaries, read the practical guide to multi-provider LLM routing.

── more in #large-language-models 4 stories · sorted by recency
── more on @dvara open source 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/routing-is-not-failo…] indexed:0 read:3min 2026-09-28 · —