cd /news/ai-infrastructure/aurora-ai-gateway-fork-for-multi-ip-… Β· home β€Ί topics β€Ί ai-infrastructure β€Ί article
[ARTICLE Β· art-123342] src=github.com β†— pub= topic=ai-infrastructure verified=true sentiment=Β· neutral

Aurora: AI gateway fork for multi-IP setups, 55x faster than LiteLLM

A developer known as entitybtw has released Aurora, a self-hosted AI gateway fork of aurorallm/aurora that claims to be 55x faster than LiteLLM and supports 14 provider types with OpenAI and Anthropic compatibility. The fork adds a redesigned dashboard, provider CRUD, fallback chains, provider pools, and a Session Hub, all managed via UI with persistence across restarts. The project is written in Go, licensed under Apache 2.0, and is available as a Docker image `entbtw/aurora`.

read21 min views2 publishedSep 8, 2026
Aurora: AI gateway fork for multi-IP setups, 55x faster than LiteLLM
Image: Michielbdejong (auto-discovered)

One API for every AI provider. Self-hosted. No vendor lock-in.

14 provider types β€’ OpenAI & Anthropic compatible β€’ Go β€’ Apache 2.0 β€’ Built for raw speed

Full guides, written for this fork.

Guide What it covers
Getting Started first run, build, config, basic usage, OpenAI-compatible client
Deployment Docker / Docker Compose, persistent state files, multi-IP host networking
Multi-account pools end-to-end: load-balanced accounts with distinct, stable client identities
Session Hub header transformation & session mapping engine, header modes, API reference, dashboard
Docker image published image entbtw/aurora , tags, how to build & publish

Quick deploy:

docker pull entbtw/aurora:latest
docker run -d --name aurora -p 8080:8080 -e AURORA_MASTER_KEY="your-secure-key" entbtw/aurora:latest

See Deployment for production (persistent config & state, network_mode: host for multi-IP).

Dashboard-driven operations β€” no more .env-only workflows for the things you change most. Everything below is managed from the UI and persists across restarts.

Warning: This fork contains custom features not present in the original aurorallm/aurora. Some features (dashboard redesign, session hub, UI enhancements) were vibecoded and may contain rough edges. Designed for advanced API integration workflows β€” use at your own discretion.

  • Redesigned dashboard β€” fullCatppuccin theme, mobile-responsive, compact/touch-friendly layout, clean auth/logo/sidebar, sharedSearchInput fix in audit logs & usage.
  • Provider CRUD β€” manage providers from the UI (base URL, API key, models, type). Per-providerbind_ip ,pool_only , runtime enable/disable, live rename, duplicate protection. Status shows if a key is setwithout exposing it . OpenRouter list is now anallowlist ;vLLM type added to the dashboard (was.env -only).
  • Custom User-Agent β€” set a customUser-Agent header per provider for upstream attribution (e.g. OpenRouter recommends this for credits).
  • Auto-fetch models toggle β€” disable automatic/models discovery per provider to use only explicitly configured model lists.
  • Fallback chains β€” edit rules in the UI, applied atruntime ; callable by name, exposed in/v1/models , order preserved on toggle/edit/delete.
  • Provider pools β€” create/edit/delete with member selection andweighted / round-robin strategies; health-aware members,pool_only models, live registry rebuild.
  • Response headers β€” configurableX-Actual-Provider /X-Actual-Model /X-Requested /X-Fallback-Chain , per-header toggles, custom headers, success/error/always modes, emitted on429 /401 .
  • Persistence β€” state saved toconfigs/provider-overrides.json ,configs/pool-overrides.json ,configs/fallback.json (env-overridable); Docker volumes keep it across recreation.
  • Session Hub β€” header transformation engine with per-provider/pool session mapping, inboundβ†’outbound unique ID generation, disk persistence with live toggle, and pool-aware binding via UI (seeSession Hub below).

Aurora sits between your app and LLM providers. Your app sends requests using the standard OpenAI or Anthropic SDK β€” Aurora routes them to whichever provider you've configured. One format handles everything β€” you dont need to worry about provider-specific formats.

client = OpenAI(base_url="https://api.openai.com/v1", api_key="sk-...")

client = OpenAI(base_url="http://localhost:8080/v1", api_key="your-aurora-key")

No SDK changes. No format changes. Just swap the base_url.

  • 14 provider types β€” OpenAI, Anthropic, Gemini, Groq, DeepSeek, OpenRouter, xAI, Z.ai, MiniMax, Azure OpenAI, Oracle, Ollama, vLLM, Jina

  • Auto-discovery β€” set an API key as an env var, restart, provider + all its models appear automatically

  • Auto-fetch toggle β€” disable per-provider model auto-discovery to use only explicitly configured model lists

  • Custom User-Agent β€” set a customUser-Agent header per provider for upstream attribution or branding

  • Provider pools β€” group multiple keys/endpoints, load-balance with round-robin or weighted distribution, health-aware failover

  • Model aliases β€” rename/remap any model to a custom identifier across the entire gateway

  • Model overrides β€” enable or disable specific models per user path, persisted via dashboard oruser_pricing.yaml

  • Fallback β€” automatic failover on 5xx/429, or manual rules (from config or external JSON) mapping failed provider+model to backups

  • Resilience β€” exponential backoff with jitter, circuit breaker per provider (closed β†’ open β†’ half-open), per-provider override of global retry/circuit-breaker settings

  • Multiple instances β€” runOPENAI_EAST_API_KEY andOPENAI_WEST_API_KEY as separate providers

  • Custom base URLs β€” override any provider's endpoint (corporate proxies, regional endpoints)

  • Passthrough β€”/p/{provider}/* for full upstream API access (not just chat completions); filter which provider types get passthrough routes

  • Config-driven workflows β€” per-request routing, caching, guardrail, audit, usage, budget, and fallback behavior controlled by persisted workflow documents

  • OpenAI-compatible β€”/v1/chat/completions ,/v1/embeddings ,/v1/rerank ,/v1/models ,/v1/files ,/v1/batches

  • Responses API β€”/v1/responses with full CRUD, cancel, input items, compact

  • Anthropic-compatible β€”/v1/messages ,/v1/messages/count_tokens (native Anthropic wire format); optional dedicated ingress at/v1/messages

  • Streaming β€” SSE streaming for all endpoints, preserved end-to-end

  • Keep-only-aliases mode β€” hide raw provider models from/v1/models and expose only aliased names

  • Configured provider models mode β€”fallback (add listed models to auto-discovered) orallowlist (only serve explicitly listed models)

  • Exact cache β€” SHA-256 hash match on request, Redis-backed, async writes

  • Semantic cache β€” vector similarity with configurable threshold, supports Qdrant, pgvector, Pinecone, Weaviate

  • Prompt cache β€” forwardscache_control to Anthropic/OpenAI/Gemini native prompt caching; configurable modes (auto ,manual ,off ), component toggles, and minimum token threshold

  • Model registry cache β€” local filesystem + Redis, offline-safe; supports vendored JSON snapshots with per-field user pricing overrides

  • Master key β€” top-level gateway auth

  • Managed API keys β€” scoped, rate-limited, per-key model authorization, usage stats

  • Rate limiting β€” per-key rate limiting backed by in-memory or Redis

  • PII redaction β€” email, phone, SSN, credit card detection and masking

  • Prompt injection blocking β€” detects and blocks injection attempts

  • System prompt protection β€” inject, override, or decorate system prompts

  • Regex blocking β€” custom pattern matching with block or sanitize actions

  • Length limits β€” character/token count enforcement on requests

  • LLM-based altering β€” guardrail that rewrites message content via an auxiliary LLM call (anonymization, custom prompts)

  • Guardrail direction & ordering β€” run before provider dispatch (input ), after response (output ), or both; same-order guardrails run in parallel

  • Batch guardrails β€” apply configured guardrails to inline items in/v1/batches requests

  • Audit logging β€” full request/response capture, buffered writes, configurable retention (body/header logging, buffer size, flush interval), live SSE stream

  • Usage analytics β€” per-model token counting, cost tracking, daily aggregation by model/user-path, pricing recalculation action

  • Prometheus metrics β€”aurora_requests_total ,aurora_request_duration_seconds ,aurora_requests_in_flight , plus gateway phase timing

  • Admin dashboard β€” React SPA built into the Go binary (Catppuccin, fully mobile-responsive): full provider CRUD, fallback chains, provider pools, response-header config, plus models, aliases, guardrails, cache, usage, audit, auth keys, workflows, console, playground

  • pprof endpoints β€” Go runtime profiling at/debug/pprof/* (heap, goroutine, mutex, block, threadcreate)

  • Structured logging β€” configurable format (JSON/text), level (debug/info/warn/error), source info, service metadata

  • Token saver β€” policy-driven output compression (profiles: concise, caveman, ultra, wenyan); scoped to specific models/providers via include/exclude filters; configurable on-error behavior (allow/block)

  • Pricing management β€” per-model pricing overrides, recalculation, import/export

  • Usage budgets β€” per-key usage tracking and limits, per-request budget enforcement via workflow feature flags

  • Single binary β€”docker pull entbtw/aurora (this fork) or run from source with Go

  • CLI β€” run from source, or drive via config files + the dashboard

  • CLI tools API β€” admin REST endpoints for CLI configuration sync, gated separately

  • Swagger docs β€”/swagger/index.html (build-tag gated)

  • Config profiles β€” pre-built configs for local, local-power, and team deployments

  • 3-layer config β€” code defaults β†’ config.yaml β†’ env vars (env vars win)

Header transformation engine for API integration workflows where upstream services require unique client identifiers per account.

  • Per-provider/pool binding β€” attach transformation rules to specific providers, pools, fallbacks, or all targets (* )
  • 6 header modes β€”map (stable inboundβ†’outbound per provider),generate (fresh ID each request),passthrough ,static ,random_from_list ,remove
  • Pool-aware β€” rules bound to a pool automatically apply to all member providers
  • Inbound header forwarding β€” client session headers are forwarded through the translation layer somap mode works even when the provider path drops arbitrary inbound headers
  • Lock-free hot path β€”Apply() is a single atomic map read; benchmarked at ~495 ns/op (negligible)
  • Persistent or in-memory β€” toggled live via API or dashboard (PUT /admin/api/v1/sessionhub/storage {"mode":"disk"} )
  • Dashboard UI β€” Settings β†’ Session Hub: binding overview from live server targets (pools/providers), add rule by selecting target, live mapping viewer, storage toggle
  1. Client sends request to Aurora (e.g. with x-opencode-session: ses_abc123 )
  2. Gateway intercepts the inbound session header and stores it in request context
  3. Request is routed to a pool member (e.g. opencode-zen β†’vllm-zen-backup )
  4. Provider's outbound headerSetter fires: session hub applies rules for that provider/pool
  5. map mode: inboundses_abc123 β†’ unique outboundses_xR4f8k2m... per provider (stable, deduplicated)
  6. generate mode: fresh randomses_... per request (always unique)
  7. Additional headers (x-opencode-client ,user-agent ) are injected per rule
  8. Outbound request goes to upstream with transformed headers

Rules are persisted in configs/session-hub-rules.yaml (gitignored). Live edits via API or dashboard are auto-saved.

enabled: true
mapping_storage: disk          # "memory" or "disk"
providers:
  opencode-zen:                # matches pool name or provider name
    enabled: true
    headers:
      - name: x-opencode-session
        mode: map              # stable inbound→outbound per provider
        prefix: "ses_"
        length: 28
      - name: x-opencode-client
        mode: static
        value: cli
      - name: user-agent
        mode: static
        value: "opencode/1.18.26 ai-sdk/openai/2.0.0 runtime/bun/1.0.0"
Method Endpoint Description
GET /admin/api/v1/sessionhub/status Stats + storage_mode
GET /admin/api/v1/sessionhub/providers List bound rules
POST /admin/api/v1/sessionhub/providers Create rule
PUT /admin/api/v1/sessionhub/providers/:name Update rule
DELETE /admin/api/v1/sessionhub/providers/:name Delete rule
GET /admin/api/v1/sessionhub/mappings List live mappings
DELETE /admin/api/v1/sessionhub/mappings Clear all mappings
PUT /admin/api/v1/sessionhub/storage Toggle memory /disk
POST /admin/api/v1/sessionhub/apply Test transform
Mode Behavior
map First request generates unique outbound value per provider; subsequent requests with same inbound reuse it
generate Fresh random value every request
passthrough Original value forwarded unchanged
static Fixed value (set value: )
random_from_list Random pick from values: list
remove Strip header entirely

Start routing AI traffic in 60 seconds.

Recommended β€” Docker (published image):

docker pull entbtw/aurora:latest
docker run -d --name aurora -p 8080:8080 -e AURORA_MASTER_KEY="your-secure-key" entbtw/aurora:latest

Full examples below. For production (persistent state, multi-IP) see the Deployment guide.

The quickest way to configure providers from scratch is the dashboard: http://localhost:8080/admin/dashboard β†’ Providers β†’ Add provider. For env-var driven setups:

Linux / macOS #

AURORA_MASTER_KEY=your-secure-key \
  OPENAI_API_KEY=sk-... \
  ANTHROPIC_API_KEY=sk-ant-... \
  GEMINI_API_KEY=... \
  GROQ_API_KEY=gsk_... \
  DEEPSEEK_API_KEY=... \
  OPENROUTER_API_KEY=... \
  XAI_API_KEY=... \
  ZAI_API_KEY=... \
  MINIMAX_API_KEY=... \
  AZURE_API_KEY=... \
  ORACLE_API_KEY=... \
  OLLAMA_API_KEY=... \
  VLLM_API_KEY=... \
  JINA_API_KEY=... \
  LOGGING_ENABLED=true \
  METRICS_ENABLED=true \
  GUARDRAILS_ENABLED=true \
  TOKEN_SAVER_ENABLED=true \
  aurora

Windows PowerShell #

$env:AURORA_MASTER_KEY="your-secure-key"; `
$env:OPENAI_API_KEY="sk-..."; `
$env:ANTHROPIC_API_KEY="sk-ant-..."; `
$env:GEMINI_API_KEY="..."; `
$env:GROQ_API_KEY="gsk_..."; `
$env:DEEPSEEK_API_KEY="..."; `
$env:OPENROUTER_API_KEY="..."; `
$env:XAI_API_KEY="..."; `
$env:ZAI_API_KEY="..."; `
$env:MINIMAX_API_KEY="..."; `
$env:AZURE_API_KEY="..."; `
$env:ORACLE_API_KEY="..."; `
$env:OLLAMA_API_KEY="..."; `
$env:VLLM_API_KEY="..."; `
$env:JINA_API_KEY="..."; `
$env:LOGGING_ENABLED="true"; `
$env:METRICS_ENABLED="true"; `
$env:GUARDRAILS_ENABLED="true"; `
$env:TOKEN_SAVER_ENABLED="true"; `
aurora

Windows CMD #

set AURORA_MASTER_KEY=your-secure-key ^
  && set OPENAI_API_KEY=sk-... ^
  && set ANTHROPIC_API_KEY=sk-ant-... ^
  && set GEMINI_API_KEY=... ^
  && set GROQ_API_KEY=gsk_... ^
  && set DEEPSEEK_API_KEY=... ^
  && set OPENROUTER_API_KEY=... ^
  && set XAI_API_KEY=... ^
  && set ZAI_API_KEY=... ^
  && set MINIMAX_API_KEY=... ^
  && set AZURE_API_KEY=... ^
  && set ORACLE_API_KEY=... ^
  && set OLLAMA_API_KEY=... ^
  && set VLLM_API_KEY=... ^
  && set JINA_API_KEY=... ^
  && set LOGGING_ENABLED=true ^
  && set METRICS_ENABLED=true ^
  && set GUARDRAILS_ENABLED=true ^
  && set TOKEN_SAVER_ENABLED=true ^
  && aurora

Published image: entbtw/aurora Β· tags latest, v1.0.0.

docker pull entbtw/aurora:latest
docker run -d --name aurora -p 8080:8080 \
  -e AURORA_MASTER_KEY="your-secure-key" \
  -e OPENAI_API_KEY="sk-..." \
  -e ANTHROPIC_API_KEY="sk-ant-..." \
  -e GEMINI_API_KEY="..." \
  -e GROQ_API_KEY="gsk_..." \
  -e DEEPSEEK_API_KEY="..." \
  -e OPENROUTER_API_KEY="..." \
  -e XAI_API_KEY="..." \
  -e ZAI_API_KEY="..." \
  -e MINIMAX_API_KEY="..." \
  -e AZURE_API_KEY="..." \
  -e ORACLE_API_KEY="..." \
  -e OLLAMA_API_KEY="..." \
  -e VLLM_API_KEY="..." \
  -e JINA_API_KEY="..." \
  -e LOGGING_ENABLED=true \
  -e METRICS_ENABLED=true \
  -e GUARDRAILS_ENABLED=true \
  -e TOKEN_SAVER_ENABLED=true \
  entbtw/aurora:latest

For production setups (persistent config/state, multi-IP host networking) see the Deployment guide.

After starting, confirm the gateway is up and the dashboard loads:

curl -s http://localhost:8080/health

open http://localhost:8080/admin/dashboard

curl -s http://localhost:8080/admin/api/v1/sessionhub/status \
  -H "Authorization: Bearer your-master-key"

If health returns {"status":"ok"} β€” the gateway is running. Now add a provider via the dashboard or env vars, then test a model call:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer your-master-key" \
  -d '{"model":"groq/llama-4-scout-17b-16e-instruct","messages":[{"role":"user","content":"Hello!"}]}'

curl http://localhost:8080/v1/messages \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer your-master-key" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "anthropic/claude-sonnet-5-20260630",
    "max_tokens": 1024,
    "stream": true,
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

curl http://localhost:8080/v1/embeddings \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer your-master-key" \
  -d '{"model":"openai/text-embedding-3-small","input":"Hello world"}'

curl http://localhost:8080/v1/rerank \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer your-master-key" \
  -d '{"model":"jina/jina-reranker-v2-base-multilingual","query":"test","documents":["doc1","doc2"]}'

Dashboard: http://localhost:8080/admin/dashboard

Docs (this fork): Getting Started Β· Deployment Β· Session Hub Β· Docker image

Source: github.com/entitybtw/aurora Β· Image: hub.docker.com/r/entbtw/aurora

Base project (upstream): aurorallm/aurora Β· aurorallm.online/docs

Providers are auto-discovered from environment variables. Set any provider's _API_KEY and restart β€” the provider and its default models appear automatically.

Security note: The env var names below are documentation references. Actual secrets go into your .env file (in .gitignore) or your deployment secrets manager β€” never commit them.

Provider Env var Default base URL Requires base URL API key required Default models
OpenAI OPENAI_API_KEY https://api.openai.com/v1 No Yes gpt-5.6-sol ,gpt-5.6-luna
Anthropic ANTHROPIC_API_KEY https://api.anthropic.com/v1 No Yes claude-sonnet-5 ,claude-fable-5
Google Gemini GEMINI_API_KEY https://generativelanguage.googleapis.com/v1beta/openai No Yes gemini-3.1-pro ,gemini-3.5-flash
Groq GROQ_API_KEY https://api.groq.com/openai/v1 No Yes llama-4-scout-17b ,llama-4-maverick-17b ,qwen3-32b
DeepSeek DEEPSEEK_API_KEY https://api.deepseek.com No Yes deepseek-v4-pro ,deepseek-v4-flash
OpenRouter OPENROUTER_API_KEY https://openrouter.ai/api/v1 No Yes 300+ models
xAI (Grok) XAI_API_KEY https://api.x.ai/v1 No Yes grok-4.5 ,grok-4.3
Z.ai ZAI_API_KEY https://api.z.ai/api/paas/v4 No Yes glm-5.2
MiniMax MINIMAX_API_KEY https://api.minimax.io/v1 No Yes minimax-m3
Azure OpenAI AZURE_API_KEY β€” Yes Yes Your deployments
Oracle ORACLE_API_KEY β€” Yes Yes cohere.command-r-plus
Ollama OLLAMA_API_KEY http://localhost:11434/v1 No No (optional) Any local model
vLLM VLLM_API_KEY http://localhost:8000/v1 No No (optional) Any served model
Jina (reranker) JINA_API_KEY β€” Yes Yes jina-embeddings-v3

Every provider supports *_MODELS to override auto-discovered models:

OPENAI_MODELS=gpt-5.6-sol,gpt-5.6-terra,gpt-5.6-luna

Custom base URL:

OPENAI_BASE_URL=https://my-corp-openai-proxy.example.com/v1

YAML provider config supports additional options:

providers:
  openai:
    type: openai
    api_key: "${OPENAI_API_KEY}"
    base_url: "https://api.openai.com/v1"
    user_agent: "MyApp/1.0"
    auto_fetch_models: false
    models:
      - gpt-4o
      - gpt-4o-mini

Multiple instances of the same provider (underscores become hyphens in the provider name):

OPENAI_EAST_API_KEY=sk-...     # β†’ provider: openai-east
OPENAI_WEST_API_KEY=sk-...     # β†’ provider: openai-west

Azure requires API version:

AZURE_API_VERSION=2024-10-21

OpenRouter extras:

OPENROUTER_SITE_URL=https://github.com/entitybtw/aurora
OPENROUTER_APP_NAME=Aurora Gateway

The gateway loads settings in this priority order (later wins):

code defaults β†’ config.yaml β†’ .env / environment variables

Generated by aurora init, every section of config.yaml is documented inline:

Section What it controls
server Port, base path, master key, passthrough, Anthropic ingress
admin Dashboard API and UI
models Discovery, overrides, allowlisting
storage SQLite (default), PostgreSQL, or MongoDB
logging Audit logging of requests/responses
usage Token tracking, pricing, retention
metrics Prometheus endpoint
guardrails Content safety filters
cache Model cache, response cache (exact + semantic)
combos Multi-model combo definitions
token_saver Output compression
fallback Provider failover rules
resilience Retry + circuit breaker
workflows Policy-based request routing

Pre-built configs in configs/editions/:

Profile File Use case
OSS oss.env.example Minimal local β€” SQLite, no Redis
OSS Local Power oss.local-power.env.example SQLite + Redis exact cache
OSS Team oss.team.env.example Postgres + Redis + Qdrant β€” full team deployment
export AURORA_CONFIG_PATH=configs/editions/oss.team.example.yaml

Server & Security #

Env var Default Description
PORT 8080 HTTP listening port
BASE_PATH / URL path prefix to mount under
AURORA_MASTER_KEY "" Master API key for auth
BODY_SIZE_LIMIT 10M Max request body size
SWAGGER_ENABLED false Enable Swagger UI at /swagger/index.html
PPROF_ENABLED false Enable pprof at /debug/pprof/
ENABLE_PASSTHROUGH_ROUTES true Provider-native passthrough at /p/{provider}
ALLOW_PASSTHROUGH_V1_ALIAS true Allow /p/{provider}/v1/... alias routes
ENABLED_PASSTHROUGH_PROVIDERS openai,anthropic,openrouter,zai,vllm Provider types for passthrough
ENABLE_ANTHROPIC_INGRESS false Expose /v1/messages for native Anthropic clients
DISABLE_REQUEST_LOGGING false Turn off request logging
DISABLE_REQUEST_BODY_SNAPSHOT false Don't snapshot request bodies
DISABLE_PASSTHROUGH_SEMANTIC_ENRICHMENT false Disable semantic enrichment on passthrough

HTTP Client & Proxy #

Env var Default Description
HTTP_TIMEOUT 600 Upstream request timeout (seconds)
HTTP_RESPONSE_HEADER_TIMEOUT 600 Timeout for upstream response headers
HTTP_PROXY β€” HTTP proxy URL for upstream calls
HTTPS_PROXY β€” HTTPS proxy URL
NO_PROXY β€” Hosts to exclude from proxy

Storage #

Env var Default Description
STORAGE_TYPE sqlite Backend: sqlite ,postgresql , ormongodb
SQLITE_PATH data/aurora.db SQLite database file path
POSTGRES_URL β€” PostgreSQL connection string
POSTGRES_MAX_CONNS 10 PostgreSQL connection pool max
MONGODB_URL β€” MongoDB connection string
MONGODB_DATABASE aurora MongoDB database name

Model Registry #

Env var Default Description
MODEL_LIST_URL https://raw.githubusercontent.com/aurorallm/aurora/refs/heads/main/docs-assets/assets/models.json External model metadata registry
MODEL_LIST_LOCAL_PATH data/models.local.json Local model registry snapshot path
MODEL_LIST_USER_OVERRIDES_PATH data/user_pricing.yaml User pricing override file
MODELS_ENABLED_BY_DEFAULT true Default enabled state for provider models
MODEL_OVERRIDES_ENABLED true Allow per-model overrides
KEEP_ONLY_ALIASES_AT_MODELS_ENDPOINT false Hide provider models, show only aliases
CONFIGURED_PROVIDER_MODELS_MODE fallback fallback orallowlist

Caching #

Model cache:

Env var Default Description
CACHE_REFRESH_INTERVAL 3600 Model registry cache refresh (seconds)
AURORA_CACHE_DIR .cache Local filesystem cache directory
REDIS_URL β€” Redis connection URL (enables Redis-backed model cache)
REDIS_KEY_MODELS aurora:models Redis key for model cache
REDIS_TTL_MODELS 86400 Redis model cache TTL (seconds)

Response cache (exact match):

Env var Default Description
RESPONSE_CACHE_SIMPLE_ENABLED false Enable Redis exact-response cache
REDIS_KEY_RESPONSES aurora:response: Redis key prefix for responses
REDIS_TTL_RESPONSES 3600 Response cache TTL (seconds)

Semantic cache (vector similarity):

Env var Default Description
SEMANTIC_CACHE_ENABLED false Enable semantic cache
SEMANTIC_CACHE_THRESHOLD 0.92 Similarity threshold (0-1)
SEMANTIC_CACHE_PROMPT_SIMILARITY 0.90 Prompt similarity threshold
SEMANTIC_CACHE_TTL 3600 Entry TTL (seconds)
SEMANTIC_CACHE_MAX_CONV_MESSAGES 3 Recent conversation messages to embed
SEMANTIC_CACHE_EXCLUDE_SYSTEM_PROMPT false Exclude system prompt from cache key
SEMANTIC_CACHE_EMBEDDER_PROVIDER openai Embedder provider name
SEMANTIC_CACHE_EMBEDDER_MODEL text-embedding-3-small Embedder model
SEMANTIC_CACHE_VECTOR_STORE_TYPE qdrant Backend: qdrant ,pgvector ,pinecone ,weaviate
SEMANTIC_CACHE_QDRANT_URL http://localhost:6333 Qdrant URL
SEMANTIC_CACHE_QDRANT_COLLECTION aurora_semantic Qdrant collection name
SEMANTIC_CACHE_QDRANT_API_KEY β€” Qdrant API key
SEMANTIC_CACHE_PGVECTOR_URL β€” pgvector connection string
SEMANTIC_CACHE_PGVECTOR_TABLE aurora_semantic_cache pgvector table name
SEMANTIC_CACHE_PGVECTOR_DIMENSION 1536 pgvector embedding dimension
SEMANTIC_CACHE_PINECONE_HOST β€” Pinecone host URL
SEMANTIC_CACHE_PINECONE_API_KEY β€” Pinecone API key
SEMANTIC_CACHE_PINECONE_NAMESPACE β€” Pinecone namespace
SEMANTIC_CACHE_PINECONE_DIMENSION 1536 Pinecone embedding dimension
SEMANTIC_CACHE_WEAVIATE_URL β€” Weaviate URL
SEMANTIC_CACHE_WEAVIATE_CLASS AuroraSemanticCache Weaviate class name
SEMANTIC_CACHE_WEAVIATE_API_KEY β€” Weaviate API key

Audit Logging #

Env var Default Description
LOGGING_ENABLED false Enable audit log to storage
LOGGING_LOG_BODIES true Log request/response bodies
LOGGING_LOG_HEADERS true Log headers (sensitive headers redacted)
LOGGING_ONLY_MODEL_INTERACTIONS true Skip health/metrics/admin endpoints
LOGGING_BUFFER_SIZE 1000 In-memory queue capacity
LOGGING_FLUSH_INTERVAL 5 Flush interval (seconds)
LOGGING_RETENTION_DAYS 30 Auto-delete after N days (0 = forever)

Usage Tracking #

Env var Default Description
USAGE_ENABLED true Enable token usage tracking
USAGE_PRICING_RECALCULATION_ENABLED true Allow admin pricing recalculation
ENFORCE_RETURNING_USAGE_DATA true Add stream_options.include_usage=true to streaming requests
USAGE_BUFFER_SIZE 1000 In-memory queue capacity
USAGE_FLUSH_INTERVAL 5 Flush interval (seconds)
USAGE_RETENTION_DAYS 90 Auto-delete after N days (0 = forever)

Guardrails #

Env var Default Description
GUARDRAILS_ENABLED false Enable content safety filters globally
ENABLE_GUARDRAILS_FOR_BATCH_PROCESSING false Apply guardrails to /v1/batches items

Metrics #

Env var Default Description
METRICS_ENABLED false Enable Prometheus /metrics endpoint
METRICS_ENDPOINT /metrics Metrics endpoint path

Token Saver #

Env var Default Description
TOKEN_SAVER_ENABLED false Enable output compression
TOKEN_SAVER_ENDPOINTS chat_completions Endpoints to apply it to
TOKEN_SAVER_APPLY_STREAMING true Apply to streaming responses
TOKEN_SAVER_OUTPUT_ENABLED false Enable output style/profile
TOKEN_SAVER_OUTPUT_PROFILE concise Profile: concise ,caveman ,ultra ,wenyan
TOKEN_SAVER_MODELS_INCLUDE β€” Models to include (comma-separated)
TOKEN_SAVER_MODELS_EXCLUDE β€” Models to exclude
TOKEN_SAVER_PROVIDERS_INCLUDE β€” Providers to include
TOKEN_SAVER_PROVIDERS_EXCLUDE β€” Providers to exclude
TOKEN_SAVER_ON_ERROR allow Behavior on error: allow orblock
TOKEN_SAVER_EMIT_HEADERS true Emit token-saver headers in response
TOKEN_SAVER_AUDIT_ENABLED true Log token-saver actions

Resilience #

Env var Default Description
RETRY_MAX_RETRIES 3 Upstream retry count
RETRY_INITIAL_BACKOFF 1s Initial backoff duration
RETRY_MAX_BACKOFF 30s Maximum backoff duration
RETRY_BACKOFF_FACTOR 2.0 Exponential backoff multiplier
RETRY_JITTER_FACTOR 0.1 Random jitter fraction
CIRCUIT_BREAKER_FAILURE_THRESHOLD 5 Failures before circuit opens
CIRCUIT_BREAKER_SUCCESS_THRESHOLD 2 Successes before circuit closes
CIRCUIT_BREAKER_TIMEOUT 30s Time before half-open retry

Fallback #

Env var Default Description
FEATURE_FALLBACK_MODE manual Fallback mode: auto ,manual , oroff
FALLBACK_MANUAL_RULES_PATH β€” Path to manual fallback rules JSON

Admin & Features #

Env var Default Description
ADMIN_ENDPOINTS_ENABLED true Enable /admin/api/v1/* REST endpoints
ADMIN_UI_ENABLED true Enable /admin/dashboard UI
COMBOS_ENABLED true Enable combo model calls
CLI_TOOLS_ENABLED true Enable CLI tools integration
CLI_TOOLS_APPLY_ENABLED false Allow admin/API to apply tool changes
WORKFLOW_REFRESH_INTERVAL 1m Workflow refresh interval from storage
EDITION β€” Edition identifier (Enterprise use)

Config file path #

Env var Default Description
AURORA_CONFIG_PATH configs/config.yaml Override path to config YAML

Run the built binary directly (from source: go build -o aurora ./apps/aurora, then ./aurora). The npm iaurora wrapper is the upstream package and isn't republished by this fork.

Command Description
aurora Start the gateway server (default port 8080)
aurora init Scaffold config.yaml ,.env ,data/ in current directory
aurora models sync Download upstream model registry to local file
aurora models diff Show pricing diff between upstream and local snapshot
aurora models show Print effective pricing for a model after merging overrides
aurora -version Print version information
aurora -help Show all CLI options and config reference
aurora -help-json Dump env var schema as JSON
aurora/
β”œβ”€β”€ apps/              # Application entrypoints
β”œβ”€β”€ internal/          # Core packages (providers, gateway, storage, guardrails, etc.)
β”œβ”€β”€ dashboard-ui/      # React admin dashboard (Vite)
β”œβ”€β”€ configs/           # Configuration profiles and examples
β”œβ”€β”€ documentation/     # Markdown docs (Getting Started, Deployment, Session Hub, Docker)
β”œβ”€β”€ docs-assets/       # Images, models.json, assets
β”œβ”€β”€ monitoring/        # Prometheus + Grafana configs
β”œβ”€β”€ bench-results/     # Benchmark data
β”œβ”€β”€ release/           # Release scripts
└── scripts/           # Build and utility scripts

This project is licensed under the Apache 2.0 License β€” see the LICENSE file for details.

Community fork of Aurora. Session Hub features and multi-account integration built by entitybtw. The upstream project is built by the Aurora team.

── more in #ai-infrastructure 4 stories Β· sorted by recency
── more on @entitybtw 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/aurora-ai-gateway-fo…] indexed:0 read:21min 2026-09-08 Β· β€”