{"slug": "editorial-re-verification-weights-biases-inference", "title": "Editorial re-verification: Weights & Biases Inference", "summary": "Weights & Biases Serverless Inference now offers API and playground access to open-source LLMs hosted on CoreWeave through an existing W&B account with no separate provider keys, and supports bringing your own LoRA weights for fine-tuned models. Published per-1M-token rates include Meta Llama 3.1 8B at $0.22 input/$0.22 output, Meta Llama 3.3 70B at $0.71/$0.71, OpenAI GPT OSS 120B at $0.03/$0.17, NVIDIA Nemotron 3.5 Lightning at $0.07/$0.20 with a $0.04 cache-hit rate, and DeepSeek V4-Pro at $1.15/$2.55 with a $0.20 cache-hit rate. Inference spend is billed separately from platform subscriptions, can be capped with a monthly W&B Inference budget, and the free allowance is stated only as 'Free credits for a limited time' on the Free plan plus a $5/mo credit on Pro, with no dollar amount or expiry published as of 2026-09-16.", "body_md": "# > Weights & Biases Inference\n\n[wandb.ai](https://docs.wandb.ai/weave/reference/service-api/inference/inference-router-openrouter-models)\n\nWeights & Biases Serverless Inference gives API and playground access to open-source LLMs hosted on CoreWeave, using an existing W&B account and no separate provider keys, with support for bringing your own LoRA weights for fine-tuned models. Token rates are listed per 1M tokens on the W&B pricing page: Meta Llama 3.1 8B at $0.22 input / $0.22 output, Meta Llama 3.3 70B at $0.71 / $0.71, OpenAI GPT OSS 120B at $0.03 / $0.17, NVIDIA Nemotron 3.5 Lightning at $0.07 / $0.20 with a $0.04 cache-hit rate, and DeepSeek V4-Pro at $1.15 / $2.55 with a $0.20 cache-hit rate. Inference spend is separate from platform subscriptions and can be capped with a monthly W&B Inference budget.\n\n### // pros\n\n- Input, output and cache-hit prices are published per 1M tokens for a long open-model list, with newer and legacy models priced side by side and context windows shown\n- Cache-hit pricing is exposed as its own column, so prompt-cache savings can be estimated before adoption rather than discovered on the invoice\n- Observability is bundled: inference runs on CoreWeave infrastructure with tracing, evaluation and monitoring through W&B Weave without extra instrumentation\n- Serverless LoRA inference lets you deploy fine-tuned weights without standing up serving infrastructure per iteration, which suits train-then-serve loops\n- A monthly inference budget can be set that stops serving requests once crossed, plus an email threshold alert, giving a hard spend ceiling\n\n### // cons\n\n- The free inference allowance is not quantified on official pages: the plan comparison only says 'Free credits for a limited time' on Free and Pro tiers\n- Inference is billed in addition to Enterprise and Pro licences, so it is not covered by an existing W&B subscription\n- ARIA adds an Agent Token Rate of $0.50 per million tokens on top of model API pricing for every token type, which raises the effective cost of agentic runs\n- Budget enforcement is explicitly not instantaneous: the FAQ warns there may be a delay in enforcing the limit and that the customer is responsible for any overage\n\nSuits teams already using W&B for experiment tracking or Weave evaluation, since one account covers open-model inference, tracing and evaluation, and the per-token rates are competitive on small and mid-size open models. Set the monthly inference budget before pointing production traffic at it, and treat the unquantified 'limited time' credits as a trial rather than a costed allowance.\n\nimported analysis — independent third party, not Uprouter editorial · verify before relying on it\n\nFree tier present — value not quantified\n\nThis provider offers a free tier, but its terms don’t convert cleanly into a USD figure (e.g. request-count limits or capacity-dependent pools). We refuse to print a misleading zero here. See the note below and the official pricing page for the exact limits.\n\nnote: Weights & Biases states only that Inference includes 'Free credits for a limited time' on the Free plan and a $5/mo credit on Pro, with no dollar amount or expiry published for the free grant (checked 2026-09-16).\n\n| Plan | Type | Monthly | Input /1M | Output /1M | Note | \n|---|---|---|---|---|---|\n| Pay-as-you-go | payg | — | Free | Free | Cheapest tracked model rate (OpenRouter-normalized). | \n\n[official pricing page](https://wandb.ai/site/pricing/inference)\n\nPricing normalized from public sources — always verify with the provider.\n\n### // model prices\n\n| Model | Input /1M | Output /1M | Context | \n|---|---|---|---|\n| [gpt-oss-120b](https://www.uprouter.online/models/gpt-oss-120b) | Free | Free | 131k | \n| [gpt-oss:20b](https://www.uprouter.online/models/gpt-oss-20b) | Free | Free | 131k | \n| [Qwen3 235B A22B Thinking 2507](https://www.uprouter.online/models/qwen/qwen3-235b-a22b-thinking-2507) | Free | $0.0000 | 131k | \n| [GLM 4.5](https://www.uprouter.online/models/z-ai/glm-4.5) | $0.0000 | $0.0000 | 131k | \n| [DeepSeek V3.1](https://www.uprouter.online/models/deepseek/deepseek-chat-v3.1) | $0.0000 | $0.0000 | 164k | \n\n## $ Does Weights & Biases Inference have a free tier?\n\nYes — Weights & Biases Inference offers a free tier. Always confirm current limits on Weights & Biases Inference's official pricing page — free tiers change without notice.\n\n## $ How much does Weights & Biases Inference cost per 1M tokens?\n\nThe cheapest model we track at Weights & Biases Inference is Free per 1M input tokens and Free per 1M output tokens. This is normalized from Weights & Biases Inference's published pricing — verify with the provider before purchasing, since prices change frequently.\n\n## $ Is Weights & Biases Inference safe to use?\n\nUprouter rates Weights & Biases Inference as low risk. Risk level: moderate. Rates are published and a customer-set monthly budget can halt serving, which limits runaway cost, but the FAQ notes enforcement can lag and that overages remain the customer's responsibility. The free inference grant is described only as credits 'for a limited time' with no amount or expiry on official pages, so it cannot be relied on for capacity planning; data flows through W&B's hosted platform and CoreWeave rather than the customer's own infrastructure. Risk ratings are editorial, evidence-based, and never influenced by affiliate relationships.\n\n## $ Can I use Weights & Biases Inference through Uprouter Connect?\n\nYes — Weights & Biases Inference is Connect-compatible. Add your Weights & Biases Inference API key in Uprouter Connect (encrypted at rest) and route requests to it through one OpenAI- and Claude-compatible API, optionally behind failover aliases. Uprouter meters the traffic in compute at the provider's listed rate.\n\nRaw community sentiment, shown unblended. Votes never feed into objective data fields like price, risk, or status — and never into sort order.\n\nNo published reviews yet — be the first.\n\n## // more_informationprice history · live status · risk & confidence · code snippets\n\n### price_history\n\nNo recorded changes yet.\n\n### live_status\n\ncurrent: unknown\nNo probes recorded yet.\n\nUprouter runs its own probes; best-effort, not a guarantee. Probe results depend on our network location and cadence — for production, always use the provider’s own status page.\n\n### risk_&_confidence\n\nRisk level: moderate. Rates are published and a customer-set monthly budget can halt serving, which limits runaway cost, but the FAQ notes enforcement can lag and that overages remain the customer's responsibility. The free inference grant is described only as credits 'for a limited time' with no amount or expiry on official pages, so it cannot be relied on for capacity planning; data flows through W&B's hosted platform and CoreWeave rather than the customer's own infrastructure.\n\nConfidence reflects how much verifiable evidence (official docs, probe history, community reports) backs this entry. Risk is our editorial assessment of operator reliability and terms — not financial advice.\n\n### code_snippets\n\n```\ncurl https://www.uprouter.online/api/connect/v1/chat/completions \\  -H \"Authorization: Bearer upr_live_YOUR_KEY\" \\  -H \"Content-Type: application/json\" \\  -d '{\n    \"model\": \"z-ai/glm-4.5\",\n    \"messages\": [{ \"role\": \"user\", \"content\": \"Hello via Weights & Biases Inference\" }]\n  }'\n```\n\nMetered in Uprouter compute — the usage log shows which upstream served each request and its exact cost.\n\n[Baichuanfree $11.20pricing n/aconnectMedium risk](https://www.uprouter.online/s/baichuan)\n\n[Baidu (ERNIE)free tier (n/q)in from Free/1MconnectMedium risk](https://www.uprouter.online/s/baidu)\n\n[Baidu Qianfanfree tier (n/q)in from Free/1MconnectMedium risk](https://www.uprouter.online/s/qianfan)\n\n[Codex Cloudfree $0in from $0.0000/1MconnectLow risk](https://www.uprouter.online/s/codex-cloud)\n\nRanked by neutral similarity signals only — live status, Connect compatibility, free-tier shape and free-credit proximity. Commercial relationships never influence this list.", "url": "https://wpnews.pro/news/editorial-re-verification-weights-biases-inference", "canonical_source": "https://www.uprouter.online/s/wandb", "published_at": "2026-09-16 21:31:24+00:00", "updated_at": "2026-09-16 22:23:38.484028+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-tools", "ai-products"], "entities": ["Weights & Biases", "CoreWeave", "Meta Llama 3.1 8B", "Meta Llama 3.3 70B", "OpenAI GPT OSS 120B", "NVIDIA Nemotron 3.5 Lightning", "DeepSeek V4-Pro", "W&B Weave"], "alternates": {"html": "https://wpnews.pro/news/editorial-re-verification-weights-biases-inference", "markdown": "https://wpnews.pro/news/editorial-re-verification-weights-biases-inference.md", "text": "https://wpnews.pro/news/editorial-re-verification-weights-biases-inference.txt", "jsonld": "https://wpnews.pro/news/editorial-re-verification-weights-biases-inference.jsonld"}}