cd/entity/LiteLLM· home entities LiteLLM
grep -l @litellm /news/*.json | wc -l → 160

LiteLLM

mentions 160 type Organization page 6/8 feed RSS

// recent coverage 160 mentions

07:17
2026-06-27
byteiota.com
ai-policy

Multi-Provider AI Gateway: Build It Before the Next Ban

On June 12, the US government ordered Anthropic to disable Fable 5 globally within six hours, leaving developers without access and highlighting the risk of regulatory takedowns. The following day, GP…

01:00
2026-06-26
agentgateway.dev
ai-infrastructure

Benchmarking Agentgateway vs LiteLLM Part 2: Fixed Throughput

AgentGateway outperformed LiteLLM in a fixed-throughput benchmark at 3,000 QPS, sustaining 2,998.94 QPS with sub-millisecond P99 latency while LiteLLM achieved only 2,465.89 QPS with 30.6 ms P99 laten…

00:00
2026-06-26
agentgateway.dev
ai-infrastructure

Benchmarking Agentgateway vs LiteLLM

A benchmark comparing agentgateway and LiteLLM found agentgateway handled over 11× more requests per second (36,933 QPS vs 3,198 QPS) with sub-2 ms P99 latency while consuming under 30 MB of RAM, vers…

12:37
2026-06-25
dev.to
ai-tools

Lite-Harness SDK

LiteLLM launched the Lite-Harness SDK, enabling developers to swap between AI agent harnesses such as Claude Code and Codex without rewriting application code. The SDK provides a unified query interfa…

06:20
2026-06-24
github.com
large-language-models

I built an LLM router that doesn't use an LLM

Developer Lore released Wayfinder, an open-source LLM router that determines whether to send a prompt to a local or cloud model by analyzing structural features like length, headings, and code, withou…

19:13
2026-06-20
dev.to
large-language-models

We Cut Our LLM API Bill 30% With Four Lines of YAML

A developer at a company handling thousands of LLM calls per hour cut their API bill by 30% using semantic caching. By embedding prompts and checking cosine similarity against cached responses, they a…

13:26
2026-06-19
dev.to
large-language-models

Fault-injecting our LLM provider to trust Bifrost fallbacks

Buildkite ran a game day that fault-injected OpenAI with 429s and 500s to test whether Bifrost's fallback config would reroute requests for an LLM-backed build-failure summariser. After fixing a retry…

16:29
2026-06-18
devashish.me
large-language-models

Two Qwen3 models on one DGX Spark: the residency math

Alibaba's Qwen3-80B and Qwen3-4B models were successfully co-located on a single NVIDIA DGX Spark using vLLM containers behind a LiteLLM proxy, but the 80B model's inability to emit tool calls in auto…

← prev page 6 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics