cd/entity/DeepInfra· home entities DeepInfra
grep -l @deepinfra /news/*.json | wc -l → 27

DeepInfra

mentions 27 type Organization page 1/2 feed RSS

// recent coverage 27 mentions

22:21
2026-08-18
tokenstead.ai
large-language-models

Qwen3 30B A3B - cheapest: DeepInfra $0.12/M input

DeepInfra offers the cheapest API pricing for Qwen3 30B A3B at $0.12 per million input tokens, with Alibaba at $0.13 and NextBit at $0.14, according to live pricing data refreshed about 19 hours ago v…

22:21
2026-08-18
tokenstead.ai
artificial-intelligence

Gemma 4 26B A4B - cheapest: DekaLLM $0.06/M input

Google's Gemma 4 26B A4B, a 25.2B-parameter multimodal mixture-of-experts model with 3.8B active parameters and 128 experts, is now available with DekaLLM offering the cheapest API pricing at $0.06 pe…

00:00
2026-08-14
oskrim.github.io
large-language-models

Deepseek V4 Flash 0731 latency numbers from nine providers

A one-time snapshot of DeepSeek V4 Flash 0731 latency across nine inference providers found Baseten fastest with 3,980 decode tok/s and 7.72 s total p99, while Azure ran an older checkpoint and Scalew…

15:03
2026-08-13
the-decoder.com
artificial-intelligence

Ling 3.0 Flash is the smartest open model at its size

Ant Group's inclusionAI released Ling 3.0 Flash, an open-weights AI model that scores 38 points on the Artificial Analysis Intelligence Index, making it the smartest open model under 124 billion total…

21:08
2026-07-21
tokenstead.ai
large-language-models

DeepSeek V3 0324 - cheapest: DeepInfra $0.24/M input

DeepSeek V3 0324, a 671B-parameter Mixture-of-Experts model with 37B active per token and a 128k context window, is available on DeepInfra at $0.24 per million input tokens, the cheapest listed provid…

21:08
2026-07-21
tokenstead.ai
large-language-models

Qwen3.6 27B - cheapest: Morph $0.29/M input

Morph offers the cheapest API pricing for Qwen3.6 27B at $0.29 per million input tokens, according to live provider data refreshed about one hour ago via OpenRouter. The 27-billion-parameter model, re…

21:08
2026-07-21
tokenstead.ai
artificial-intelligence

Mistral Nemo 12B - cheapest: DekaLLM $0.02/M input

DekaLLM offers the cheapest inference for Mistral Nemo 12B at $0.02 per million input tokens, delivering 3,444 points per dollar based on a general score of 62. The 12.2-billion-parameter model suppor…

18:47
2026-07-21
blog.mempko.com
artificial-intelligence

Your Agentic Workflow's Cache Keepalive Costs 8x Too Much

A new measurement study of prompt cache keepalive economics across Anthropic, OpenAI, Gemini, and DeepSeek finds that the common practice of pinging every 30 seconds costs 8x more than necessary, and …

19:08
2026-07-16
byteiota.com
artificial-intelligence

Meta Compute: What Developers Need to Know (2026)

Meta is reportedly launching Meta Compute, a cloud service offering hosted Llama model APIs and raw GPU rental, directly competing with AWS, Azure, and Google Cloud. Bloomberg reported the plans on Ju…

15:02
2026-07-08
letsdatascience.com
ai-infrastructure

DeepInfra Opens Toronto AI Inference Cluster

DeepInfra opened a 1.7 MW Toronto data center on July 8, its ninth site and first outside the United States, hosting over 1,000 NVIDIA Blackwell B300 GPUs for low-latency AI inference. The expansion s…

00:00
2026-06-28
runagentrun.co.uk
large-language-models

The LLM tier that actually fits your work

Two new 2026 comparisons from DeepInfra and GMI Cloud conclude that the gap between open and closed LLMs has narrowed to 5-10% on overall capability, with no clean leaderboard existing. Closed models …

20:13
2026-06-25
artificialanalysis.ai
large-language-models

GLM-5.2 (Max) API Provider Benchmarking and Analysis

A new benchmark of 14 API providers for the GLM-5.2 (max) model reveals Fireworks leads in output speed (261.5 t/s) and low latency (9.81s), while GMI (FP8) offers the lowest blended price at $0.72 pe…

10:43
2026-06-19
github.com
ai-agents

Beast – governed output gateway for AI coding agents

Beast, a governed output gateway for AI coding agents, intercepts inputs and outputs between agents and LLM providers to enforce output contracts and repair non-compliant patches, achieving 100% task …

07:57
2026-06-18
dev.to
large-language-models

Nemotron 3 Ultra went live June 4. Here's the call that works.

NVIDIA released Nemotron 3 Ultra on June 4, 2026, a 550-billion-parameter open-weights model that achieves the highest intelligence score among US open models. The model uses a hybrid Mamba-Transforme…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics