cd/entity/Baseten· home› entities› Baseten
grep -l @baseten /news/*.json | wc -l → 82

Baseten

mentions 82 type Organization page 2/5 feed RSS

// recent coverage 82 mentions

11:48
2026-09-09
testingcatalog.com
large-language-models

Inception launches Mercury 2.5 at 1,107 tokens per second

Inception released Mercury 2.5, which it calls the most capable diffusion LLM on the market and the largest it has ever trained, reporting a 40% intelligence gain over Mercury 2 at 1,107 tokens per se…

23:03
2026-09-08
trackllm.net
ai-infrastructure

z-ai/glm-5.3-flash @ baseten/fp8: B3IT change (TV 0.61)

OpenRouter retired behavioral monitoring of the z-ai/glm-5.3-flash endpoint hosted at baseten/fp8 on 2026-09-17 after re-initialization repeatedly timed out following a detected change, leaving no log…

20:14
2026-09-08
inceptionlabs.ai
artificial-intelligence

Mercury 2.5

Inception AI released Mercury 2.5, its most capable production model, claiming a 40% increase in intelligence over Mercury 2 with speeds of 1,107 tokens per second and a context window of 260K tokens,…

21:50
2026-09-06
frontierroles.com
ai-infrastructure

Staff Software Engineer, Model Infrastructure — Harvey

Harvey, an AI company for legal and professional services, is hiring a Staff Software Engineer for Model Infrastructure in San Francisco with a salary of $231k–340k/yr. The role involves leading the d…

02:46
2026-09-02
forgeeks.net
machine-learning

LLM inference now has two ways to get cheaper

Baseten's technical breakdown of LLM inference efficiency identifies two categories of engineering choices: those that trade latency for throughput, such as batch sizing, tensor parallelism, expert pa…

03:15
2026-08-17
vibeleaderboard.ai
artificial-intelligence

Inference now sets the bill, and open weights are closing the gap

OpenAI's compute chief Sachin Katti said inference will account for over 80 percent of all AI compute spend, with a roadmap to 30GW, while Crusoe's Chase Lockmiller reported that power, not GPU supply…

00:00
2026-08-14
oskrim.github.io
large-language-models

Deepseek V4 Flash 0731 latency numbers from nine providers

A one-time snapshot of DeepSeek V4 Flash 0731 latency across nine inference providers found Baseten fastest with 3,980 decode tok/s and 7.72 s total p99, while Azure ran an older checkpoint and Scalew…

18:39
2026-08-13
firecrawl.dev
large-language-models

What Is Kimi K3? A Complete Developer Guide for 2026

Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight model, on July 27, 2026, making it the largest open-weight model as of August 2026. It features 104B active parameters per token, a 1…

22:10
2026-08-11
frontierroles.com
ai-infrastructure

Software Engineer - AI Developer Productivity — Baseten

Baseten, an AI infrastructure company that raised a $1.5B Series F led by Altimeter Capital, Conviction Partners, and Spark Capital, is hiring a Software Engineer for AI Developer Productivity in San …

13:00
2026-08-11
coderabbit.ai
machine-learning

Teaching NVIDIA Nemotron 3.5 Lightning to route code reviews

CodeRabbit, NVIDIA, and Baseten post-trained NVIDIA Nemotron 3.5 Lightning to route code reviews, achieving 80.7% exact route agreement (up from 75.8% for the GPT baseline) and cutting estimated infer…

12:45
2026-08-09
read.technically.dev
artificial-intelligence

Dispatch: Kimi K3 licensing, Liquid AI, and stacked PRs on GitHub

Moonshot AI released the weights for its Kimi K3 model with 2.8 trillion parameters and a 1M token context window under a new Kimi K3 License, which allows inference providers like Modal, Baseten, Fir…

00:00
2026-08-06
huggingface.co
artificial-intelligence

Baseten on Hugging Face Inference Providers 🔥

Hugging Face has added Baseten as a supported Inference Provider on its Hub, enabling serverless access to open-weight LLMs such as Kimi K3, DeepSeek V4 Flash, and GLM-5.2 for conversational and text-…

10:11
2026-08-05
byteiota.com
artificial-intelligence

Kimi K3 Open Weights: Self-Hosting Reality Check

Moonshot AI released the weights for Kimi K3 on July 27, a 2.8-trillion-parameter Mixture-of-Experts model with a 1-million-token context window, scoring third globally behind Claude Fable 5 Max and G…

12:20
2026-08-04
machinelearningmastery.com
artificial-intelligence

Static vs. Dynamic vs. Continuous Batching in LLM Inference

IBM's article explains that static, dynamic, and continuous batching are methods to improve GPU utilization in large language model inference, with continuous batching scheduling at the token level to…

← prev page 2 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics