We built the new fastest API for GLM-5.2
Baseten has built the fastest API for GLM-5.2, achieving peak speeds of 280 tokens per second and average speeds around 100 tokens per second, more than double the performance of the launch-day API as…
Baseten has built the fastest API for GLM-5.2, achieving peak speeds of 280 tokens per second and average speeds around 100 tokens per second, more than double the performance of the launch-day API as…
Helion, PyTorch's high-level DSL for writing performance-portable ML kernels, partnered with Google to build a TPU backend that compiles Helion kernels to Pallas, enabling PyTorch-friendly TPU kernel …
Unsloth achieves up to 7.3x training speedup over standard Transformers for MoE models like gpt-oss-20B on an NVIDIA B200, according to published benchmarks, while Axolotl delivers up to 1.45x speedup…
AMD Instinct MI355X GPUs with ATOM and ATOMesh achieve competitive inference performance for MiniMax-M3, a 428-billion-parameter multimodal MoE model, outperforming NVIDIA B200 and B300 in per-GPU thr…
Self-hosting open-weight large language models can save millions of dollars compared to using inference providers, according to a practical guide by Cline that analyzes the economics using Kimi K2.6 a…
The Soofi Consortium, coordinated by KI Bundesverband and funded by the German Federal Ministry for Economic Affairs and Energy, released Soofi S 30B-A3B, an open hybrid Mamba-Transformer Mixture-of-E…
A developer released the first open-source training kernels for MiniMax Sparse Attention (MSA) on Hopper and Blackwell GPUs, enabling efficient million-token training with sparse attention. The kernel…
Researchers introduced cuTile Rust, a tile-based system for safe, idiomatic GPU kernel authoring in Rust that extends Rust's ownership discipline to GPU kernels. On the NVIDIA B200 GPU, cuTile Rust ac…
AWS invented Parallel-EAGLE (P-EAGLE), a speculative decoding method that parallelizes draft token generation, achieving up to 1.69x throughput speedup over vanilla EAGLE frameworks. Amazon SageMaker …
Recursive's automated AI research system achieved state-of-the-art results on three benchmarks: fixed-budget language model training, small-model training speed, and GPU kernel optimization. The syste…