PyTorch 2.14 Release Blog
PyTorch 2.14, released by the PyTorch team, introduces NVGEMM for CuTeDSL-generated CUTLASS kernels, a new nccl2 backend for distributed training, fault tolerance as a first-class concept, native line…
PyTorch 2.14, released by the PyTorch team, introduces NVGEMM for CuTeDSL-generated CUTLASS kernels, a new nccl2 backend for distributed training, fault tolerance as a first-class concept, native line…
PyTorch Conference North America 2026, scheduled for October 20-21, will feature agentic AI and next-generation intelligence across keynotes and sessions, including talks by Sara Hooker of Adaption on…
PyTorch Conference North America 2026, held October 20–21 in San Jose, CA, will feature vLLM across multiple sessions on KV cache management, disaggregated serving, hardware portability, kernel optimi…
PyTorch Conference North America 2026, scheduled for October 20–21 in San Jose, CA, will feature Core PyTorch sessions covering compiler and runtime internals, distributed communication, device portab…
The PyTorch Ecosystem Working Group welcomed 10 new projects to the PyTorch Ecosystem Landscape, including Perforated, AReaL, TorchJD, RLinf, Miles, SMG, FiftyOne, TokenSpeed, VisualTorch, and TorchSu…
PyTorch Conference North America 2026 announced its keynote speaker lineup for the October 20–21 event in San Jose, California, featuring sessions on PyTorch updates, native PyTorch on Trainium, agent…
IBM researchers demonstrated that AI agents can write small runtime adapters enabling stock HuggingFace Transformers models to run on IBM's Spyre AI accelerator, achieving full enablement for thousand…
AMD and PyTorch upstreamed FP8 training optimizations for AMD Instinct GPUs into TorchAO and TorchTitan, delivering a 13.4% throughput gain over BF16 on Llama3-8B dense models and recovering 89% of FP…
Meta introduced Muse Glimmer, an open-weight 30-billion-parameter model distilled from Muse Spark for on-device agentic workflows, with ExecuTorch adding end-to-end support for running it on NVIDIA GP…
PyTorch Conference North America 2026 will be held in San Jose, California, on October 20–21, 2026, with keynote speakers including Mark Collier, Executive Director of the PyTorch Foundation; Mazin Gi…
The inaugural Santa Cruz PyTorch Meetup, organized by Red Hat's Steve and UCSC OSPO's Stephanie Lieggi, drew 45 local engineers, students, and leaders for GPU/CUDA talks and lightning presentations on…
Meta's FBTriton infrastructure powers custom GPU compiler innovations like TLX and autoWS while staying synced with upstream Triton using agentic ingestion and a stratified L1/L2/L3 validation framewo…
The PyTorch Foundation announced a design contest for the 2026 flare pin for PyTorch Conference North America, with the winning entrant receiving one complimentary ticket to the conference in San Jose…
PyTorch serves as both a reference language and an implementation language for deep learning, with reference implementations in plain PyTorch used to verify the correctness of optimized production ker…
Helion, PyTorch's high-level DSL for writing performance-portable ML kernels, partnered with Google to build a TPU backend that compiles Helion kernels to Pallas, enabling PyTorch-friendly TPU kernel …
The PyTorch Foundation, which expanded into a multi-project foundation in April 2025, now hosts six projects including PyTorch, vLLM, DeepSpeed, Ray, Helion, and Safetensors. PyTorch released version …
The PyTorch Conference North America schedule is live, with the event taking place in San Jose on October 20–21, featuring sessions on training, inference, compiler innovations, responsible AI, and th…
The PyTorch-Triton 3.7 release introduces the Triton Plugin Extensions system, a framework for dynamically loading custom compiler passes, dialects, and DSL extensions into upstream Triton at runtime …
PyTorch 2.13 has been released with performance improvements including FlexAttention on Apple Silicon with up to 12x speedups, a CuTeDSL backend for Inductor, and fused nn.LinearCrossEntropyLoss reduc…
AMD has brought PyTorch Monarch to its Instinct GPUs with ROCm, enabling single-controller distributed training for large language models. The port addresses reliability challenges at scale by providi…