Zilliz has a Loonatic storage engine
Zilliz announced a new storage engine called Loon that supports real-time search, large-scale discovery, and batch analytics from a single copy of vector data on low-cost object storage, enabling onli…
Zilliz announced a new storage engine called Loon that supports real-time search, large-scale discovery, and batch analytics from a single copy of vector data on low-cost object storage, enabling onli…
AMD Strix Halo cluster setup guide details how to configure a two-node system linked via Intel E810 RoCE v2 for distributed vLLM inference using Tensor Parallelism. The guide covers hardware prerequis…
Researchers from Anyscale and NVIDIA demonstrated a method to scale robot policy evaluation by disaggregating GPU-heavy simulation and policy inference workloads using Ray and Isaac Lab. The approach …
Metaflow, an open-source ML/AI orchestration framework, ranks first in every category of the Cloud Native Computing Foundation's latest Technology Radar report, with 51% of surveyed users highly likel…
Ray Serve LLM, in partnership with Google Kubernetes Engine, announced major performance improvements achieving up to 4.4x higher throughput on prefill-heavy workloads and 24x higher on decode-heavy w…
The PyTorch Foundation has opened nominations for its 2026 Contributor Awards, recognizing individuals who strengthen projects like PyTorch, vLLM, DeepSpeed, Ray, Helion, and Safetensors through techn…
Engineers achieved up to 67% cost savings and 2.7x better goodput by using Prefill-Decode disaggregation with Ray and vLLM on AMD MI325X GPUs, separating prefill and decode phases onto dedicated hardw…
Alibaba's 1.7B parameter Qwen3-TTS voice cloning model was fine-tuned using Fully Sharded Data Parallel (FSDP) with PyTorch and Ray, demonstrating memory-efficient distributed training across 4 GPUs. …
Ray Serve LLM and vLLM on AMD MI325X achieve up to 67% cost savings by disaggregating prefill and decode phases in LLM serving, separating them onto dedicated GPUs to eliminate interference and improv…
Anyscale released agent skills for debugging Ray workloads, including /anyscale-platform-fix and /anyscale-platform-inspect, which automate troubleshooting of failing pipelines. A user demonstrated fi…
Two NVIDIA DGX Spark units were clustered over a 200 GbE link to serve the Qwen3-30B-A3B-Thinking model using Ray and vLLM with tensor parallelism. The setup required pinning all transport layers to t…
Adyen trained a Transaction Foundation Model on 51 trillion tokens using Ray, while Xoople, Criteo, and BMW shared their own scaling AI stories at Anyscale's Ray Day London event. The event highlighte…
At Anyscale's Ray Day: NYC event, Torc Robotics reported achieving 90% GPU utilization by consolidating its fragmented multimodal AI stack onto Ray, up from 30-40%. Discord detailed its ML platform ev…
Anyscale on Azure entered public preview, allowing Azure customers to provision the AI compute platform powered by Ray inside their own tenancy. Co-engineered with Microsoft, the integration inherits …
Anyscale released Agent Skills, a token-efficient tool for building Ray pipelines, and introduced a new maturity model for ML operations. The skills act as first responders across build, deploy, and o…
Ion Stoica, a UC Berkeley professor and co-founder of Databricks, delivered a keynote at Alumni House tracing how research problems led to foundational technologies like Apache Spark and Ray, which po…