cd/entity/NCCL· home› entities› NCCL
grep -l @nccl /news/*.json | wc -l → 28

NCCL

mentions 28 type Organization page 1/2 feed RSS

// recent coverage 28 mentions

02:20
2026-09-26
quasiben.github.io
ai-infrastructure

Faster Transport on Cloud Infra

AWS EFA's SRD transport moves CUDA buffers roughly 40x faster than tuned TCP, about 13x faster through a shuffle, and roughly 2x faster on the PDS-H Q9 benchmark, according to a blog post recording co…

02:01
2026-09-23
academy.dair.ai
artificial-intelligence

Wiki Foundation Model for Complex Agentic Reasoning

Researchers Junnan Dong, Linhao Luo, Senlei Zhang and colleagues at Tencent Youtu Lab and Monash University proposed WFM, a Wiki Foundation Model for encoding and retrieving LLM Wiki knowledge bases t…

16:06
2026-09-17
byteiota.com
machine-learning

PyTorch 2.14: Fault-Tolerant Training and Apple Silicon Fixes

PyTorch 2.14 shipped on September 15 with 2,995 commits from 487 contributors, adding fault-tolerant distributed training as a first-class framework concept via a rewritten NCCL backend (nccl2) with i…

04:00
2026-09-17
machinebrief.com
ai-agents

WFM: Wiki Foundation Model for Complex Agentic Reasoning

Researchers proposed WFM, a Wiki Foundation Model for scalable, agent-native knowledge representation and retrieval, detailed in arXiv paper 2609.18182v1. WFM formalizes a Wiki Graph schema coupling d…

02:03
2026-09-17
skypilot.ai
machine-learning

RL Is Everything, Everywhere, All at Once

Reinforcement learning has become the standard final stage of training frontier language models, with GRPO on verifiable rewards now the standard recipe since DeepSeek-R1 popularized it, according to …

17:01
2026-09-15
promptcube3.com
ai-infrastructure

NVLink 6 handles failures so AI factories don't stop

Nvidia's NVLink 6 interconnect adds multi-layer resiliency designed to keep large GPU clusters running through link failures instead of crashing entire training and inference jobs. The system reroutes…

05:46
2026-07-29
inmyhead.is
artificial-intelligence

Preempting the Prefill

A new paper, FlowPrefill by Hsieh et al., proposes preempting long LLM inference prefills mid-forward-pass to rescue urgent requests that would otherwise miss their time-to-first-token (TTFT) service-…

16:52
2026-07-27
i-programmer.info
artificial-intelligence

Programming Massively Parallel Processors, 5th Ed(Morgan Kaufmann)

The 5th edition of 'Programming Massively Parallel Processors' by Wen-mei W. Hwu, David B. Kirk, and Izzat El Hajj, published by Morgan Kaufmann, introduces new chapters on filtering, wavefront parall…

23:00
2026-07-01
databricks.com
ai-infrastructure

How we keep GPUs reliable across Databricks AI

Databricks AI engineers detailed how they maintain GPU reliability at scale, describing failure modes including crashed jobs, silent slowdowns, and numerical corruption, and outlining a multi-stage he…

17:20
2026-06-27
uccl-project.github.io
developer-tools

rdmatop: Cross-Provider Htop for RDMA Traffic

The UCCL team released rdmatop, a real-time terminal UI that monitors RDMA traffic across any Linux device including NVIDIA ConnectX, AWS EFA, and Broadcom NICs. The tool reads RDMA netlink to provide…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics