CUDA Multi-Process Service
NVIDIA's CUDA Multi-Process Service (MPS) enables concurrent kernel execution across multiple processes, improving GPU utilization compared to default time-slicing, according to a technical blog post …
NVIDIA's CUDA Multi-Process Service (MPS) enables concurrent kernel execution across multiple processes, improving GPU utilization compared to default time-slicing, according to a technical blog post …
A technical explainer on distributed training and inference details how work is divided across CPUs, GPUs, and multi-node clusters, distinguishing data, pipeline, and tensor parallelism and sharding, …
Extropic, a Boston-based hardware startup, is hiring junior ML scientists for its Thermo ML Resident program, offering a salary of $75,000–$200,000 per year for on-site work. Residents will collaborat…
A developer detailed the construction of an end-to-end imitation learning pipeline for robotic manipulation, covering dataset loading, normalization, model architecture, and training. The tutorial emp…
A developer documented writing GPU kernels in Triton to understand PyTorch operations, starting with vector add and fused ReLU/dropout, highlighting the benefit of fused kernels in reducing memory rou…
Nagarro, a digital product engineering company with 15,000+ experts across 26 countries, is hiring a Senior Staff Engineer - Data Science in Guadalajara, Mexico, requiring 8+ years of experience in da…
TraceML, an open-source PyTorch training diagnosis tool, reproduced a known LeRobot training regression and found that recurring DataLoader fetches took over five seconds, while most fetches were quic…
A developer's source-code walkthrough of vLLM 0.22's V1 execution path traces a single offline inference request from the LLM.generate() API through inter-process communication, scheduling, GPU execut…
A user reports that Qwen-Image-Edit-2511 struggles with one-to-one replacement of multiple identical objects in a scene, often removing all but one instance or misaligning viewpoints despite reference…
CellularFlow, a continual-learning LLM architecture from developer celcilin, replaces dense feed-forward networks with Multi-Head Associative DNA Memory Banks and an Episodic Memory Slot Buffer, achie…
Onur Satici of SpiralDB presented a data-loading pipeline that streams columnar data from S3 directly to GPUs at 13 gigabits per second, bypassing NVMe and CPU decompression bottlenecks, using the ope…
NanoLM Studio V4, a local desktop workbench for building small decoder-only language models, has been publicly released by its developer. The Tk-based application integrates document ingestion, ByteLe…
A developer known as No Saved DATA introduced Neve, a new programming language designed to unify high-level Python-like syntax with low-level efficiency for deep learning. The language features LLVM J…
42dot, the Global Software Center of Hyundai Motor Group, is hiring a Staff VLA Engineer in Sunnyvale, California, with a salary range of $189,000–$311,000 per year, to research and prototype next-gen…
Developer gpjt has uploaded PyTorch-compatible versions of all his JAX-trained GPT-2 models to the Hugging Face Hub, including models from his blog series on writing an LLM from scratch and Chinchilla…
A developer explains why Python is the preferred language in the data field, citing its simple syntax, powerful libraries such as pandas and PyTorch, and a large supportive community. The post also pr…
Indicate, a new open-source tool for transliterating 12+ Indic languages to and from English, was released on Hacker News, offering both a PyTorch-based local model and LLM backends with auto-detectio…
WiCi One, a wireless GPU device from an unnamed company, is now available for pre-order at $1,999 USD for early signups, down from $2,599, and claims to deliver real-time AI inference and gaming perfo…
PyTorch 2.14, released by the PyTorch team, introduces NVGEMM for CuTeDSL-generated CUTLASS kernels, a new nccl2 backend for distributed training, fault tolerance as a first-class concept, native line…
An engineer analyzed 559 public CLAUDE.md files and found that 62.9% of the content is documentation rather than actionable rules, with only 43.5% of actual rules being mechanically checkable. The fin…