Popping the GPU Bubble
Moondream HQ reveals that GPUs often sit idle during AI model inference due to CPU overhead, a phenomenon called the 'GPU bubble.' The company's Photon system uses pipelined decoding to overlap CPU an…
Moondream HQ reveals that GPUs often sit idle during AI model inference due to CPU overhead, a phenomenon called the 'GPU bubble.' The company's Photon system uses pipelined decoding to overlap CPU an…
DeepSpark's confidence-scheduled verifier reduces GPU waste in AI agent inference by skipping verification for low-probability tokens, dynamically adjusting thresholds under load to trade minimal accu…
Researchers accidentally discovered that ordinary CMOS transistors can function as artificial neurons and synapses, potentially enabling neuromorphic computing that is vastly more energy-efficient tha…
NANOG 97 in Bellevue, Washington, focused on AI's impact on data center design and network operations, with Cogent's Dave Schaeffer presenting statistics showing US AI infrastructure spending will rea…
Speculative decoding uses a small draft model to guess tokens and a large model to verify them, cutting AI agent latency without losing quality. The technique, formalized by Google and DeepMind in 202…
DeepSeek released DeepSpark, a speculative decoding method that accelerates large language model inference by 50–400% without retraining or quality loss. The technique uses a small draft model to gene…
Cerebras, an AI chip startup that bet against GPUs by designing a wafer-scale processor, went public in May 2026 at a $56 billion valuation, the largest US tech IPO since Uber. The company's confident…
A developer built a hands-free computer interface using a webcam and microphone, combining head tracking with MediaPipe FaceMesh and voice commands. The system enables mouse control via head gestures,…
Google's Gemma 3 270M language model can be fine-tuned for structured data extraction using PyTorch and Hugging Face libraries, according to a tutorial that walks beginners through the process of teac…
A developer patched llama.cpp to improve prompt processing throughput by 20% when using Multi-Token Prediction (MTP) on GPU, recovering performance lost to MTP overhead. The proof-of-concept processes…
MacBook GPU memory allocation depends on total system memory: up to 36GB allows 66% GPU usage, while 36GB or more allows 75%. Users can increase allocation via Terminal for AI models. Shared memory el…
AI clusters often underperform despite powerful GPUs because the GPUs are idle due to bottlenecks in data loading, CPU preprocessing, network communication, or storage contention. A developer explains…
A Hacker News user proposes that AI companies install GPU clusters in individual households and pay residents hundreds or thousands of dollars monthly, framing the idea as a potential source of univer…
Extropic unveiled thermodynamic computing hardware and algorithms that run generative AI workloads using radically less energy than GPUs. The company released its `thrml` library and plans to build a …
Unconventional AI released Un-0, an image generator that uses simulated coupled oscillators instead of neural network layers, achieving an FID of 6.74 on ImageNet 64x64. The model validates that physi…
Red Alice AI released the first official benchmark of its Version 2 architecture, reporting a 200x performance gain in the RedTensor engine. The upgrade introduces a PyTorch-backed TorchTensor backend…
Micron Technology reported record revenue of $41.46 billion for the quarter ending June 24, 2026, a 346% year-over-year surge driven by AI demand for high-bandwidth memory, which is sold out through 2…
Prefill/decode disaggregation separates the two phases of LLM inference—prefill (compute-bound) and decode (memory-bound)—onto different GPUs to avoid the performance compromise of running both on the…
AI infrastructure is entering a new phase focused on rack-scale system composition for agentic AI workflows, where CPUs play critical orchestration roles alongside accelerators. The shift from single-…
A developer built a custom Inference Optimization Engine on an NVIDIA RTX 4050 GPU to analyze how PyTorch, ONNX, and TensorRT interact with hardware, revealing that model deployment and optimization c…