Long Live the Short King: Why 4-hi HBM Wins
Nvidia's next-generation Rubin Ultra accelerator will ship with 192GB of HBM per GPU, down from 288GB in standard Rubin and B300, according to SemiAnalysis, which first reported the change. SemiAnalys…
Nvidia's next-generation Rubin Ultra accelerator will ship with 192GB of HBM per GPU, down from 288GB in standard Rubin and B300, according to SemiAnalysis, which first reported the change. SemiAnalys…
NVIDIA Research's Efficient AI Team and Singapore Lab released Sol-H3, an inference stack that generated a five-second, audio-synced MiniMax H3 video clip in 1.653 seconds on eight NVIDIA B300 GPUs, f…
A developer is seeking advice on running full-parameter AI models from flash storage for a security-focused local AI project aimed at elite-level code analysis and vulnerability detection, citing plan…
A developer detailed deployment recipes for Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter Mixture-of-Experts model with 95B active parameters, using vLLM. The guide covers verified GPU configurations, i…
An audit by computer science researchers finds that NVIDIA's Blackwell Ultra GPU (B300) nominally supports INT8 tensor-core compute at a 30:1 FP8-to-INT8 ratio, but the format is undeployable by defau…
NVIDIA's flagship workstation GPU, the RTX PRO 6000 "Blackwell," now retails for $16,000, making it the most expensive non-server GPU the company sells. The price hike is driven by memory shortages, a…
Runware, the AI inference startup founded by Flaviu Radulescu and Ioana Hreninciuc, launched Sonic Inference Pods, modular data centers in 20-foot shipping containers that pack about 1,200 GPUs and on…
Benchmarking by an unnamed engineer shows AMD's MI355X GPU delivers lower cost per million tokens than Nvidia's B300 for serving Kimi K3, due to memory bandwidth and quantization fitting the model on …
Wafer, an AI infrastructure provider founded by Emilio Andere and Steven Arellano, reported that running Kimi K3 on eight AMD MI355X GPUs delivered 48 tokens per second per dollar, outperforming Nvidi…
AI News Digest for August 2 covers 29 stories, including Kimi K3 on MI355X offering better performance per dollar than B300, Figure AI's F.03 ladder climb raising autonomy or OSHA concerns, and Google…
AMD's MI355X GPU achieved 48 tokens/sec per dollar running the 2.8T-parameter Kimi K3 model, outperforming Nvidia's B300 (33 tokens/sec) and B200 (7 tokens/sec), according to benchmarks from Wafer. Th…
Wafer's benchmark shows AMD's MI355X GPUs deliver 952 tok/s/node on Kimi K3, outperforming NVIDIA's B300 on performance per dollar at 48 tok/s/$ versus 33 tok/s/$, despite B300's higher aggregate thro…
LLM inference profitability depends on the trade-off between batch size and GPU count, which determines token latency and cost per million tokens. Applying this model to Kimi K3, which requires at lea…
IREN Limited reported that demand for its GPU cloud services now exceeds supply, with contracts pricing above $15 million per megawatt. The company, formerly known as Iris Energy, has secured a multi-…
Thinking Machines Lab, Inc. released Inkling, a general-purpose multimodal model with 975 billion total parameters and 41 billion active parameters, on July 15, 2026 under an Apache 2.0 license. The m…
Supermicro expanded its Data Center Building Block Solutions liquid cooling lineup with a ten-model Rear Door Heat Exchanger portfolio spanning 10kW to 120kW at the door level and up to 240kW at the r…
NVIDIA announced the BioNeMo Agent Toolkit to accelerate biomolecular structure prediction and co-folding with OpenFold3, using GPU-accelerated MSA generation and multi-GPU scaling to enable virtual s…
Fable AI has returned after a brief suspension, but the incident has set a precedent for export controls and rapid takedowns of AI models. Google is developing an 'Audio Memory' feature for Pixel phon…
Amazon SageMaker AI now supports NVIDIA Blackwell GPUs with P6-B200 instances, enabling optimized training for large AI models through expanded memory and new precision formats. The integration allows…
Nvidia's banned AI chips are selling for double their US price on China's black market, with B300 servers fetching about 7 million yuan ($1 million) after tightened US export controls. At least $1 bil…