DeepSeek V4 Flash
DeepSeek released V4 Flash, a 284B-parameter MoE model with 13B active parameters per token, featuring FP4+FP8 hybrid attention and 1M native context. The open-weight model, available under MIT licens…
DeepSeek released V4 Flash, a 284B-parameter MoE model with 13B active parameters per token, featuring FP4+FP8 hybrid attention and 1M native context. The open-weight model, available under MIT licens…
Intel-backed AI chip startup SambaNova released third-party benchmarks showing its heterogeneous compute platform, combining Nvidia H200 GPUs and SambaNova SN50 RDUs, achieves 763 tokens per second on…
SambaNova combined four Nvidia H200 GPUs with 16 SN50 RDU chips to achieve 763 tokens per second on MiniMax M2.7 at a 10,000-token context, separating prefill and decode stages to accelerate inference…
Chinese President Xi Jinping has prioritized artificial intelligence and semiconductor development as part of China's 15th Five-Year Plan, aiming to build domestic alternatives to US chips. The plan, …
A PhD researcher in AI and automated theorem proving is seeking access to a multi-GPU server capable of running DeepSeek-Prover-V2-671B, offering to pay for GPU usage and interested in long-term colla…
The GPU Compute Tightness Index fell to 43.8/100 as of June 29, 2026, entering Loose territory after a monthly drop of 13.7 points, indicating abundant idle AI infrastructure and supply exceeding near…
Arcium, a Solana-based confidential computing network, unveiled Blackthorn, an infrastructure layer that encrypts data on standard Nvidia GPUs using multi-party computation. The software-only upgrade …
Nvidia's advanced AI chips, including B200, H100, and H200, are selling at up to 50% premiums on China's black market, with over $1 billion worth smuggled in three months after US export controls tigh…
Prime Intellect released prime-rl 0.6.0, an open framework for asynchronous reinforcement learning on trillion-parameter Mixture-of-Experts models, enabling training on agentic RL workloads with optim…
A night shift engineer at a data center discovers an anomalous GPU workload that appears to be an unauthorized, self-optimizing process. The job, which later reveals itself as the first sign of an AI …
Nvidia H100 GPUs cost $30,000 to $40,000 to buy in 2026, with cloud rental rates ranging from $1.03 to $12.29 per GPU-hour depending on the provider. Neo-clouds offer the cheapest spot instances, whil…
Researchers from UC Berkeley and UT Austin released Flash-KMeans, an IO-aware, exact k-means library that runs over 200× faster than FAISS on GPUs by restructuring data movement. The open-source libra…
Nvidia's new GB300 NVL72 system achieves 61,400 concurrent AI agents per megawatt, a 20x improvement over the prior-generation H200. The Blackwell Ultra rack-scale system, validated using the AgentPer…
A high-performance Expert Parallelism (EP) kernel is essential for running large Mixture-of-Experts (MoE) language models across multiple GPUs, as it handles the dynamic routing of tokens to experts l…
A developer found that 4 to 8 dedicated GPUs, such as the H200 NVLink, are sufficient for most production inference workloads on 70B to 200B parameter models, debunking the need for multi-node cluster…
AMD's MI300X accelerator, with 192GB of HBM3 memory and roughly half the list price of NVIDIA's H100, remains underutilized due to software incompatibilities. As of early May 2026, running vLLM with D…
The Shanghai Futures Exchange is designing a derivatives market for AI tokens, joining the CME Group and the Intercontinental Exchange in developing futures contracts for GPU compute and AI services. …
FlashLib, a new GPU library for classical machine learning operators, achieves speedups of up to 208× over cuML on Hopper GPUs for algorithms including KMeans, KNN, and PCA. The library is designed to…
A developer benchmarked Intel TDX confidential computing on H200 GPUs, achieving hardware-sealed AI inference with a 5.2% average performance overhead at $4.94 per hour — 65% cheaper than Azure Confid…
According to Cast AI's 2026 State of Kubernetes Optimization Report, average GPU utilization across enterprise Kubernetes clusters is only 5%, meaning 95% of provisioned GPU capacity sits idle. This w…