cd/entity/H200· home entities H200
grep -l @h200 /news/*.json | wc -l → 62

H200

mentions 62 type Organization page 1/4 feed RSS

// recent coverage 62 mentions

03:06
2026-08-19
insideai.news
ai-chips

Nvidia H200 Chips Reach China in Small Shipments, FT Reports

Nvidia's H200 processors have reached mainland China in limited quantities, with ByteDance and Tencent each receiving about 10,000 units in recent weeks, according to the Financial Times. U.S. regulat…

16:45
2026-08-18
promptcube3.com
ai-infrastructure

Stopping your laptop from killing your AI agents mid-task is a

Modal, the cloud compute platform, announced a new feature that lets AI agents run inside full KVM virtual machines, providing kernel-level access and real GPU drivers, with hardware ranging from 1 vC…

20:09
2026-08-15
byteiota.com
artificial-intelligence

Cloudflare Unweight: Lossless LLM Compression on H100

Cloudflare released Unweight, a lossless compression system that reduces BF16 MLP weights in LLMs by 15–22% while producing bit-identical outputs, freeing roughly 3 GB of VRAM on Llama-3.1-8B running …

05:22
2026-08-14
dev.to
ai-infrastructure

NVIDIA GPU roadmap explained: from A100 to H200 and beyond

NVIDIA's data center GPU roadmap has advanced through four major architectures—Ampere, Hopper, Blackwell, and Vera Rubin—each delivering more memory, faster interconnects, and lower-precision compute …

17:37
2026-08-11
runtimewire.com
artificial-intelligence

Nvidia ships Nemotron 3.5 Lightning for single-GPU AI agents

Nvidia released Nemotron 3.5 Lightning on August 11, an open-weight reasoning model with 30 billion total parameters and 3 billion active per token, designed to run agent workloads on a single Nvidia …

04:00
2026-08-11
arxiv.org
machine-learning

Training Variable Long Sequences with Data-Centric Parallel

Researchers introduced Data-Centric Parallel (DCP), a method that dynamically adjusts runtime settings based on each batch's sequence length to train deep learning models on variable long sequences, a…

23:08
2026-08-03
sourcefeed.dev
artificial-intelligence

Cloudflare's Quantization Math Is Right. The Disclosure Isn't

Cloudflare published an engineering post detailing how it quantizes models on Workers AI, including Moonshot's Kimi K2.6 with FP8 KV cache and Z.ai's GLM 5.2 with INT4 weights, achieving up to 41% hig…

21:09
2026-08-03
sourcefeed.dev
artificial-intelligence

Quantize the Decode, Not the Prefill

Cloudflare's production benchmarks for Moonshot's Kimi K2.6 and Z.ai's GLM 5.2 show that quantizing the decode phase, not the prefill, yields the biggest cost and throughput gains for trillion-paramet…

01:31
2026-08-02
softwareseni.com
artificial-intelligence

How to Model the True Cost per Token Across GPU Architectures

Nvidia's Jensen Huang announced at GTC that Vera Rubin delivers approximately 10x more inference throughput per watt compared to Blackwell (B200), but the comparison point matters because most organiz…

page 1 / 4 next →
// co-occurs with top 8 entities
// topics top 6 topics