cd/entity/H100· home› entities› H100
grep -l @h100 /news/*.json | wc -l → 225

H100

mentions 225 type Organization page 4/12 feed RSS

// recent coverage 225 mentions

05:53
2026-08-24
dev.to
ai-infrastructure

Stop Comparing GPU Clouds Only by $/Hour

A developer argues that comparing GPU cloud providers by hourly rate alone is misleading, and that the true cost metric is cost per unit of work done, factoring in utilization, data transfer, provisio…

18:01
2026-08-23
pub.towardsai.net
large-language-models

Tuning vLLM: What Every Setting Does to the Arithmetic

VLLM v0.27.1, the open-source inference engine, has unified its scheduler around a fixed token budget per step, making chunked prefill, prefix caching, and speculative decoding stackable by default, a…

06:37
2026-08-23
tradestie.com
ai-infrastructure

The market is underpricing memory bandwidth

Compute grew 106x since NVIDIA's P100 in 2016, while memory bandwidth grew only 10.9x, a 10x decline in bytes-per-FLOP that has left AI accelerators bandwidth-starved, according to a chip inventory of…

00:00
2026-08-23
tomtunguz.com
ai-infrastructure

The AI Bullwhip

The AI infrastructure supply chain is experiencing a multi-year bullwhip effect, with shortages cascading from GPUs to memory, CPUs, and storage, and data center construction costs soaring to over $20…

14:51
2026-08-22
promptcube3.com
ai-infrastructure

Nvidia is doubling down on physical infrastructure by partnering

Nvidia is partnering with data center developer Cloverleaf to build AI-native facilities optimized for its H100 and upcoming Blackwell GPUs, addressing power density and liquid cooling needs. The coll…

00:19
2026-08-22
github.com
developer-tools

Mockup Nvidia GPUs on Linux systems

A new open-source tool, Mock Nvidia GPUs, lets Linux developers simulate NVIDIA GPU telemetry for monitoring software by intercepting NVML and nvidia-smi calls, without emulating CUDA or actual GPU ha…

21:26
2026-08-21
promptcube3.com
artificial-intelligence

Nvidia's latest demo proves the inference stack matters more

Nvidia's latest demo shows that optimizing the inference stack, not the model weights, is the key to performance, achieving 4.2× higher throughput at half the latency on identical H100 hardware with L…

21:13
2026-08-21
promptcube3.com
artificial-intelligence

Zhipu's Mythos benchmark leak suggests GLM-4.

Zhipu AI's leaked Mythos benchmark suggests its upcoming GLM-4 model scores 87.2% on MMLU-Pro, 78.5% on GPQA-Diamond, and 99.1% on a 128k-context retrieval test, potentially outperforming GPT-4o on MM…

18:01
2026-08-20
promptcube3.com
ai-safety

OpenAI admits safety monitoring eats 20% of inference compute

OpenAI has admitted that safety monitoring consumes 20% of inference compute, a significant operational cost that translates to $6,000 annually per H100 GPU. The admission signals that supervision is …

15:14
2026-08-20
promptcube3.com
artificial-intelligence

Rethinking LLM scaling after Jie Tang's latest breakdown

Jie Tang's team at Z.ai argues that LLM scaling should optimize for fixed inference budgets, favoring smaller, deeper architectures trained longer over wider models. Their ablation shows a 7B model tr…

07:49
2026-08-20
promptcube3.com
artificial-intelligence

H200 chips quietly entering China despite export controls

Nvidia's H200 AI chips are quietly entering China through gray-market channels despite U.S. export controls, with volumes in the hundreds, enough for cluster bring-ups and inference optimization but n…

19:02
2026-08-19
promptcube3.com
artificial-intelligence

Cerebras WSE-3 smokes H100 on Llama 3 70B inference at a

Cerebras Systems claims its Wafer-Scale Engine 3 (WSE-3) delivers 1,800 tokens per second at 23 kW for Llama 3 70B inference, versus about 850 tokens per second at 56 kW for an 8×H100 DGX system, yiel…

17:59
2026-08-19
cryptobriefing.com
ai-infrastructure

CFTC seeks public comment on derivatives markets for AI compute

The Commodity Futures Trading Commission has filed a Request for Comment on derivatives contracts tied to computing capacity, seeking feedback on cash market liquidity, manipulation risks, customer pr…

← prev page 4 / 12 next →
// co-occurs with top 8 entities
// topics top 6 topics