cd/entity/H100· home› entities› H100
grep -l @h100 /news/*.json | wc -l → 225

H100

mentions 225 type Organization page 1/12 feed RSS

// recent coverage 225 mentions

16:21
2026-10-03
jaehun.me
large-language-models

The KV Cache Is the New Memory Wall

A Systematization of Knowledge paper, "The KV Cache Is the New Memory Wall" (arXiv:2609.30854) by Tejinder Singh of Dell Technologies, unifies five families of KV cache reduction techniques — quantiza…

20:12
2026-10-02
siliconangle.com
ai-infrastructure

IBM and CoreWeave co-design controls for agent workloads

IBM Research and CoreWeave Inc. co-designed identity management and workload isolation controls for agent workloads, extending IBM's internal identity systems into CoreWeave and using CoreWeave Sandbo…

01:53
2026-10-01
intelligence.mts.now
ai-infrastructure

The Physical Stack Behind AI

A new attributed record of AI compute infrastructure catalogs 83 facilities, 482 GPU clusters and 2,935 price observations, finding that official H100 rental prices range from $4.57 to $18.37 per acce…

17:47
2026-09-25
dev.to
ai-infrastructure

Where to rent or buy h100 and b200 gpus for ai startups

An engineering analysis of GPU infrastructure economics finds that owning an 8x H100 SXM5-class server carries roughly $310,000 in upfront capital plus about $37,500 in electricity and $4,300 per year…

09:00
2026-09-25
dev.to
ai-infrastructure

How to Monitor GPU Usage and Temperature on a Remote Server

A developer published a command-line guide for monitoring GPU utilization and temperature on remote servers, using nvidia-smi queries, cron-scheduled CSV logging, webhook-based temperature alerts, and…

14:08
2026-09-24
huggingface.co
large-language-models

Accelerating vision-language models with LFM2.5-VL-DSpark

LiquidAI released LFM2.5-VL-DSpark, a 279.5M-parameter speculative decoding drafter for its LFM2.5-VL-3B vision-language model that adds 8.9% to the target model's parameter count while delivering dec…

02:45
2026-09-24
jax-ml.github.io
machine-learning

Sharded matrices and how to multiply them

A new installment in the "How To Scale Your Model" series lays out a theory of sharded matrix multiplication for training large ML models across thousands of accelerators, using named-axis notation to…

00:00
2026-09-23
mindstudio.ai
artificial-intelligence

Bonsai 2 27B: A 27B Model That Runs in Under 6GB on a Laptop

Prism ML released Bonsai 2 27B, a ternary-quantized version of Qwen3.8-27B that stores weights as -1, 0, or +1 and fits in roughly 6 to 8.6GB while retaining 98.2% of the original FP16 model's benchma…

13:45
2026-09-22
data.ornn.com
ai-infrastructure

The Economics of Open-Weight Inference

Ornn Data reported that the cheapest qualifying open-weight model on the Artificial Analysis Intelligence Index completes a task at roughly one fifth the cost of a comparable closed model, and that se…

19:58
2026-09-21
computeassay.com
ai-infrastructure

A graded registry of GPU cloud prices and fabric disclosure

Compute Assay launched ASSAY-1, a free, versioned specification, registry, and weekly report that defines the GPU-hour as a graded deliverable good, after finding a 3.8× price spread across H100 capac…

page 1 / 12 next →
// co-occurs with top 8 entities
// topics top 6 topics