cd/entity/H200· home› entities› H200
grep -l @h200 /news/*.json | wc -l → 83

H200

mentions 83 type Organization page 1/5 feed RSS

// recent coverage 83 mentions

13:45
2026-09-22
data.ornn.com
ai-infrastructure

The Economics of Open-Weight Inference

Ornn Data reported that the cheapest qualifying open-weight model on the Artificial Analysis Intelligence Index completes a task at roughly one fifth the cost of a comparable closed model, and that se…

18:33
2026-09-21
dev.to
large-language-models

High-Throughput LLM Inference & Training: A Deep Dive into vLLM

An engineer at g factor detailed how vLLM's PagedAttention and continuous iteration-level batching solve the memory-bandwidth bottleneck in production LLM inference, drawing on benchmarks run on dedic…

04:00
2026-09-18
arxiv.org
ai-agents

Do AI Agents Understand Computer Architecture?

A new arXiv paper (2609.19387v1) introduces AutoTuring, a benchmark that gives the same AI agent the same 15-dimensional accelerator design space twice — once as named architectural knobs with simulat…

00:00
2026-09-17
mindstudio.ai
artificial-intelligence

Agnes-3.0-Flash Preview: Specs and Benchmarks of the Open Model

Agnes AI released Agnes-3.0-Flash Preview, a 33-billion-parameter open-weight multimodal language model under an Apache 2.0 license with a 262,144-token context window and a hybrid attention architect…

01:35
2026-09-04
forum.level1techs.com
ai-infrastructure

GH200 as a personal workstation

NVIDIA's GH200 Grace Hopper Superchip modules and baseboards are appearing on eBay at prices below conventional PCIe H200 variants, prompting hobbyists to explore building custom HPC workstations from…

21:07
2026-08-26
cryptobriefing.com
ai-infrastructure

AWS, Nvidia to supply 2M GPUs for AI infrastructure expansion

AWS and Nvidia announced plans to supply an additional 2 million GPUs to enhance infrastructure for agentic and physical AI systems, building on AWS's existing deployment of over 1 million Nvidia GPUs…

17:09
2026-08-26
sourcefeed.dev
artificial-intelligence

Google's 4.7x Qwen 3.5 Speedup Is a Sharding Story

Google engineers reported a 4.7x faster prefill and 3.1x faster decode for Qwen 3.5-397B-A17B on Ironwood TPUs between April and June, achieved by using data parallelism for attention layers and exper…

06:37
2026-08-23
tradestie.com
ai-infrastructure

The market is underpricing memory bandwidth

Compute grew 106x since NVIDIA's P100 in 2016, while memory bandwidth grew only 10.9x, a 10x decline in bytes-per-FLOP that has left AI accelerators bandwidth-starved, according to a chip inventory of…

19:03
2026-08-22
cryptobriefing.com
ai-infrastructure

Nvidia notifies customers of over 15% price hikes on AI products

Nvidia has notified customers and supply chain partners of price increases exceeding 15% on AI-related GPU products, driven by rising high-bandwidth memory costs and surging AI demand. Server GPUs lik…

00:00
2026-08-21
seangoedecke.com
artificial-intelligence

Readers can't identify watermarked AI text

A quiz by software engineer Sean Godecke found that readers could not identify AI watermarked text, with 73 participants averaging 3.4 out of 10 correct guesses, close to random chance. The quiz used …

page 1 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics