cd/entity/NVIDIA H200· home entities NVIDIA H200
grep -l @nvidia h200 /news/*.json | wc -l → 16

NVIDIA H200

mentions 16 type Person feed RSS

// recent coverage 16 mentions

00:00
2026-08-17
mindstudio.ai
artificial-intelligence

Qwen3.8-27B AEON Uncensored: How This Abliteration Actually Works

A community release called Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16, built on Alibaba's Qwen3.8-27B by the team AEON-7, removes the model's refusal behavior via abliteration and reports a mean KL div…

01:03
2026-08-16
changyi.fun
machine-learning

Attention Through Arithmetic Intensity

A technical analysis by an unnamed author, inspired by a post from Su Jianlin, derives the arithmetic intensity (FLOPs per byte) of attention during single-token decode, finding MHA has an AI of exact…

03:09
2026-08-12
sourcefeed.dev
artificial-intelligence

Mistral 3 Was Never About the Leaderboard

Mistral released the Mistral 3 family on December 2, 2025, under Apache 2.0, including nine dense Ministral 3 models (3B, 8B, 14B) and the sparse mixture-of-experts Mistral Large 3 with 675B total and…

13:30
2026-08-10
cast.ai
artificial-intelligence

LLM Inference Cost Optimization: Run AI Inference for Less

Cast AI benchmark testing shows that continuous batching at batch size 8 reduces Llama 3.1 70B inference cost on a single H100 from approximately $0.60-$0.80 per million tokens to $0.15-$0.25 per mill…

18:49
2026-07-16
kimi.com
artificial-intelligence

Kimi K3: Open Frontier Intelligence

Kimi released Kimi K3, a 2.8-trillion-parameter open model with native vision and a 1-million-token context window, claiming it is the world's first open 3T-class model. While trailing proprietary mod…

10:01
2026-06-25
discuss.huggingface.co
large-language-models

Deepseek? Qwen?

A single H200 GPU with 141GB HBM3e cannot comfortably run DeepSeek V4 Flash (284B total, 13B active parameters) due to VRAM constraints, even with 2TB system RAM for offloading. The model requires an …

01:13
2026-06-14
byteiota.com
large-language-models

DiffusionGemma: Google’s 4x Faster Text Diffusion Model

Google DeepMind released DiffusionGemma on June 10, 2026, a 26B open-weight text diffusion model that generates 256 tokens simultaneously, achieving up to 1,008 tokens per second on an H100—4-5x faste…

// co-occurs with top 8 entities
// topics top 6 topics