cd/entity/A100· home entities A100
grep -l @a100 /news/*.json | wc -l → 46

A100

mentions 46 type Organization page 1/3 feed RSS

// recent coverage 46 mentions

05:22
2026-08-14
dev.to
ai-infrastructure

NVIDIA GPU roadmap explained: from A100 to H200 and beyond

NVIDIA's data center GPU roadmap has advanced through four major architectures—Ampere, Hopper, Blackwell, and Vera Rubin—each delivering more memory, faster interconnects, and lower-precision compute …

21:32
2026-08-13
promptcube3.com
artificial-intelligence

Reproducing 2

A new analysis of 2,200 AI research reproducibility cases finds that most papers fail to replicate due to hardware variance, hyper-parameter sensitivity, and dependency issues, with the missing link o…

20:12
2026-08-12
byteiota.com
artificial-intelligence

Kubernetes 1.37: DRA Slices One GPU Into Many Pods

Kubernetes 1.37, releasing August 26, brings pod-level resources to stable, DRA device taints to GA, and advances GPU partitioning, addressing the 5% average GPU utilization measured by Cast AI across…

11:48
2026-08-05
gist.github.com
machine-learning

Cloud Training on RunPod: A Field Guide to the Edge Cases

AlphaPebble Labs engineers detailed a field guide for training AI models on RunPod's rented GPU infrastructure, highlighting edge cases such as the SSH gateway acting as a console rather than an exec …

20:57
2026-07-31
dev.to
machine-learning

Why INT4 Weight-Only Quantization Doesn't Speed Up Prefill

A developer's analysis shows that INT4 weight-only quantization speeds up decode but not prefill, because prefill is compute-bound while decode is memory-bound. The crossover point where a GEMM become…

10:40
2026-07-17
cast.ai
artificial-intelligence

Fractional GPUs and GPU Rightsizing: Stop Wasting Whole Cards

Average GPU utilization in production Kubernetes clusters is just 5%, according to the Cast AI 2026 State of Kubernetes Optimization Report, meaning 95 cents of every GPU dollar goes to idle silicon. …

10:09
2026-07-15
byteiota.com
large-language-models

NVIDIA Nemotron TwoTower: 2.42x Faster LLM Inference

NVIDIA released Nemotron-Labs-TwoTower on July 1, achieving 2.42 times faster inference throughput at 98.7% of the baseline model's benchmark quality by adding a second neural network tower trained on…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics