cd/entity/NVIDIA B200· home entities NVIDIA B200
grep -l @nvidia b200 /news/*.json | wc -l → 21

NVIDIA B200

mentions 21 type Person page 1/2 feed RSS

// recent coverage 21 mentions

17:19
2026-09-13
github.com
ai-infrastructure

OpenGEMM: Open-source B200 GEMM kernels

OpenGEMM, an open-source project from developer aramesh10, released CUDA GEMM kernels for NVIDIA's B200 GPU (sm_100a), installable via pip and requiring PyTorch 2.8+ and CUDA 12.9+. The library emits …

18:27
2026-08-23
spheron.network
artificial-intelligence

Etched Sohu vs. Nvidia: Transformer ASIC vs. GPU (2026)

Etched AI, a chip startup founded in 2022, exited stealth on June 30, 2026, unveiling its transformer-only ASIC, Sohu, which the company claims delivers 500,000 tokens per second on an 8-chip server f…

04:00
2026-08-19
arxiv.org
artificial-intelligence

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

KernelArc, a multi-agent framework for autonomous GPU kernel optimization, achieved first-place rankings on representative L1, L2, Quantization, and FlashInfer tasks at the public SOL-ExecBench leader…

21:57
2026-08-16
promptcube3.com
artificial-intelligence

AI debt bubbles are going to force the Fed's hand again

A looming AI debt bubble, driven by massive infrastructure spending on GPU clusters and data centers, threatens systemic financial stability, according to analysis of the tech sector's capital expendi…

13:08
2026-08-15
sourcefeed.dev
artificial-intelligence

What a 232x AI Kernel Speedup Actually Proves

A solo developer with no professional GPU background placed 12th of 183 in GPU MODE's batched QR decomposition contest, beating the cuSolver-backed baseline by 232x (419,000 µs down to 1,805 µs on NVI…

02:51
2026-07-26
baseten.co
artificial-intelligence

We built the new fastest API for GLM-5.2

Baseten has built the fastest API for GLM-5.2, achieving peak speeds of 280 tokens per second and average speeds around 100 tokens per second, more than double the performance of the launch-day API as…

17:22
2026-07-23
pytorch.org
machine-learning

Helion on TPU: Towards Hardware Heterogeneous Kernel Authoring

Helion, PyTorch's high-level DSL for writing performance-portable ML kernels, partnered with Google to build a TPU backend that compiles Helion kernels to Pallas, enabling PyTorch-friendly TPU kernel …

18:00
2026-07-16
cline.ghost.io
large-language-models

How to Save Millions by Self-Hosting LLMs

Self-hosting open-weight large language models can save millions of dollars compared to using inference providers, according to a practical guide by Cline that analyzes the economics using Kimi K2.6 a…

06:33
2026-06-17
arxiv.org
machine-learning

Fearless Concurrency on the GPU

Researchers introduced cuTile Rust, a tile-based system for safe, idiomatic GPU kernel authoring in Rust that extends Rust's ownership discipline to GPU kernels. On the NVIDIA B200 GPU, cuTile Rust ac…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics