cd/entity/H100· home› entities› H100
grep -l @h100 /news/*.json | wc -l → 225

H100

mentions 225 type Organization page 3/12 feed RSS

// recent coverage 225 mentions

00:00
2026-09-10
mindstudio.ai
ai-products

Nex-N2.5 Mini Hands-On: Testing Next AGI's Agentic Model

Next AGI released Nex-N2.5 Mini, a multimodal agentic model built on a Qwen3.5 mixture-of-experts backbone, under an Apache 2.0 license on Hugging Face, requiring two 80GB-class GPUs to run. In hands-…

00:00
2026-09-10
mindstudio.ai
ai-products

How to Run Nex-N2.5 Mini Locally on RunPod (Dual H100 Setup)

Nex AGI's Nex-N2.5 Mini, the smaller model in the company's N2.5 agentic family, requires two 80GB-class GPUs such as H100s and consumed roughly 66GB of VRAM per card in testing on RunPod, according t…

20:53
2026-09-09
newsletter.semianalysis.com
artificial-intelligence

Where Does a Robot Think – On-Device vs Datacenter Inference

Physical AI is moving AI beyond screens into robots, but the field is split on where inference should run: on-device versus in datacenters. Figure runs its Helix model entirely onboard, while Physical…

20:11
2026-09-09
byteiota.com
artificial-intelligence

Kubernetes 1.37 Gang Scheduling Beta: Fix GPU Training Deadlocks

Kubernetes 1.37 “Garhwal,” released August 26, 2026, includes gang scheduling in Beta, which prevents partial-job deadlocks that can leave idle GPUs costing around $467 per event. The feature, detaile…

00:00
2026-09-09
sailresearch.com
ai-infrastructure

Chasing Speed of Light on TPU v6e

In a blog post, the team at an AI infrastructure company detailed their optimization of Gemma 4 31B on Google's TPU v6e, boosting prefill efficiency from ~32% to ~63% MFU. They highlighted the chip's …

09:00
2026-09-08
cohere.com
artificial-intelligence

Inside the megakernel serving engine for North Mini Code

Cohere released a serving engine for its North Mini Code model built around a decode megakernel that runs 1.25x to 1.41x faster than vLLM end-to-end on a single H100 with BF16 precision. The engine, a…

00:46
2026-09-03
github.com
ai-infrastructure

Hot reload vLLM and sglang configs

A new open-source tool called trimtab lets operators change SGLang and vLLM scheduler settings live, cutting configuration changes from a 1-7 minute redeploy to about 15 ms, with zero dropped requests…

09:26
2026-08-31
martinalderson.com
artificial-intelligence

What GLM-5.3 Flash running on Chinese hardware means

Z.AI confirmed that its GLM-5.3 Flash model runs entirely on Chinese-manufactured hardware, though the company did not name the chipmaker or publish performance figures. The analysis assumes HiSilicon…

17:45
2026-08-29
promptcube3.com
ai-infrastructure

Nvidia is winning the AI race by fixing data center bottlenecks

Nvidia is winning the AI race by focusing on data center interconnect and networking technologies rather than just raw chip performance, according to an analysis. The company's NVLink, Mellanox-based …

23:48
2026-08-28
martinalderson.com
artificial-intelligence

What GLM-5.3 Flash running on Chinese hardware actually means

Z.AI confirmed that its GLM-5.3 Flash model runs entirely on Chinese-manufactured hardware, but the company did not name the chipmaker or publish performance figures. The analysis assumes HiSilicon pa…

17:08
2026-08-28
promptcube3.com
ai-infrastructure

Hardware lifecycles for AI chips are moving way faster than

A commentary argues that AI chip hardware lifecycles are longer than the hype suggests, with a three-year baseline for high-end silicon, citing deployment patterns and depreciation schedules. The piec…

01:15
2026-08-28
databricks.com
machine-learning

Fast, fault-tolerant PyTorch training on AI Runtime

Databricks' AI Runtime introduces APIs for fast, fault-tolerant PyTorch training, emphasizing that GPU failures are expected at scale and that data pipeline and checkpointing choices determine goodput…

04:00
2026-08-27
arxiv.org
large-language-models

DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

A new benchmark, DataKernelBench, evaluates whether large language models can optimize database queries on GPUs, achieving up to 2.11x speedup over torch.compile on TPC-H SF10 with an H100 GPU. The be…

17:09
2026-08-26
promptcube3.com
ai-infrastructure

Nvidia's massive $500B financing move raises serious questions

Nvidia's $500 billion financing platform, which lets AI startups and data center operators purchase H100s, B200s, and future Blackwell architectures through structured financing, raises serious concer…

← prev page 3 / 12 next →
// co-occurs with top 8 entities
// topics top 6 topics