cd/entity/H100· home› entities› H100
grep -l @h100 /news/*.json | wc -l → 225

H100

mentions 225 type Organization page 2/12 feed RSS

// recent coverage 225 mentions

18:33
2026-09-21
dev.to
large-language-models

High-Throughput LLM Inference & Training: A Deep Dive into vLLM

An engineer at g factor detailed how vLLM's PagedAttention and continuous iteration-level batching solve the memory-bandwidth bottleneck in production LLM inference, drawing on benchmarks run on dedic…

00:00
2026-09-18
digitalapplied.com
ai-research

Small AI Models Trained for One Job: When They Win

Engineer Rohan Bansal reported on September 16, 2026 that a 4-billion-parameter model trained on roughly 500 frontier-model example runs plus reinforcement learning against four Postgres database cont…

00:00
2026-09-17
mindstudio.ai
artificial-intelligence

Agnes-3.0-Flash Preview: Specs and Benchmarks of the Open Model

Agnes AI released Agnes-3.0-Flash Preview, a 33-billion-parameter open-weight multimodal language model under an Apache 2.0 license with a 262,144-token context window and a hybrid attention architect…

18:59
2026-09-16
aws.amazon.com
ai-infrastructure

Fault tolerant distributed training on Amazon EKS using NVRx

NVIDIA's Resiliency Extension (NVRx) can be integrated into PyTorch Fully Sharded Data Parallel training on Amazon Elastic Kubernetes Service to cut idle GPU time, with synchronous checkpointing consu…

18:01
2026-09-15
blog.neurometric.ai
artificial-intelligence

Compounding Inference Is As Powerful As Compounding Interest

Shopify's ML team fine-tuned a 0.8B-parameter model for buyer-profile generation that scored 84.6 against GPT-5.6-sol's 83.0 on the company's judge, according to a chart CEO Tobi Lütke posted on Septe…

00:00
2026-09-15
mindstudio.ai
large-language-models

Run GLM 5.3 Flash Locally: GSQ and RCO Quantization Explained

An independent research group in Austria has developed two quantization techniques, GSQ and RCO, that compress Z.ai's 320-billion-parameter GLM 5.3 Flash vision-language model from roughly 320GB at fu…

16:46
2026-09-14
promptcube3.com
machine-learning

NVIDIA Transformer Engine makes JAX MoE training actually fast

NVIDIA Transformer Engine is now integrated with JAX to accelerate dropless Mixture of Experts (MoE) training, targeting the conditional computation bottleneck in architectures used by DeepSeek, Qwen,…

21:53
2026-09-13
discuss.huggingface.co
ai-infrastructure

Alternative GPU Compute Infrastructure Model for Large Pipelines

ApexSovereign Infrastructure Protocol launched ApexSovereign, a bare-metal GPU compute service it says cuts model fine-tuning and inference overhead by up to 40% by scanning global spot-market pricing…

15:34
2026-09-13
github.com
ai-infrastructure

Show HN: OSS Modal Alternative

Beam released Beta9, an open-source runtime for serverless AI workloads that it positions as an alternative to Modal, offering container cold starts in under a second and fan-out to hundreds of contai…

00:00
2026-09-13
mindstudio.ai
ai-products

How to Self-Host Nex-N2.5 with SGLang and Docker

Nex-AGI released its Nex-N2.5 family of open-weight agentic models in three sizes — mini, Pro, and Max — with a prebuilt Docker image running a customized SGLang fork (nexagi/sglang:v0.5.18-nex-patch)…

← prev page 2 / 12 next →
// co-occurs with top 8 entities
// topics top 6 topics