cd/entity/H100· home entities H100
grep -l @h100 /news/*.json | wc -l → 149

H100

mentions 149 type Organization page 5/8 feed RSS

// recent coverage 149 mentions

14:10
2026-07-12
byteiota.com
ai-infrastructure

CoreWeave’s $35B Debt Loop and Your GPU Cloud Risk

Nvidia has engineered a circular financing loop with GPU cloud providers CoreWeave and Nebius, selling them GPUs, investing billions, and guaranteeing to buy unsold capacity, creating a $35 billion de…

04:54
2026-07-11
machinebrief.com
large-language-models

Accelerating SwiGLU: A breakthrough for Large Language Models

Researchers introduced two CUTLASS-based kernels that accelerate the SwiGLU activation function in large language models by up to 2.47x on NVIDIA H100 GPUs, shifting workloads from memory-bound to com…

13:42
2026-07-10
infoq.com
artificial-intelligence

Presentation: Chaos Engineering GPU Clusters

Bryan Oliver, a Thoughtworks Radar contributor and author, presented on chaos engineering for GPU clusters at a recent event, highlighting the complexity and scale of modern AI hardware such as NVIDIA…

14:12
2026-07-09
tokenstead.ai
artificial-intelligence

DeepSeek V4 Flash

DeepSeek released V4 Flash, a 284B-parameter MoE model with 13B active parameters per token, featuring FP4+FP8 hybrid attention and 1M native context. The open-weight model, available under MIT licens…

23:50
2026-07-08
letsdatascience.com
artificial-intelligence

Hugging Face Speeds Transformers Inference in vLLM

Hugging Face announced that its Transformers modeling backend for vLLM now matches or exceeds native vLLM speed for compatible architectures, reducing the need for separate hand-optimized serving port…

10:04
2026-07-08
lesswrong.com
artificial-intelligence

How slower does takeoff go with 10× less compute?

A new model estimates that a 10x reduction in R&D compute for an AGI company would slow AI progress by about 6x in the median case, with an 80% confidence interval of 3.5x to 8x. The model accounts fo…

02:11
2026-07-07
byteiota.com
artificial-intelligence

Cerebras + GPT-5.6 Sol: 750 tok/s Changes Your Agent Latency

Cerebras is running OpenAI's GPT-5.6 Sol at 750 tokens per second on its WSE-3 chip, a 10x throughput improvement over Nvidia H100 clusters. This drastically reduces agentic workflow latency, but acce…

14:30
2026-07-06
cast.ai
artificial-intelligence

The 7 Kubernetes Cost Drivers Most Teams Miss

Cast AI's 2026 analysis of tens of thousands of production clusters on AWS, GCP, and Azure found average CPU utilization at 8% fleet-wide, down from 10% the prior year. The company identified seven ke…

22:33
2026-07-03
letsdatascience.com
artificial-intelligence

Austria's MUSICA Deploys 1,088 NVIDIA H100 GPUs

Austria's MUSICA supercomputer, featuring 1,088 NVIDIA H100 GPUs delivering 45.11 petaflops, was inaugurated on July 3, 2026, marking an eightfold increase over previous systems. The EUR 45 million sy…

11:00
2026-07-02
science.slashdot.org
ai-infrastructure

The Space-Based Data Center Hype Machine Is Already In Orbit

IEEE Spectrum reports that orbital data centers remain economically and technically impractical despite Elon Musk's predictions, citing massive scaling challenges in launch capacity, satellite manufac…

12:00
2026-07-01
spectrum.ieee.org
artificial-intelligence

The Orbital Data Center Hype Machine Is Already in Orbit

SpaceX founder Elon Musk claimed at Davos that orbital data centers will be the lowest-cost AI option within two to three years, and SpaceX filed an FCC application for up to 1 million satellites. How…

09:07
2026-07-01
glukhov.org
large-language-models

Speculative Decoding: 20-50% Faster LLM Inference

Speculative decoding accelerates large language model inference by 20-50% without quality loss, using a draft-verify mechanism that generates multiple tokens per forward pass. The technique amortizes …

← prev page 5 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics