04:51
2026-08-10
newsletter.semianalysis.com
artificial-intelligence
Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceX
TileRT's persistent engine on NVIDIA GPUs achieves up to 500 tokens/s/user on the InferenceX GLM5 FP8 744B benchmark on a single B200 decode server, approximately 3Γ faster than GB300 NVL72 running trβ¦