cd/entity/DFlashΒ· homeβ€Ί entitiesβ€Ί DFlash
grep -l @dflash /news/*.json | wc -l β†’ 23

DFlash

mentions 23 type Organization page 2/2 feed RSS

// recent coverage 23 mentions

23:16
2026-06-12
letsdatascience.com
large-language-models

Xiaomi MiMo Hits 1,000 Tokens-Per-Second Inference

Xiaomi's MiMo-V2.5-Pro-UltraSpeed, a 1.02-trillion-parameter MoE model, achieved 1,000 tokens per second inference on standard cloud GPUs using FP4 quantization, DFlash speculative decoding, and TileR…

00:00
2026-06-08
fergusfinn.com
artificial-intelligence

The economics of speculative decoding

Speculative decoding, a lossless inference optimisation that predicts future tokens to reduce latency, faces new economic constraints as modern mixture-of-experts (MoE) architectures replace dense tra…

← prev page 2 / 2
// co-occurs with top 8 entities
// topics top 6 topics