cd/entity/TurboPrefill· home entities TurboPrefill
grep -l @turboprefill /news/*.json | wc -l → 3

TurboPrefill

mentions 3 type Organization feed RSS

// recent coverage 3 mentions

07:32
2026-07-25
devpost.com
large-language-models

TurboPrefill: 3.27× Prefill Speedup in Llama.cpp

A new scheduling technique called TurboPrefill achieves up to 3.27× prefill speedup in llama.cpp by pipelining prompt microbatches across multiple GPUs, reducing inter-GPU communication bottlenecks. D…

22:09
2026-06-20
github.com
artificial-intelligence

Show HN: VLMs Can Respond Twice as Fast Without Losing Quality

A new scheduling technique called TurboPrefill reduces waiting time for Vision Language Models by nearly half, from 9.0 to 4.6 seconds, without changing model weights or architecture. The optimization…

// co-occurs with top 8 entities
// topics top 6 topics