cd/entity/MTP· home entities MTP
grep -l @mtp /news/*.json | wc -l → 6

MTP

mentions 6 type Organization feed RSS

// recent coverage 6 mentions

18:44
2026-08-25
forum.level1techs.com
large-language-models

Self-hosting LLMs, my journey so far, and knowledge desired.

A user self-hosting large language models on an AMD Radeon RX 9700 reports achieving only ~22 tokens/s with Qwen 3.8 Q5_K_m, despite fitting 132k Q8_0 context, and seeks tips for improving performance…

07:00
2026-07-07
dotnetperls.com
large-language-models

Notes on MTP, EAGLE-3 and DFlash

A developer reports that speculative decoding techniques MTP, EAGLE-3, and DFlash can significantly speed up local inference of large language models in llama-cpp. Testing on an NVidia 3060 RTX 12 GB …

00:00
2026-05-31
cefboud.com
large-language-models

Exploring Speculative Decoding: From Concept to Implementation

Speculative decoding optimizes LLM inference by using a cheap draft model to predict multiple tokens, which are then verified in a single forward pass of the target model, reducing memory-bandwidth bo…

16:00
2026-05-27
dev.to
large-language-models

Why your quantized LLM loses its MTP heads and how to keep them

A developer discovered that standard quantization pipelines for large language models silently discard multi-token prediction (MTP) heads, causing speculative decoding speedups to vanish despite the b…

// co-occurs with top 8 entities
// topics top 6 topics