cd/entity/DFlash· home› entities› DFlash
grep -l @dflash /news/*.json | wc -l → 28

DFlash

mentions 28 type Organization page 1/2 feed RSS

// recent coverage 28 mentions

18:09
2026-10-01
dev.to
ai-infrastructure

Evaluating Speculative Decoding in vLLM on AMD MI300X GPUs

A developer evaluated speculative decoding in vLLM on AMD Instinct MI300X and MI355X GPUs, testing draft methods including MTP, EAGLE-3, DFlash, and DSpark that propose candidate tokens verified in a …

09:26
2026-09-07
vllm.ai
large-language-models

Speculative Decoding in vLLM on AMD GPUs

AMD's experiments with speculative decoding in vLLM on AMD Instinct MI300X and MI355X GPUs using the ROCm platform show that output-token throughput gains vary by drafting method, proposal length, mod…

20:28
2026-08-11
piszczek.pl
artificial-intelligence

DFlash changes what tokens per second means

A configuration using the DFlash speculative decoding drafter with Meta Muse Glimmer 30B on an NVIDIA RTX PRO 4000 Blackwell SFF GPU reached 84.64 tokens per second on a coding task, a 4.50 times spee…

00:00
2026-08-10
runagentrun.co.uk
artificial-intelligence

Muse Glimmer lands as a 30B local agent

Meta Superintelligence Labs released Muse Glimmer, a 30-billion-parameter multimodal open model under Apache 2.0, designed for local agent workloads on consumer hardware. The model scores 76 on SWE-Be…

13:24
2026-07-10
arxiv.org
large-language-models

DominoTree

Researchers introduced DominoTree, a training-free best-first draft tree method for speculative decoding that uses Domino's conditional correction to achieve up to 6.6x speedup over autoregressive dec…

07:00
2026-07-07
dotnetperls.com
large-language-models

Notes on MTP, EAGLE-3 and DFlash

A developer reports that speculative decoding techniques MTP, EAGLE-3, and DFlash can significantly speed up local inference of large language models in llama-cpp. Testing on an NVidia 3060 RTX 12 GB …

07:00
2026-07-02
dotnetperls.com
large-language-models

DFlash for Local LLM Inference

Z-Lab's DFlash technique uses diffusion models to accelerate LLM token generation through speculative decoding, achieving up to 123 tokens per second for code generation in tests with Qwen 3 8B on lla…

20:16
2026-06-27
github.com
machine-learning

GitHub DeepSeek-AI/DeepSpec

DeepSeek-AI released DeepSpec, an open-source codebase for training and evaluating draft models for speculative decoding, supporting three draft model algorithms (DSpark, DFlash, Eagle3) and requiring…

13:27
2026-06-27
byteiota.com
large-language-models

DeepSeek DSpark Goes Live with 80% Inference Speed Gains

DeepSeek released DSpark, a speculative decoding framework now live in its DeepSeek-V4 Flash and Pro production API, delivering 51 to 400 percent throughput gains and up to 80 percent latency reductio…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics