cd/entity/DFlash· home entities DFlash
grep -l @dflash /news/*.json | wc -l → 23

DFlash

mentions 23 type Organization page 1/2 feed RSS

// recent coverage 23 mentions

20:28
2026-08-11
piszczek.pl
artificial-intelligence

DFlash changes what tokens per second means

A configuration using the DFlash speculative decoding drafter with Meta Muse Glimmer 30B on an NVIDIA RTX PRO 4000 Blackwell SFF GPU reached 84.64 tokens per second on a coding task, a 4.50 times spee…

00:00
2026-08-10
runagentrun.co.uk
artificial-intelligence

Muse Glimmer lands as a 30B local agent

Meta Superintelligence Labs released Muse Glimmer, a 30-billion-parameter multimodal open model under Apache 2.0, designed for local agent workloads on consumer hardware. The model scores 76 on SWE-Be…

13:24
2026-07-10
arxiv.org
large-language-models

DominoTree

Researchers introduced DominoTree, a training-free best-first draft tree method for speculative decoding that uses Domino's conditional correction to achieve up to 6.6x speedup over autoregressive dec…

07:00
2026-07-07
dotnetperls.com
large-language-models

Notes on MTP, EAGLE-3 and DFlash

A developer reports that speculative decoding techniques MTP, EAGLE-3, and DFlash can significantly speed up local inference of large language models in llama-cpp. Testing on an NVidia 3060 RTX 12 GB …

07:00
2026-07-02
dotnetperls.com
large-language-models

DFlash for Local LLM Inference

Z-Lab's DFlash technique uses diffusion models to accelerate LLM token generation through speculative decoding, achieving up to 123 tokens per second for code generation in tests with Qwen 3 8B on lla…

20:16
2026-06-27
github.com
machine-learning

GitHub DeepSeek-AI/DeepSpec

DeepSeek-AI released DeepSpec, an open-source codebase for training and evaluating draft models for speculative decoding, supporting three draft model algorithms (DSpark, DFlash, Eagle3) and requiring…

13:27
2026-06-27
byteiota.com
large-language-models

DeepSeek DSpark Goes Live with 80% Inference Speed Gains

DeepSeek released DSpark, a speculative decoding framework now live in its DeepSeek-V4 Flash and Pro production API, delivering 51 to 400 percent throughput gains and up to 80 percent latency reductio…

09:28
2026-06-22
blog.doubleword.ai
large-language-models

Anatomy of a Diffusion Language Model

Diffusion Language Models (dLLMs) offer a faster alternative to autoregressive LLMs by generating multiple tokens in parallel, but face consistency challenges. Recent innovations like DFlash, Diffusio…

00:00
2026-06-22
fergusfinn.com
large-language-models

Adaptive speculative decoding: picking draft lengths at runtime

Researchers have developed adaptive speculative decoding, a method that dynamically selects draft lengths at runtime to optimize token generation efficiency in large language models. The approach addr…

00:17
2026-06-20
modal.com
large-language-models

Speculation Is All You Need

Modal Labs released state-of-the-art DFlash speculators for Qwen 3.5 and Qwen 3.6 models on Hugging Face, achieving 5-20% additional speedups and enabling Qwen 3.5 122B-A10B to run at over 1000 tok/s …

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics