03:36
2026-09-14
ianbarber.blog
large-language-models
Agents love prefill
DeepSeek's V4.1 Flash technical report introduces a "Causal Encoder–Decoder" architecture that runs only the first 20 of the model's 40 layers during prefill, cutting prefill compute in half for its 5…