Agents love prefill DeepSeek's V4.1 Flash technical report introduces a "Causal Encoder–Decoder" architecture that runs only the first 20 of the model's 40 layers during prefill, cutting prefill compute in half for its 552B-parameter MoE model with 16B active parameters during decode and 8B active during prefill. The design, inspired by Microsoft's YOCO, has the first decoder layer project KV from the final encoder state and the other nineteen decoder layers reuse it, so prefill can stop halfway while still supplying everything the decoder needs. DeepSeek also combines Compressed Sparse Attention (CSA2) and FP4 KV storage to cut cached-input cost by four, charging cached input at 2% of the uncached-input price versus roughly 10% at most other labs. LLM inference has two stages: prefill, where the prompt is processed and the KV cache is built, and decode, where the model auto-regressively generates tokens. In a chat use case the two are somewhat close in size. The user writes a prompt, the model reasons about it then generates an answer, which is probably longer than the prompt. That is no longer where the FLOPs go. Everything is agentic now even the chats , so the loop looks more like: - you put in a query - the model generates a tool call - the tool call runs - the tool output is appended and prefilled, and the model decides what to do next Chat had this back and forth too, but prefix caching meant that you didn’t have to re-prefill what had already been generated. Tool outputs are lengthy, uncached, new content: file contents, terminal dumps, a fly’s connectome, etc. etc. This was clearly On The Mind of folks at DeepSeek. From their v4.1 Flash technical report https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek V41 Tech Report.pdf : “The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy.” Heavy enough that they made some interesting architectural changes. DSv4.1 Flash is a 552B parameter MoE with 16B params active… during decode. For prefill they just run the first half of the model, where only 8B are active 1 b8d59ac1-b8a1-4b1a-b843-7ceb3584852f They call this “Causal Encoder–Decoder”, inspired by Microsoft’s YOCO https://arxiv.org/abs/2405.05254 . For T5 fans, it isn’t an encoder-decoder in the old