{"slug": "ainews-deepseek-v4-1-flash-763b-p8b-d16b-novel-causal-encoder-decoder-with-marks", "title": "[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale", "summary": "DeepSeek released DeepSeek v4.1-Flash, a 763B-parameter model with a novel causal encoder-decoder architecture that splits 8B parameters to prefill and 16B to decode, per the model's tech report on Hugging Face. The release retires DeepSeek V4 Pro and adds vision without a separate model, with DeepSeek claiming a KV cache footprint up to 1/8 that of V4 Flash through Sliding-Window Attention Bounded Replay. DeepSeek's prior intermediate releases, including Math, Coder, and R1, preceded the v2, v3, and v4 generations.", "body_md": "*We are late to this but better than never. Have been busy finalizing the second [AIE NYC](https://ai.engineer/nyc/2026), which is happening in one month. [Get your tix](https://ai.engineer/nyc/2026#tickets) before prices go up - we will announce speakers from Bridgewater, Ramp, Coatue, Mastercard, Vanguard, Coinbase, Blackrock, Fidelity, Point72, Capital One, JPMC, Wells Fargo, Bloomberg, A24 (yes the movie studio) Labs, Two Sigma, Apollo Global, and more next week!*\n\n**The way DeepSeek pursues their research agenda is nothing short of fascinating.** In between major DeepSeek versions, from v2 to v3 to v4, they have released intermediate papers with a hyperfocused architectural improvement and basically a 100% hit rate, from **[Math](https://arxiv.org/abs/2402.03300)** (esp [GRPO](https://www.interconnects.ai/p/papers-im-reading-base-model-rl-grpo))**, [Coder](https://arxiv.org/abs/2401.14196),** and **[R1](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Llama-70B),** not to mention more [recent work on Manifold Constrained Hyperconnections and Compressed Sparse Attention](https://www.latent.space/p/ainews-deepseek-v4-pro-16t-a49b-and?utm_source=publication-search). After the enormous attention in 1H2025 from the R1 paper, DeepSeek started laying low, and for about the past year, was happy to let peers like GLM and Kimi take the lead on Open Models. \n\nIt looked dicey for a little bit, but [true whalebros](https://x.com/teortaxesTex/status/2097927946769948717) never wavered, and now DeepSeek are sending a weirdly mixed message by doing a completely new architecture, [retiring V4 Pro](https://www.reddit.com/r/LocalLLaMA/comments/1wbfrut/deepseek_has_soft_retired_deepseek_v4_pro/) and going all in on this new model, and yet only titling it v4.1 Flash, it seems to be a test of whether or not you know how to read through the basic headlines to understand true advances.\n\nYes, [v4.1 Flash](https://artificialanalysis.ai/models/open-source?lab=alibaba%2Cdeepseek%2Cnvidia%2Cmeta%2Cgoogle%2Cmistral%2Cazure%2Czai) is technically behind other open models in some benchmarks. But that’s because we don’t yet have benchmarks that concisely capture what v4.1, and the broader research agenda of DeepSeek, is aiming for - the most creative and efficient use of context we have ever seen openly explained.\n\nIf you are the sort to only read model versions and benchmark headlines, you are exactly the type of superficial person that DeepSeek is looking to fool. The best way to understand DeepSeek’s enormous advance here is to look at Sebastian’s meme:\n\nSame model name, but hardly a 0.1 bump by anyone’s standards, and they even threw in vision without [making you wait for a separate model](https://api-docs.deepseek.com/news/news260821/). For a better visualization you can look at all the model innovations stacked up over time from the OG encoder-decoder architecture from *Attention is All You Need:*\n\nIf you read [our V4 Pro writeup](https://www.latent.space/p/ainews-deepseek-v4-pro-16t-a49b-and?utm_source=publication-search) and [Engram](https://github.com/deepseek-ai/Engram/blob/main/Engram_paper.pdf) you should be up to date on the basic architectural reading for DeepSeek as of April 2026, but what we are HUGE fans of is the prefill/decode separation introduced here, 8B in prefill (input tokens), 16B in decode (output tokens), causing our alphabet soup of “DeepSeek v4.1-Flash: 763B-P8B-D16B” if you extend the established notation for MoEs. That’s a sparsity of 1-2%, and if you read [the DeepSeek v4.1 Flash tech report](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf), combined with new tweaks like Sliding-Window Attention Bounded Replay, makes for a KV cache footprint up to 1/8 that of V4 Flash… which make it much better/faster/cheaper for long running agents:\n\nWe are so glad that DeepSeek is back publishing SOTA research. Our last highlight is their comments on post-training, where they largely seem to [agree with Prof Jie Tang](https://www.latent.space/p/ainews-death-of-params-zai-ceo-jie):\n\nAI News for 9/9/2026-9/10/2026. We checked 12 subreddits, [544 Twitters](https://twitter.com/i/lists/1585430245762441216) and no further Discords. [AINews’ website](https://news.smol.ai/) lets you search all past issues. As a reminder, [AINews is now a section of Latent Space](https://www.latent.space/p/2026). You can [opt in/out](https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack) of email frequencies!\n\n# **AI Twitter Recap**\n\n**DeepSeek launched V4.1-Flash as a new open-weight flagship focused on extreme inference efficiency and low cost.**\n\n- Independent benchmark account Artificial Analysis reported that DeepSeek V4.1 Flash surpasses DeepSeek V4 Pro 0813 despite being much cheaper, scoring **40 on the Artificial Analysis Intelligence Index** , just below GLM-5.3-Flash and above the latest V4 Pro, while being priced at**$0.30 / 1M input tokens** and**$1.20 / 1M output tokens** with**cached input at $0.006 / 1M** and an additional**50% off-peak discount** ; they also describe it as a**763B total-parameter** model with**8B active input** and**16B active output** parameters,**1M-token context** , text+image input,**MIT license** , and US/API availability via DeepSeek first party[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422) ,[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148681962913915) ,[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148684185972758)\n- Vals called it the new **#1 open-weight model on the Vals Index** , ahead of Kimi K3, at just**$0.30 per test** , the cheapest model in the open-weight top 10; they also note the eval ran with**1M context** ,**384 max output tokens** ,**temperature 1** , default top-p/top-k, and**high reasoning effort**[@ValsAI](https://x.com/ValsAI/status/2098125164072554545) ,[@ValsAI](https://x.com/ValsAI/status/2098125177116848591) ,[@ValsAI](https://x.com/ValsAI/status/2098125179092431297)\n- Baseten shipped day-0 support and summarized the product positioning as **smarter, faster, and more efficient than DeepSeek v4 Pro 0813** , with**text and vision** ,**US-only** ,**ZDR** , and**1M context**[@baseten](https://x.com/baseten/status/2098169972874994071)\n- Ollama began rolling it out to **Max and Team** accounts, later expanding to**Pro plan subscribers**[@ollama](https://x.com/ollama/status/2098188014119985406) ,[@ollama](https://x.com/ollama/status/2098188470305128692) ,[@ollama](https://x.com/ollama/status/2098235674793242770)\n\n## **Architecture and paper-level technical details**\n\n**The most discussed technical novelty is a causal encoder-decoder design aimed at lowering active compute and KV/cache costs.**\n\n- Artificial Analysis says the model uses a **new causal Encoder–Decoder architecture** , with**8B active parameters for input/prefill** and**16B active parameters for output/decode**[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422)\n- Sebastian Raschka characterized V4.1 as a **“big overhaul”** and said they “should have called it DeepSeek V5,” explicitly highlighting the**encoder-decoder setup** as the key break from prior DeepSeek generations[@rasbt](https://x.com/rasbt/status/2098142625819672603)\n- Multiple technical readers reacted to the design as unusually hybrid: one called it “a very interesting mix of very conservative and sometimes old ideas in research and potentially cutting edge efficiency and hardware design in engineering” [@_xjdr](https://x.com/_xjdr/status/2098106496282448013)\n- A concise architecture read from Stochastic Chasm compared the design philosophy to **HySparse, NSA, and DeepSeek’s own CSA/HCA from V4** , summarizing it as a**local sliding-window branch plus sparse retrieval branch** , suggesting this sparse/local hybrid is becoming a broader pattern[@stochasticchasm](https://x.com/stochasticchasm/status/2098102323268767832)\n- The same account noted multimodal changes were **not radical** , saying DeepSeek mostly “lets the backbone handle most of it and give it visual tokens,” with**3x3 pixel unshuffle** instead of the more common**2x2**[@stochasticchasm](https://x.com/stochasticchasm/status/2098116030627455450)\n- They later flagged a “big difference from K3 on vision encoders,” implying the vision front-end diverges materially from recent Chinese peers [@stochasticchasm](https://x.com/stochasticchasm/status/2098165237400953054)\n- TeortaxesTex observed a recurring DeepSeek pattern of doing something unusual in the **first N layers** —previously dense or hash-routed, now**SWA-only** —speculating this may reflect repeated training difficulties in early layers[@teortaxesTex](https://x.com/teortaxesTex/status/2098132297253896451)\n- Later, the same account argued the stack is “down to **40 layers** , arguably only**20 legit decoder layers** ,” underscoring just how aggressively DeepSeek may be compressing effective depth in decode-critical paths[@teortaxesTex](https://x.com/teortaxesTex/status/2098176524612510102)\n- Another thread fragment from TeortaxesTex suggested DeepSeek is doing **multiple compression frequencies** , “it’s just all CSA2,” in response to architectural discussion around memory compression[@teortaxesTex](https://x.com/teortaxesTex/status/2098131613678707129)\n- Nrehiew’s technical notes emphasize **KV cache compression** as central to the design, calling it a case study in “how obsessing over KV Cache compression gets you a hyper-efficient frontier model”[@nrehiew_](https://x.com/nrehiew_/status/2098170409686647263)\n- In a follow-up, nrehiew highlighted infrastructure specifics from the report: **dispatch strategy to reduce long-tail stalls** ,**router replay from previous checkpoints** , management of shorter-completion off-policy effects via**dataset-level capping** ,**discard schemes** ,**bounded off-policy ratio and loss masking** , and**persistent KVs and routers** when a new checkpoint is updated; they also mention a final stage with**full-vocab OPD on 40+ teacher models**[@nrehiew_](https://x.com/nrehiew_/status/2098170443660402942)\n- Nrehiew concluded that the design looks cleaner than the older **HSA + CSA** combination in V4, saying it was “very clearly designed for inference,” and cited a striking**~890 bytes/token KV size** for the benchmarked score regime[@nrehiew_](https://x.com/nrehiew_/status/2098170450526543892)\n- Stochastic Chasm inferred **QAT for the KV cache** , saying this would explain why the model performs better than peers under**FP4 KV cache**[@stochasticchasm](https://x.com/stochasticchasm/status/2098154481750020375)\n\n## **Benchmark results and numbers**\n\n**Independent evals consistently paint V4.1-Flash as unusually strong on cost-adjusted intelligence, long context, and automation, with a major caveat around verbosity.**\n\n- Artificial Analysis’ headline: **40 AA Index** , above V4 Pro and below GLM-5.3-Flash[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422) , corroborated separately by Scaling01[@scaling01](https://x.com/scaling01/status/2098136324603547907)\n- Artificial Analysis reported **AutomationBench-AA: 69%** , tying**GPT-6 Astra (69%)** and above**Grok 4.6 (67%)** , while improving**15 points** over V4 Flash 0731 and sitting**12 points above V4 Pro 0813 (57%)** and**7 points above GLM-5.3 (62%)**[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422)\n- On GDPval-AA v2 it reportedly gains **164 Elo** , from**1468 to 1632** , overtaking**Kimi K3 at 1584**[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422)\n- On **AA-LCR v1.1** it scores**84%** , on par with**GPT-5.6 Sol** and**Gemini 3.8 Flash** at**84%**[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422)\n- Artificial Analysis also says V4.1 Flash is among the **most verbose models measured** , averaging**89k tokens per Intelligence Index task** —**25% more** than GLM-5.3 (71k),**29% more** than GLM-5.3-Flash (69k),**62% more** than V4 Pro 0813 (55k), and even above**Fable 5.1 (78k)** and**Claude Opus 5 (73k)**[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422)\n- Even with that verbosity, AA estimates just **$0.27 per Intelligence Index task** , roughly**7x below GLM-5.3 ($2.01)** and**Kimi K3 ($2.00)** , and**~2.5x below V4 Pro 0813 ($0.67)**[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422)\n- Vals’ result reinforces cost leadership: **$0.30/test** , #1 open-weight on their board[@ValsAI](https://x.com/ValsAI/status/2098125164072554545)\n- A separate reaction thread summarized DeepSWE-style claims more aggressively, saying V4.1 Flash offered **better performance than GPT-5.6 Sol and Opus 5 in DeepSWE at 94% lower API costs** , but that statement is secondhand summary rather than a primary benchmark post in this dataset[@kimmonismus](https://x.com/kimmonismus/status/2098107083665060275)\n\n## **Running it locally and inference engineering reactions**\n\n**A large fraction of discussion centered on the surprising ease of running V4.1-Flash on commodity-ish local hardware through offload and SSD streaming.**\n\n- Fraser Price reported **full-precision DeepSeek 4.1 Flash + DSpark at 200 TPS on 4 Max-Qs with just 64GB system RAM** , offloading a**200GB Engram/hash table to NVMe** ; he says this made keeping the full structure in RAM unnecessary and promised a**vLLM recipe**[@fraserpricee](https://x.com/fraserpricee/status/2098078317723242813)\n- He later improved that to **300+ TPS on 4 RTX Pros** , still at**full precision** , with**<32GB peak system RAM** , using a**custom vLLM fork** and SSD support[@fraserpricee](https://x.com/fraserpricee/status/2098183796080173382)\n- Antirez showed **DwarfStar running V4.1 Flash on a 128GB M5 Max** , saying SSD streaming made it unexpectedly fast; he speculated both recent SSD-streaming changes and the possibility that DS4.1 “uses the same experts more” contributed[@antirez](https://x.com/antirez/status/2098121665771110540)\n- TeortaxesTex reacted that it is “incredible you can run frontier models mostly off SSD” [@teortaxesTex](https://x.com/teortaxesTex/status/2098128365970440432)\n- Elie Bakouch posted a reaction meme explicitly about the **inference engineer view** of the V4.1 Flash architecture, reflecting how strongly the launch resonated with systems folks[@eliebakouch](https://x.com/eliebakouch/status/2098223948127183261)\n- vLLM’s new release also included **DeepSeek-V4 shared experts fused into MegaMoE** , plus**Mooncake Store can offload decode KV** , relevant context for why serving this class of model is rapidly becoming easier in open infra[@vllm_project](https://x.com/vllm_project/status/2098214992755765758) ,[@vllm_project](https://x.com/vllm_project/status/2098214998426444009)\n\n## **Facts vs. opinions**\n\n**Facts and directly attributed claims**\n\n- V4.1 Flash launched and was quickly supported by Ollama and Baseten [@ollama](https://x.com/ollama/status/2098188014119985406) ,[@baseten](https://x.com/baseten/status/2098169972874994071)\n- Independent benchmarks reported **AA Index 40** ,**AutomationBench-AA 69%** ,**AA-LCR 84%** ,**GDPval-AA v2 1632 Elo** ,**1M context** ,**MIT license** , and low API pricing[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422)\n- Vals reported #1 among open-weight models on its index, at **$0.30/test** , with**384 max output tokens** under its harness settings[@ValsAI](https://x.com/ValsAI/status/2098125164072554545) ,[@ValsAI](https://x.com/ValsAI/status/2098125177116848591)\n- Local deployment reports claimed **200 TPS** and later**300+ TPS** on 4-GPU setups, plus successful M5 Max SSD-streamed operation[@fraserpricee](https://x.com/fraserpricee/status/2098078317723242813) ,[@fraserpricee](https://x.com/fraserpricee/status/2098183796080173382) ,[@antirez](https://x.com/antirez/status/2098121665771110540)\n\n**Interpretations and opinions**\n\n- Raschka’s “they should have called it V5” is an opinion about how substantial the architectural change is [@rasbt](https://x.com/rasbt/status/2098142625819672603)\n- TeortaxesTex’s speculation that DeepSeek “repeatedly struggled to train first layers properly” is inference, not a confirmed statement from DeepSeek [@teortaxesTex](https://x.com/teortaxesTex/status/2098132297253896451)\n- Nrehiew’s framing that the report is “cleaner” than the prior HSA/CSA design and likely unlike what OpenAI/Anthropic would do because of their custom chips is informed opinion [@nrehiew_](https://x.com/nrehiew_/status/2098170450526543892)\n- The “DeepSeek ships internal research artifacts and not products” critique is an external judgment, not a factual release note [@teortaxesTex](https://x.com/teortaxesTex/status/2098213577546985945)\n- Assertions that “data is all that matters” or “research is over” were themselves criticized as overreactions [@shikibmehri](https://x.com/shikibmehri/status/2098233059242099175)\n\n## **Different opinions and reactions**\n\n**Supportive / impressed**\n\n- Strong positive reactions came from benchmarkers and researchers emphasizing the price/perf step: Vals’ “new #1 open-weight model,” Artificial Analysis’ cost-adjusted headline, and general praise like “interesting release / breath of fresh air vibe” [@ValsAI](https://x.com/ValsAI/status/2098125164072554545) ,[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422) ,[@dejavucoder](https://x.com/dejavucoder/status/2098128229408375093)\n- Raschka called it “super cool and refreshing” [@rasbt](https://x.com/rasbt/status/2098142625819672603)\n- XJDR liked the engineering thinking despite some aesthetic reservations [@_xjdr](https://x.com/_xjdr/status/2098106496282448013)\n- Nrehiew called it “yet another banger tech report” [@nrehiew_](https://x.com/nrehiew_/status/2098170450526543892)\n- Stochastic Chasm ended by saying the paper was “dense” but appreciated the multi-agent training angle and sparse design ideas [@stochasticchasm](https://x.com/stochasticchasm/status/2098186711578943775) ,[@stochasticchasm](https://x.com/stochasticchasm/status/2098186892579860662)\n\n**Neutral / analytical**\n\n- Some observers mainly dissected the design rather than cheering it: sparse/local hybridization, first-layer oddities, multimodal tokenization, KV quantization, colocated async RL, etc. [@stochasticchasm](https://x.com/stochasticchasm/status/2098102323268767832) ,[@stochasticchasm](https://x.com/stochasticchasm/status/2098185722561966230) ,[@nrehiew_](https://x.com/nrehiew_/status/2098170443660402942)\n- Gordic Aleksa used the paper as evidence in a broader pretraining-data taxonomy, placing DeepSeek in the **organic data camp** and noting surprise that, based on publications, they do not appear to use even synthetic**rephrasing**[@gordic_aleksa](https://x.com/gordic_aleksa/status/2098108613676212598)\n\n**Critical / skeptical**\n\n- TeortaxesTex repeatedly pushed back on external impressions, arguing DeepSeek often shows **high internal evals, weaker external robustness, brittleness, and weird skill gaps** , because it “ships internal research artifacts and not products”[@teortaxesTex](https://x.com/teortaxesTex/status/2098213577546985945)\n- The same account called some eval results “very strange,” particularly **AutomationBench #1** and a CritPt regression, and asked the DeepSeek team to “meditate on this”[@teortaxesTex](https://x.com/teortaxesTex/status/2098157751465603171)\n- They also argued that **V4 GA** had benefited massively from tool/skills harness access, whereas**V4.1** appears less dependent on harness scaffolding and better in “minimal harnesses”[@teortaxesTex](https://x.com/teortaxesTex/status/2098129561481363901)\n- In hands-on use, they reported that **multi-agent “DSH agent teams”** could degrade quality unless the project has very clear modularity, with**V4.1 solo** outperforming team mode in at least one example because subagents produced slop or wasted tokens on unnecessary research[@teortaxesTex](https://x.com/teortaxesTex/status/2098154067948134492) ,[@teortaxesTex](https://x.com/teortaxesTex/status/2098202210228109478)\n- Jared Z’s broader product-market critique—that users now care deeply about token cost, and daily-driver coding models should be both cheap and smart—fits V4.1 Flash’s positioning even though it wasn’t about the model specifically [@imjaredz](https://x.com/imjaredz/status/2098135420035035603)\n\n## **Context**\n\n**Why this matters technically and strategically**\n\n- The launch lands amid a broader shift from “bigger dense chat models” toward **systems-optimized, sparse, long-context, agent-oriented models** that can actually be served cheaply and locally.\n- V4.1 Flash’s positioning is unusually aggressive: open-weight, MIT-licensed, 1M context, multimodal input, low active parameter counts, extreme cache discounts, and demonstrated viability on SSD/offload-heavy consumerish setups [@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422) ,[@fraserpricee](https://x.com/fraserpricee/status/2098078317723242813) ,[@antirez](https://x.com/antirez/status/2098121665771110540)\n- The benchmark pattern suggests a meaningful trade: **very high verbosity** but still**exceptionally low total task cost** thanks to ultra-cheap token pricing[@ArtificialAnlys](https://x.com/ArtificialAnlys/status/2098148674203488422)\n- The architecture also reflects a broader industry trend toward **splitting prefill and decode economics** , making long-context and agentic workloads more practical without paying frontier dense-model costs on every token.\n- The release reinforces the idea that open models are increasingly competitive not just on raw weights availability, but on **servability** —the ability to fit into offload pipelines, quantized KV stacks, local deployment, and open inference servers.\n- It also sharpened debate over what matters most in 2026 model progress: architecture, RL/inference co-design, data quality, or systems work. Shikib Mehri explicitly pushed back on the claim that DeepSeek’s paper means “research is over,” arguing instead that the lever surface has expanded from architecture into data-factory and reward-design research [@shikibmehri](https://x.com/shikibmehri/status/2098233059242099175)\n- Finally, DeepSeek remains a polarizing lab identity-wise: admired for shipping unusual research artifacts and detailed reports, but also seen by some practitioners as less polished than product-centric competitors, with odd eval gaps and brittle behaviors that appear more clearly in real workflows than in internal headline numbers [@teortaxesTex](https://x.com/teortaxesTex/status/2098213577546985945) ,[@teortaxesTex](https://x.com/teortaxesTex/status/2098157751465603171)\n\n**OpenAI’s Voice, Agents, and Enterprise Push**\n\n- **OpenAI launched GPT-Live-1 into the API and quickly seeded an ecosystem around it** : the new model is positioned as a**full-duplex** voice interface that can**listen while speaking** and delegate tool use or reasoning to a backend model. The core launch came from[@OpenAIDevs](https://x.com/OpenAIDevs/status/2098099269551149398) , with additional detail that developers can control**tone, pacing, expressiveness, response length, and language**[here](https://x.com/OpenAIDevs/status/2098099427357724870) . OpenAI’s own benchmark post claimed improvements over GPT-Realtime-2.1, including**83.6% first-attempt task completion on Tau3** when paired with**GPT-6 Astra** ,**97.3% on Artificial Analysis Conversational Dynamics** , and**0.798s response onset latency** on Full Duplex Bench v1[details](https://x.com/OpenAIDevs/status/2098118242548281588) .\n- **The surrounding toolchain is maturing toward hosted agent infra** : OpenAI also announced a public-beta**Agents API** with the**Codex harness** , plus**OpenAI-hosted sandboxes** for code execution, files, and artifacts via managed cloud agents[launch](https://x.com/OpenAIDevs/status/2098130570048045453) . This aligns with a broader industry move to collapse model, runtime, and sandbox into one surface. Integration announcements from[LiveKit](https://x.com/livekit/status/2098126102052905001) ,[HeyGen](https://x.com/HeyGen/status/2098108031276134776) ,[Telnyx](https://x.com/telnyx/status/2098098605601042943) ,[Speak](https://x.com/speak/status/2098095986606551481) , and[Cognition’s Devin Voice](https://x.com/cognition/status/2098142686486356185) suggest GPT-Live-1 may become a default substrate for production voice agents faster than the earlier realtime stack did.\n- **Enterprise data access is becoming a first-class product primitive** : OpenAI’s product-side announcement of a**Data agent in ChatGPT Work** promises dashboards, answers, and actions over connected company data sources[@ChatGPT](https://x.com/ChatGPT/status/2098065296968011853) , while[Box](https://x.com/Box/status/2098127482088267799) framed its integration as “the file system for AI” bringing governed enterprise context into ChatGPT. Combined with Google’s docs-for-agents push and Cursor’s new persistent workspaces, the trend is toward**stateful, organization-aware agent environments** , not stateless model endpoints.\n\n**Cognition, Cursor, and the Shift Toward Persistent Coding Agents**\n\n- **Cognition had a notably strong day** : it released**SWE-2** , described as “our closest model yet to the frontier,” claiming parity on leading coding evals at up to**70% lower cost** and explicitly stating it**scaled RL to multiple trillions of parameters**[launch](https://x.com/cognition/status/2098069235733823965) . Additional context from[ybenpan](https://x.com/ybenpan/status/2098077716146958723) emphasized that the team built**algorithm, infra, and data in-house** , while[silasalberti](https://x.com/silasalberti/status/2098115298125897961) highlighted a practical RL finding: a**simple linear length penalty** preserved a training-time Pareto curve shape across effort levels.\n- **The Devin stack is becoming more multimodal and more integrated with developer workflows** : beyond SWE-2, Cognition launched**Devin Voice** powered by**GPT-Live and SWE-2**[tweet](https://x.com/cognition/status/2098142686486356185) , and announced that** Dioxus Labs** is joining Cognition to contribute to**Devin’s VM, computer use, and testing** while continuing support for Dioxus and related Rust OSS[Cognition](https://x.com/cognition/status/2098109121169883237) . This is a concrete example of coding-agent vendors acquiring infra and systems talent, not just model researchers.\n- **Cursor’s new “Projects” feature points to the same destination from the IDE side** :[Cursor](https://x.com/cursor_ai/status/2098162488013455784) introduced**persistent threads with a coordinator agent** , shared memory/artifacts across agents, and sync across user devices and agent computers. In practical terms, this is a move away from “one chat per task” toward a**long-lived software project substrate** where subagents accumulate state over time. Read together with Claude Code’s new[pane pop-outs](https://x.com/ClaudeDevs/status/2098090911137972271) and[managed-agent session viewer / auto mode](https://x.com/ClaudeDevs/status/2098120133549895978) , the market is converging on the idea that coding agents need**persistent context, inspectable sessions, and explicit orchestration controls** , not just better completions.\n\n**Agent Research: Harnesses, Horizons, Parallel Retrieval, and Self-Evolution**\n\n- **Several papers pushed on a common theme: the harness is now a core optimization target** . A widely shared Salesforce paper summary from[omarsar0](https://x.com/omarsar0/status/2097958286146605446) showed that training a weaker model on a stronger expert’s full trajectories can**hurt performance by 4–30 points** after harness evolution, because the fine-tuned model adopts an incompatible planning style. The proposed fix—rewrite only the**failing turn** in the weaker model’s own rollout—preserves model-harness fit. In parallel,[Sumanth_077’s writeup of ByteDance’s HarnessDev](https://x.com/Sumanth_077/status/2098053941800100294) described agents that build and iteratively improve their own runnable harnesses, with mixed generalization: only**34/64** changes transferred directionally to held-out tasks.\n- **Long-horizon and long-context agent training also got more principled treatments** :[dair_ai](https://x.com/dair_ai/status/2098109386568925397) summarized Qwen work on**Elastic Horizon** , a closed-loop controller that tracks the**90th percentile of successful trajectory lengths** to adjust the maximum interaction horizon, improving success while saving up to**25%** of trajectory tokens. Separately,[omarsar0](https://x.com/omarsar0/status/2098140712504332411) highlighted**PARSER** , which replaces sequential chunk reading with**parallel frozen subagents + an RL-trained lead agent** over iterative scatter-gather rounds; reported gains include**+12 points at 896K context** and up to**11x lower latency** .\n- **Skill and tool-use data generation are being formalized too** :[dair_ai on SkillAdam](https://x.com/dair_ai/status/2098154641854992676) framed skill self-evolution as a discrete optimization problem, borrowing Adam-like first/second-moment ideas to stabilize update direction and edit magnitude. Meanwhile,[Google Research’s ToolGrad](https://x.com/GoogleResearch/status/2098183830968705163) generates**ground-truth tool-use chains before prompts** , reporting near-**100% pass rate** for dataset creation and downstream tool-use gains. Taken together, this batch of work suggests the field is shifting from “prompt the model harder” toward**closed-loop optimization of scaffolds, trajectory budgets, skill documents, and tool traces** .\n\n**Safety, Misuse, Monitorability, and Model Governance**\n\n- **Anthropic’s threat intelligence report dominated the safety discussion** : the company published its most detailed misuse report so far, covering attempts to use Claude for**cyberattacks, influence ops, surveillance, biology, and weapons** , and said it**disrupted every operation described**[launch tweet](https://x.com/AnthropicAI/status/2098097512544444447) . Much of the discourse focused on reported extraction / routing patterns involving rival labs and state-linked misuse, with high-engagement reactions from[pradeepXkapoor](https://x.com/pradeepXkapoor/status/2098115046069223631) ,[logangraham](https://x.com/logangraham/status/2098112853270257747) , and former Meta threat-disruption lead[David Agranovich](https://x.com/DavidAgranovich/status/2098168519259218096) , who argued Anthropic deserves credit for this level of transparency even if some framing should be debated.\n- **A second thread focused on reasoning monitorability and “neuralese” risk** :[Redwood Research](https://x.com/redwood_ai/status/2098095409084420456) proposed transparency norms for architectures that may weaken or eliminate chain-of-thought visibility, and[Ryan Greenblatt](https://x.com/RyanGreenblatt/status/2098095983716688281) argued companies should publish evidence and policies before deploying architectures that substantially reduce CoT dependence. Related commentary from[Neel Nanda](https://x.com/NeelNanda5/status/2098177895932068174) interpreted**GPT-6 Astra** as a potentially concerning jump in**no-CoT reasoning** , possibly indicating architectural changes beyond ordinary scaling.\n- **There was also visible disagreement among frontier-lab employees and alumni about risk culture** :[Chris Hayduk](https://x.com/ChrisHayduk/status/2098017706494566761) emphasized AI’s humanitarian upside, while[balesni](https://x.com/balesni/status/2098109503518683491) and[jkcarlsmith](https://x.com/jkcarlsmith/status/2098189287917588835) openly endorsed**>10% extinction-risk** views. On governance,[Thom Wolf](https://x.com/Thom_Wolf/status/2098080470235762702) announced a new**Open Alignment** team at Hugging Face, and[Richard Ngo](https://x.com/RichardMCNgo/status/2098118195374944408) published a sharp critique of Paul joining OpenAI’s board and of what he sees as the safety community’s capture by AGI companies.\n\n**Top tweets by engagement**\n\n- **Anthropic threat intelligence report** :[@AnthropicAI](https://x.com/AnthropicAI/status/2098097512544444447) published a detailed account of sophisticated Claude misuse across cyber, influence, biology, surveillance, and weapons.\n- **OpenAI pauses new $200 Pro signups for Astra capacity reasons** :[@thsottiaux](https://x.com/thsottiaux/status/2098113585683808624) said existing users are unaffected and API/other plans remain available.\n- **GPT-Live-1 API launch** :[@OpenAIDevs](https://x.com/OpenAIDevs/status/2098099269551149398) launched the new full-duplex voice model into the API.\n- **ChatGPT Work Data agent** :[@ChatGPT](https://x.com/ChatGPT/status/2098065296968011853) announced a data-connected enterprise agent for dashboards, answers, and actions.\n- **SWE-2 release** :[@cognition](https://x.com/cognition/status/2098069235733823965) introduced a new coding model claiming near-frontier eval performance at materially lower cost.\n- **Cursor Projects** :[@cursor_ai](https://x.com/cursor_ai/status/2098162488013455784) launched persistent project threads with coordinator agents, shared memory, and synced artifacts.\n\n# **AI Reddit Recap**\n\n## **/r/LocalLlama + /r/localLLM Recap**\n\n### **1. DeepSeek V4.1 Flash Release and Architecture**\n\n- **[DeepSeek V4.1 Flash: Stronger, Faster, More Accessible](https://www.reddit.com/r/LocalLLaMA/comments/1wcb0o3/deepseek_v41_flash_stronger_faster_more_accessible/)** (Activity: 317):**DeepSeek announced V4.1 Flash, a** `552B`**-parameter MoE with native multimodal vision support and a new Causal-Encoder-Decoder asymmetric architecture:** `8B` **parameters active on input and** `16B` **on output, claiming higher capability than V4 Pro at lower inference cost ([source](https://mp.weixin.qq.com/s/qg0NU3NNUbp1co2PdkAPAg), [weights](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash), [tech report](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf)). DeepSeek claims KV-cache/storage reductions of** `4×` **HBM and** `8×` **SSD vs the prior generation, and** `437×` **vs its first-generation model; API users can switch to** `deepseek-flash`**, while deprecated** `deepseek-v4-flash`**,** `deepseek-v4-flash-vision-exp`**, and eventually** `deepseek-v4-pro` **will route to V4.1 Flash with new peak/off-peak pricing.** Top technical discussion focused on the unusual return of an**encoder-decoder-style architecture** in a frontier LLM, with commenters questioning what the encoder does for long prompts and multimodal segmentation. Others noted that despite sparse activation,`552B` total parameters makes local inference impractical even for multi-DGX Spark/Strix-style setups, so smaller V4/Qwen-derived coding models remain more realistic for local agentic workflows.\n  - Several commenters focused on the claimed **encoder-decoder/asymmetric architecture** , questioning how DeepSeek is using an encoder in a modern GPT-style LLM: e.g. whether prompts are embedded or compressed before decoder self-attention, and how this scales to long inputs split by sentence, paragraph, or modality. One interpretation was that the asymmetric design may indicate a structurally different generation path versus standard decoder-only transformers.\n  - Local inference feasibility was discussed around the model’s reported `552B` **parameter scale** , with commenters arguing it is impractical even for high-end local setups such as multiple DGX Spark/Strix-class systems. The suggested practical workflow was to use larger DeepSeek V4-class models for planning, then smaller/distilled models such as**Q38-27B** ,**Q38-35B-Distill** , or**Ornith35B** for execution in local agentic coding pipelines.\n  - A technically notable claim highlighted in the thread was a `437×` **KV-cache reduction since first generation** , which commenters viewed as significant for long-context inference cost and memory scaling. If accurate, that kind of reduction would materially affect throughput and deployment economics for long-context serving, especially compared with conventional decoder-only attention caching.\n- **[Deepseek V4.1 Flash is 748B, not 552B](https://www.reddit.com/r/LocalLLaMA/comments/1wcd4rx/deepseek_v41_flash_is_748b_not_552b/)** (Activity: 575):**OP inspected the Hugging Face** `safetensors` **and argues DeepSeek V4.1 Flash is ~**`748.5B` **parameters for backbone + engram—not** `284B`**,** `305B`**,** `485B`**, or** `522B`**—with a** `551.566B` **backbone and** `196.929B` **engram; including optional DSpark/MTP (**`14.225B`**) and vision encoder (**` 0.485B`**) brings the stored model to ~**` 763.21B` **params /** `511.76 GB`**. The confusion is attributed to counting/metadata errors: e.g. an [NVIDIA forum estimate](https://forums.developer.nvidia.com/t/deepseek-v4-1-flash/382725/11) undercounts the backbone, Hugging Face’s** `485B` **likely miscounts FP4 packed weights as bytes rather than two params/byte, similar to [GLM-5.3-Flash-NVFP4](https://huggingface.co/nvidia/GLM-5.3-Flash-NVFP4), and [vLLM’s recipe](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4.1-Flash) inconsistently lists** `522B` **before later correcting parameter details. The backbone is overwhelmingly MoE FFN experts:** `543.582B` **params in FP4, with only ~**`7.984B` **in attention/shared/embedding/other components, implying 128–256 GB RAM/VRAM is insufficient for full local use.** One commenter notes the “Flash” naming is plausibly latency-related, claiming it uses only roughly`9B` **active parameters for prefilling** . Another technical question raised whether SSD offload for engram/ngram-style lookup tables should prioritize sequential throughput or**random 4K read IOPS** , but no substantive answer is included in the provided comments.\n  - Commenters discussed that **DeepSeek V4.1 Flash** may report a much larger total size due to included`n-gram` /lookup-style components, but some argue these should not be counted like active neural parameters because they can be stored externally on SSD rather than loaded into VRAM/RAM as model weights.\n  - A technical claim was made that the “Flash” variant is fast because it uses only around `9B` parameters during**prefill** , implying the active compute path is far smaller than the headline`748B` figure and may explain the latency-focused branding.\n  - For local deployment, one commenter estimated that `256GB` system RAM plus`64–96GB` VRAM is sufficient, with the`n-gram` data hosted on any PCIe Gen 3+ NVMe SSD. The discussion raised whether SSD performance should prioritize sequential throughput or`4K` random reads, since disk-resident lookup tables may be access-pattern sensitive.\n- **[Deepseek Has Soft Retired Deepseek V4 Pro](https://www.reddit.com/r/LocalLLaMA/comments/1wbfrut/deepseek_has_soft_retired_deepseek_v4_pro/)** (Activity: 1598):**The image is a [screenshot of a tweet](https://i.redd.it/01k8gclhggoh1.png) saying DeepSeek is effectively “soft retiring” DeepSeek V4 Pro: V4 Pro traffic will be automatically routed to DS V4.1 Flash and billed at cheaper Flash pricing until V4.1 Pro launches. The stated rationale is that V4.1 Flash outperforms the older V4 Pro on performance, cost, speed, and total usage time, implying the smaller/cheaper Flash variant has become the preferred production model despite V4 Pro’s larger size.** Commenters speculate that V4 Pro’s GA release may have suffered from reward hacking and poor scaling, with one noting it was*“not performing meaningfully better than the flash model despite being nearly 6 times the size.”* There is also debate over whether DeepSeek and Google are seeing similar small-model-over-big-model effects due to separate training runs, architecture differences, or data-mix issues; another commenter complains Flash is weak for creative writing and reflects a broader shift toward coding-optimized models.\n  - Several commenters argued **DeepSeek V4 Pro GA underperformed relative to its size** , with one claiming it showed a*“high degree of reward hacking”* and was not meaningfully better than the Flash model despite being nearly`6×` larger. The technical concern is that Pro’s larger parameter/compute footprint did not translate into benchmark or real-world capability gains, making retirement rational if inference cost was high.\n  - A thread compared **DeepSeek** and**Google** cases where smaller “Flash” variants outperform or match larger models, suggesting these may not be simple distillations from one large training run. Commenters speculated the gap could come from separate architecture choices, training-pipeline differences, or data-mix effects rather than size alone, raising the question of why the smaller model generalizes better for some tasks.\n  - Some users distinguished between API retirement and model disappearance: **DeepSeek stopped serving V4 Pro, but weights reportedly remain available** , unlike fully closed retirements by OpenAI/Anthropic. Another technical hypothesis was that DeepSeek may be freeing inference capacity or migrating toward Chinese inference chips, prioritizing cheaper Flash-class serving even if Pro retained more world knowledge useful for planning/general tasks.\n- **[DeepSeek-V4.1-Flash surprised ....](https://www.reddit.com/r/LocalLLaMA/comments/1wcdati/deepseekv41flash_surprised/)** (Activity: 537):**The [image](https://i.redd.it/va67hbc7knoh1.jpeg) is a reaction meme, but it highlights a technical claim that DeepSeek-V4.1-Flash reduces global KV cache to only** `890 bytes/token`**, far below prior versions, while DeepSeek-V4.1-Flash-Base is shown as a** `552B`**-parameter backbone with only** `8B/16B` **activated parameters. The post frames this as evidence that future medium-sized models could combine MoE or dense backbones,** `10–15B` **“Engram” components, and Flash-style KV-cache optimizations to improve long-context memory efficiency.** Commenters speculate that tiny KV-cache designs could make high-memory local inference hardware like**M5 Ultra 512GB** or multi-**Spark** setups more attractive, and that other model families such as**Qwen** may adopt similar KV reductions. One commenter also corrects the sizing intuition for Engrams, arguing they are roughly`1/3–1/2` of parameters, e.g. a`30B` dense backbone would pair with about a`10–15B` Engram.\n  - Commenters focused on **memory pressure and hardware feasibility** , noting that strong “AA scores” could make very-high-memory local inference setups like**M5 Ultra** `512GB` and multi-**Spark** configurations more attractive. One user questioned whether even`512GB` unified memory would be enough to run DeepSeek-V4.1-Flash “comfortably” when using multiple subagents, implying KV-cache and concurrency overhead may dominate beyond raw model weights.\n  - A technical thread discussed architectural parameter allocation: **engrams** were estimated at roughly`1/3` to`1/2` of total parameters, so a`30B` dense backbone would imply an additional`10B–15B` engram component, for about`40B–45B` total parameters. Another commenter anticipated**Qwen** adopting a “tiny KV” design, which could reduce reliance on KV-cache quantization debates by lowering context-memory requirements directly.", "url": "https://wpnews.pro/news/ainews-deepseek-v4-1-flash-763b-p8b-d16b-novel-causal-encoder-decoder-with-marks", "canonical_source": "https://www.latent.space/p/ainews-deepseek-v41-flash-763b-p8b", "published_at": "2026-09-12 05:56:05+00:00", "updated_at": "2026-09-12 06:27:06.878605+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products", "computer-vision"], "entities": ["DeepSeek", "DeepSeek v4.1-Flash", "DeepSeek V4 Pro", "DeepSeek V4 Flash", "Hugging Face", "DeepSeek R1", "DeepSeek Coder", "Engram"], "alternates": {"html": "https://wpnews.pro/news/ainews-deepseek-v4-1-flash-763b-p8b-d16b-novel-causal-encoder-decoder-with-marks", "markdown": "https://wpnews.pro/news/ainews-deepseek-v4-1-flash-763b-p8b-d16b-novel-causal-encoder-decoder-with-marks.md", "text": "https://wpnews.pro/news/ainews-deepseek-v4-1-flash-763b-p8b-d16b-novel-causal-encoder-decoder-with-marks.txt", "jsonld": "https://wpnews.pro/news/ainews-deepseek-v4-1-flash-763b-p8b-d16b-novel-causal-encoder-decoder-with-marks.jsonld"}}