cd/entity/Kimi Delta Attention· home entities Kimi Delta Attention
grep -l @kimi delta attention /news/*.json | wc -l → 27

Kimi Delta Attention

mentions 27 type Person page 1/2 feed RSS

// recent coverage 27 mentions

04:00
2026-08-31
arxiv.org
artificial-intelligence

DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

Researchers introduced DAMP (Decay-Aware Mixed-Precision Recurrent-State Quantization), a post-training quantization method for recurrent-state language models, which reduces recurrent-state storage b…

15:00
2026-08-07
dibi8.com
artificial-intelligence

Kimi K3: Moonshot AI's 2.8T-Parameter Open-Weight Frontier Model

Moonshot AI released Kimi K3, an open-weight 2.8-trillion-parameter Mixture-of-Experts model with 104B activated parameters, built on the new Kimi Delta Attention architecture and featuring a 1,048,57…

19:42
2026-08-03
newsletter.semianalysis.com
artificial-intelligence

Kimi K3, The Manos, The Mythos, The Legendos

Moonshot AI released Kimi K3, an open frontier model that swept leaderboards, featuring a hybrid attention mechanism with Kimi Delta Attention (KDA), a linear attention layer derived from DeltaNet and…

00:00
2026-08-03
andlukyane.com
artificial-intelligence

Beyond Bigger MoE: How Kimi K3 Scales Context, Depth, and Agents

Moonshot AI's Kimi K3 model scales to 2.8T total parameters with 104B activated per token, using hybrid attention, Attention Residuals, and Stable LatentMoE to support 1M-token agentic trajectories. T…

21:09
2026-08-02
github.com
artificial-intelligence

Show HN: I implemented the Kimi K3 paper from scratch in PyTorch

A developer released a PyTorch implementation of the Kimi K3 architecture from the arXiv paper 'Kimi K3: Open Frontier Intelligence' (arXiv:2607.24653v1), reproducing the paper's Table 1 parameter cou…

00:00
2026-08-01
together.ai
artificial-intelligence

Kimi K3: The Complete Developer Guide

Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weights model, the largest open-weight model ever released and the first open-source model in the 3-trillion-parameter class, designed for l…

00:01
2026-07-30
pub.towardsai.net
large-language-models

Inside Kimi K3 Technical Report

Moonshot AI's Kimi K3 technical report details a 2.8-trillion-parameter model with 104 billion activated parameters, a 1-million-token context window, and native vision, achieving 2.5x scaling efficie…

00:00
2026-07-29
kondasamy.com
artificial-intelligence

Kimi K3: What a 2.8T Open Model Changes for Engineers

Moonshot AI released Kimi K3, a 2.78-trillion-parameter mixture-of-experts model with 104.2 billion active parameters, native vision, and a 1,048,576-token context window, claiming it is the first ope…

17:30
2026-07-28
snowchord.com
artificial-intelligence

Linear Attention, Visualized

Moonshot released Kimi K3, a 2-trillion-parameter model with a 1-million-token context window, in July 2026, using its linear-attention variant Kimi Delta Attention (KDA). The model trails only Claude…

15:08
2026-07-28
sourcefeed.dev
artificial-intelligence

Linear Attention Just Graduated to Frontier Scale

Moonshot AI's Kimi Linear attention architecture, introduced in October, now powers the company's 2.8-trillion-parameter Kimi K3 flagship model released in mid-July, marking the first production deplo…

07:22
2026-07-28
canopywave.com
large-language-models

Show HN: Kimi K3 Is Now Live on Canopy Wave

Moonshot AI's Kimi K3, a flagship model with 2.8 trillion parameters and a 1M-token context window, is now live on Canopy Wave. The model, built on Kimi Delta Attention (KDA) and Attention Residuals, …

17:07
2026-07-27
ollama.com
large-language-models

Kimi-K3 on Ollama

Moonshot AI released Kimi K3, a 2.8T-parameter open-weight multimodal agentic model, on Ollama with Pro or Max subscription and extra usage credits. The model features Kimi Delta Attention and Attenti…

16:09
2026-07-27
geopolitechs.org
artificial-intelligence

Moonshot released Kimi K3 model weights and technical report

Moonshot AI released the model weights and technical report for Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model with native visual understanding and a 1-million-token context window, along …

15:10
2026-07-27
github.com
artificial-intelligence

Kimi K3 Tech Report

Moonshot AI released Kimi K3, an open-weight 2.8-trillion-parameter native multimodal agentic model with a 1-million-token context window, claiming it is the world's first open 3T-class model. Built o…

12:15
2026-07-27
byteiota.com
large-language-models

Kimi K3 Open Weights Live: Self-Host or Use the API?

Moonshot AI released the open weights of its 2.8-trillion-parameter Kimi K3 sparse Mixture-of-Experts model on Hugging Face under an Apache 2.0 license, but the 594 GB MXFP4-quantized model requires a…

02:49
2026-07-21
arxiv.org
artificial-intelligence

Kimi Linear: An Expressive, Efficient Attention Architecture

Researchers at Moonshot AI introduced Kimi Linear, a hybrid linear attention architecture that outperforms full attention across short-context, long-context, and reinforcement learning scaling regimes…

17:00
2026-07-17
usewire.io
artificial-intelligence

Kimi K3's 1M context runs mostly on linear attention

Moonshot's Kimi K3, a 2.8-trillion-parameter mixture-of-experts model released July 16, achieves a 1M-token context window with decoding up to 6.3x faster than its predecessor by using Kimi Delta Atte…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics