Naive-N0.5-Flash pairs a 309B MoE architecture with hybrid SWA-DSA attention for native 1M-token context, built on MiMo-V2.5.
What is Naive-N0.5-Flash? #
Naive-N0.5-Flash is an open-weight mixture-of-experts (MoE) language model from NaiveAI, built on top of Xiaomi’s MiMo-V2.5 base model and aimed at coding and AI R&D tasks. It has 309 billion total parameters with 15.5 billion active per token, and it supports a native 1 million token context window using a hybrid attention design that mixes Sliding-Window Attention (SWA) with a lightweight version of DeepSeek Sparse Attention (DSA). Notably, the architecture contains no full-attention layers at all. Weights are released under the MIT license, alongside an API priced at $0.10 per million input tokens, $0.40 per million output tokens, and $0.01 per million cached tokens.
TL;DR #
- Naive-N0.5-Flash is a 309B-parameter MoE model with only 15.5B active parameters per token, keeping inference cost close to a much smaller dense model.
- The model reaches a native 1M-token context window without any full-attention layers, relying entirely on local (SWA) and sparse (DSA) attention instead.
- Its architecture uses 48 transformer layers split 39 SWA to 9 DSA , organized into eight six-layer modules, with a 128-token SWA window and top-2,048 token selection for DSA.
- DSA replaces the original MLA-based design with grouped-query attention using four KV groups (GQA4) , paired with a 16-head lightweight indexer that scores the full sequence history.
- The model went through 3.25 trillion tokens of continued training (indexer warmup, sparse attention training, and learning-rate decay) specifically to adapt to the new attention structure.
- A custom inference stack called NaiveRT combines mega-kernel fusion, Programmatic Dependent Launch, and speculative decoding to hit up to 2,000 tokens/s in its fastest mode.
- The model card reports results across coding and agentic benchmarks (like SWE-Bench Pro, Terminal-Bench 2.1, and ALE-CLI) as well asAI research benchmarks (MLE-bench-30, PaperBench, NanoGPT SpeedRun), positioning it against models like GLM-5.3, Kimi-K3, and DeepSeek-V4.1-Flash.
Remy doesn't build the plumbing. It inherits it. #
Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.
Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.
How does the hybrid SWA-DSA attention work? #
The core engineering problem Naive-N0.5-Flash tries to solve is the cost of long context. Standard transformers use full attention in at least some layers, and at a million tokens that gets expensive fast, both in compute and memory bandwidth. MiMo-V2.5, the base model this work builds on, already leans heavily on Sliding-Window Attention, where most layers only attend to a local window (128 tokens here) and the per-token decoding cost doesn’t grow as context length increases. But MiMo-V2.5 still keeps a handful of full-attention layers to preserve long-range dependencies, and at million-token scale those few layers end up dominating decode-time overhead.
Naive-N0.5-Flash’s fix is to replace those global-attention layers with DeepSeek Sparse Attention (DSA) instead of removing long-range attention altogether. In DSA, a lightweight indexer scans the entire token history and scores it, then the backbone network only computes full attention over a selected subset, the top 2,048 tokens in this case. The indexer still touches every token and the full KV cache is retained in memory, but the actual attention computation and the associated memory access drop substantially because the expensive part (backbone attention) only runs over a fraction of the sequence.
Structurally, the 48 layers are organized into eight six-layer modules. Each standard module runs five SWA layers followed by one DSA layer, and the very first layer of the first module is also swapped to DSA. That gives a roughly 5:1 ratio of local to sparse attention across the network, with sink bias applied to both attention types to stabilize behavior at long range. One additional change from the original DeepSeek DSA formulation: Naive-N0.5-Flash swaps out Multi-head Latent Attention (MLA) for grouped-query attention with four KV groups, which affects how the indexer and backbone share key-value state.
Why does the active-parameter count matter? #
At 309B total parameters, Naive-N0.5-Flash sits in a similar bracket to other large frontier-scale open models, but only 15.5B parameters activate per forward pass. That’s the point of the MoE design: the model can specialize across a large pool of experts while keeping the actual per-token compute closer to that of a mid-sized dense model. In practice this is what makes serving a 300B+ parameter model at meaningful throughput feasible without needing an enormous multi-node cluster for every request.
Combined with the sparse/local attention scheme, the active-parameter budget is also why the model can push to a 1M-token context window without decode costs spiraling. Full attention at that scale, even with a small MoE, would still be attention-bound at long context. Replacing global attention with DSA keeps decode-time cost roughly flat as context grows, which is the actual engineering claim being made here, not just “big model with big context window.”
What is NaiveRT and how fast is inference? #
#
Plans first. Then code.
Remy writes the spec, manages the build, and ships the app.
NaiveRT is the inference system built specifically for Naive-N0.5-Flash. According to the model’s documentation, it combines three techniques: mega-kernel fusion (merging multiple GPU operations into fewer kernel launches to cut overhead), Programmatic Dependent Launch (a CUDA scheduling technique for overlapping dependent kernel execution), and speculative decoding (using a smaller draft process to predict tokens ahead of the main model, then verifying them in batches).
The reported throughput is 50 tokens per second per user in “Standard” mode and up to 2,000 tokens per second in “Ultrafast” mode. That’s a big spread, and it’s worth reading the gap as the difference between a mode optimized for maximum concurrent users at moderate per-user speed versus a mode optimized for raw single-stream throughput, likely leaning harder on speculative decoding and batching tricks. NaiveAI describes the build and tuning process for NaiveRT in a separate technical blog post referenced from the model card.
How was the model trained? #
Naive-N0.5-Flash didn’t train from scratch. It started from the MiMo-V2.5 base model and went through continued training totaling 3.25 trillion tokens, split into three stages. First, 50 billion tokens of “Indexer Warmup,” which presumably trains the DSA indexer to score token relevance before it’s relied on for sparse selection. Second, the bulk of the work: 3 trillion tokens of “Sparse Attention Training,” adapting the full network to operate under the new hybrid SWA-DSA structure instead of the original attention pattern. Third, 200 billion tokens of learning-rate decay to settle the model into its final state.
The stated purpose of this training run was twofold: adapt the base model to its new attention architecture, and substantially improve coding and AI R&D capability, which lines up with the benchmark categories the model card reports results against.
Is Naive-N0.5-Flash worth using for coding and AI R&D work? #
The model card reports benchmark results across two broad categories: software engineering/agentic coding tasks (including SWE-Bench Pro, DeepSWE v1.1, Terminal-Bench 2.1, ALE-CLI, FrontierSWE v1, and ProgramBench) and AI research tasks (MLE-bench-30, PaperBench, SOL-ExecBench, NanoChat AutoResearch, and NanoGPT SpeedRun). Comparison points named in the documentation include GLM-5.3 and GLM-5.3-Flash, Kimi-K3, Qwen-3.8-Max, Hy4-preview, DeepSeek-V4.1-Flash, Step-5-preview, and various Claude Opus and GPT model versions, depending on which benchmark’s published leaderboard the scores are drawn from.
The evaluation setup used Claude Code 2.1.207 with the full 1M-token context window, temperature 1.0, and top-p 0.95, with the harness restricted to basic file I/O and Bash tools, which is a reasonably realistic approximation of an autonomous coding agent rather than a synthetic single-turn benchmark. Whether the model is “worth it” for a given team depends heavily on workload: the appeal is clearly long-context agentic coding and research automation where a 1M-token window and MoE-level inference economics both matter, rather than short-context chat or general-purpose assistant use.
Frequently Asked Questions #
What does MoE mean for Naive-N0.5-Flash specifically?
It means the model has 309 billion total parameters spread across many experts, but only 15.5 billion parameters actually activate for any given token. This keeps inference compute much closer to a 15B dense model while retaining the capacity benefits of a much larger network.
How does Naive-N0.5-Flash achieve 1M-token context without full attention?
Other agents start typing. Remy starts asking. #
Scoping, trade-offs, edge cases — the real work. Before a line of code.
It replaces the global full-attention layers found in its MiMo-V2.5 base with DeepSeek Sparse Attention (DSA), where a lightweight indexer scores the entire history but the backbone only computes attention over the top 2,048 selected tokens. The remaining layers use Sliding-Window Attention with a 128-token window, so no layer in the network performs full quadratic attention.
What hardware is needed to run Naive-N0.5-Flash?
The model card states it requires FP8-capable NVIDIA GPUs, with model weights occupying roughly 315 GB, plus additional GPU memory needed for inference and the KV cache at long context lengths.
How is Naive-N0.5-Flash licensed and priced?
The weights and inference code are released under the MIT license. API access is priced at $0.10 per million input tokens, $0.40 per million output tokens, and $0.01 per million cached-token reads.
What models is Naive-N0.5-Flash benchmarked against?
The model card compares it against models including GLM-5.3, GLM-5.3-Flash, Kimi-K3, Qwen-3.8-Max, Hy4-preview, DeepSeek-V4.1-Flash, and various Claude Opus and GPT models, with scores sourced from each model’s own published blog posts, model cards, or public leaderboards.