Vast-10M: The First Frontier LLM with 10M Tokens of Context Voltropy unveiled Vast-10M, a model family it says is the first frontier LLM with a 10-million-token native context window, built on its new Voltropy Scalable Attention (VSA) algorithm. Voltropy reported that Vast-10M-Flash scores 40.20 on the ten-million-token BEAM tier, retains 82.49% of its 1M-token BEAM score at 10M tokens, and beats Claude Fable 5.1 by 10.68% at one million tokens while reaching parity with GPT-6 Astra. The family ships in three sizes — Vast-10M-Flash (based on DeepSeek V4.0 Flash), Vast-10M-Medium (based on GLM-5.2), and Vast-10M-Pro (based on DeepSeek V4.0 Pro) — with early access opening today. - A new form of attention, VSA, gives Vast-10M ten times the native context of OpenAI or Anthropic’s flagship models. - Vast-10M-Flash beats Fable 5.1 head-to-head inside Fable’s advertised context window and matches GPT-6 Astra. - Get early access to Vast-10M https://www.voltropy.com/early-access/ starting today. Today, Voltropy is unveiling its first model family: Vast-10M, which combines frontier intelligence with a ten-million-token native context window. Despite amazing breakthroughs in LLMs over the last three years, the context windows for frontier models have been stuck at one million tokens. Vast-10M breaks the million-token barrier, delivering 10× more context without sacrificing intelligence. It comes in three sizes: Vast-10MFlash Based on DeepSeek V4.0 Flash Vast-10MMedium Based on GLM-5.2 Vast-10MPro Based on DeepSeek V4.0 Pro What unites the Vast-10M family is a new algorithm, Voltropy Scalable Attention VSA , which improves intelligence at shorter contexts while unlocking unprecedented long-context recall and reasoning abilities. At one million tokens, Vast-10M-Flash beats Claude Fable 5.1 on the BEAM Beyond a Million Tokens benchmark and achieves parity with GPT-6 Astra. That is where those models stall, but Vast-10M keeps going. It retains 82.49% of its score when the context increases an order of magnitude. On the ten-million-token BEAM tier, Vast-10M-Flash scores 40.20, the highest score any raw model has ever achieved, and beyond the scores reported for most RAG systems. 10M tokens of native context in every Vast-10M model 82.49% of Vast-10M-Flash’s 1M-token BEAM score retained at 10M tokens +10.68% higher BEAM score than Claude Fable 5.1 at one million tokens Extending Context While Boosting Intelligence There have previously been two main approaches to extending the context windows of LLMs. One approach was to cram more tokens into transformers by making attention cheaper but less accurate. The other was to abandon transformers in favor of cheaper attention-free architectures like state-space models or other linear RNNs. These strategies have the same flaw: they purchase more context at the cost of less intelligence. Vast-10M eliminates this tradeoff. The Vast-10M models are all transformers, but they use a new kind of scalable attention VSA that improves reasoning at shorter contexts rather than degrading it. This improvement is clear in the BEAM benchmark, where VSA improves scores at all context lengths, including the shortest 100K tier. The size of the improvement is underscored by the fact that Vast-10M-Flash dominates DeepSeek V4.1 Flash, even though Vast-10M was based on the older DeepSeek V4.0 Flash. The addition of VSA to V4.0 Flash yielded a larger boost on BEAM than DeepSeek achieved by training a new model. Vast-10M-Flash is so capable that it can compete with the best closed models in the world inside their advertised context windows. Vast-10M-Flash beats Fable 5.1 on the one-million-token tier of BEAM, and it achieves parity with GPT-6 Astra. This does not mean that Vast-10M is universally more capable than Fable or Astra. Those models possess more raw intelligence than Vast-10M for the absolute hardest tasks, such as frontier mathematics. But Vast-10M can outperform them on consumer and enterprise workloads that require a fusion of high-precision recall and reasoning. A New Scaling Dimension Vast-10M's native context window is 10× larger than Fable 5.1's and 9.5× larger than GPT-6 Astra's. This provides unprecedented capabilities for processing large quantities of data. With Vast-10M, the entire U.S. tax code fits into context. So do earnings calls for the entire S&P 500, the transcript of a multi-month trial, or a full year of the New England Journal of Medicine . A larger context window also provides a major advantage when searching for cybersecurity vulnerabilities. Instead of looking at pieces of a repo in isolation, Vast-10M can hunt bugs caused by the interplay between distant lines of code. It can analyze the entire React codebase. Or every line of SQLite. Or the full TypeScript compiler. Each of those fits into the model's native context window. Native Context vs. External Memory A natural question is why the world needs models with ten million tokens of native context, when an entire ecosystem of external retrieval and memory systems has cropped up. This is a false dichotomy: mega-context models like Vast-10M can be combined with external retrieval systems to make those systems more powerful. When native context grows 10×, ten times as many chunks can be injected into that context by a RAG system, so the precision required to surface a particular chunk drops by an order of magnitude. Expanding the target makes it easier to hit. Native context also has many advantages over external retrieval. The chunks of text that an external retrieval system fetches are often poor predictions of what a native model will find valuable for accomplishing a task. This difference can be obscured on retrieval benchmarks that favor RAG systems, but it shows up in real-world usage, where RAG frequently produces disjointed outputs. Vast-10M also offers far greater flexibility than RAG systems. Retrieval pipelines typically must be tuned for a particular dataset, with custom ontologies or chunking strategies that cannot be transferred between tasks. A Vast-10M endpoint processes ten million tokens out of the box, without any tuning or configuration. Tasks that used to need a custom retrieval pipeline now take a single API call. Mega Context as Continual Learning Larger context has immediate benefits for today's workloads, but our lab was founded on the belief that expanding context windows is also an overlooked path to building superintelligence. We created Voltropy Scalable Attention VSA to unblock that path. The greatest shortcoming of current models is that they cannot learn effectively from experience once they are deployed. Anthropic, OpenAI, and others have tried to solve this problem by pursuing models that can retrain their weights at inference time. So far, that technique has not worked well enough to be broadly deployed in the economy. At Voltropy, we believe there is a faster path to continual learning. Transformers are already remarkably good at learning inside their context windows. They can not only remember facts but pick up new skills and even new languages. If VSA can make their context windows large enough, the models will be able to learn continuously once deployed, without needing to retrain their weights. Even the partial realization of this vision, such as VSA models capable of learning on the job for a few days, will allow AI agents to enter sectors of the economy where they have not yet made a dent. A Safer Kind of Superintelligence We plan to push VSA even further, creating models capable of processing trillions of tokens inside native context. We believe this is a capital-efficient path to training the smartest, safest models the world has ever seen. Today, transformers store their knowledge and skills together inside their weights. It is our goal to use VSA to create a new kind of transformer: one whose neurons develop deep general intelligence, but whose knowledge of the world is stored separately in a structured, context-like format where it can be edited, upgraded, and deleted on demand. The result will be transformers that behave more like traditional computers what researchers call von Neumann machines and less like human brains. This has a variety of advantages over the existing biologically inspired paradigm. It improves training efficiency, because backpropagation can be targeted to skill development rather than memorization. It slashes serving costs, because each model can be customized to store only the specific knowledge needed to perform its job, not everything it memorizes from the internet. And it provides new safety guarantees, because humans can monitor and control what models know rather than having that information hidden in their weights. Once again, even the partial realization of this vision would have a transformative effect. Knowledge and skills may be too tightly coupled to ever be completely disaggregated. But the more that we can shrink the neural surface area by shifting data from parametric storage to context, the more realistic it becomes to use techniques like mechanistic interpretability to design provably safe LLMs. From the user perspective, the result of this architecture will be something new: an LLM that is not just smart, but dependable. One that remembers everything it has ever read, and everything it has ever been told, unless you tell it to forget. An AI that is more responsible than a human employee or assistant, not less. Get Early Access to Vast-10M Vast-10M is an intermediate step towards the kinds of systems we hope to build. It retrofits VSA to base models that were trained with different forms of attention. Yet even this imperfect combination is already enough to deliver an order-of-magnitude improvement in long-context reasoning. To try ten million tokens of native context yourself, sign up here https://www.voltropy.com/early-access/ for early access to Vast-10M-Flash.