Introducing Vast-10M

The first frontier LLM with ten million tokens of context.

1M · Today's frontier Vast-10M

Today, Voltropy is unveiling its first model family: Vast-10M, which combines frontier intelligence with a ten-million-token native context window.

Breaking the million-token barrier
Context windows drawn to scale by area: Vast-10M 10 million tokens; GPT-6 Astra about 1.05 million; Claude Fable 5.1 1 million 10,000,000 Vast-10M native context ≈1,050,000 GPT-6 Astra 1,000,000 Claude Fable 5.1

Despite amazing breakthroughs in LLMs over the last three years, the context windows for frontier models have been stuck at one million tokens. Vast-10M breaks the million-token barrier, delivering 10× more context without sacrificing intelligence.

It comes in three sizes:

Flagship

Vast-10MFlash

Based on DeepSeek V4.0 Flash

Mid-size

Vast-10MMedium

Based on GLM-5.2

Largest

Vast-10MPro

Based on DeepSeek V4.0 Pro

What unites the Vast-10M family is a new algorithm, Voltropy Scalable Attention (VSA), which improves intelligence at shorter contexts while unlocking unprecedented long-context recall and reasoning abilities.

Vast-10M beats its base model at 10× the contextFig. 02
Vast-10M-Flash scores 40.20 at 10M tokens, above its DeepSeek V4.0 Flash base model at 1M tokens (39.69) 30 35 40 45 50 55 60 100K 500K 1M 10M 40.20 Vast-10M-Flash at 10M 39.69 Base model at 1M Context length (tokens, log scale) BEAM score Vast-10M-Flash DeepSeek V4.0 Flash
At ten million tokens, Vast-10M-Flash (40.20) outscores its own base model at one million tokens (39.69), achieving 101% of the base model's score at ten times the context.

At one million tokens, Vast-10M-Flash beats Claude Fable 5.1 on the BEAM (Beyond a Million Tokens) benchmark and achieves parity with GPT-6 Astra. That is where those models stall, but Vast-10M keeps going. It retains 82.49% of its score when the context increases an order of magnitude. On the ten-million-token BEAM tier, Vast-10M-Flash scores 40.20, the highest score any raw model has ever achieved, and beyond the scores reported for most RAG systems.

10Mtokens of native context in every Vast-10M model
82.49%of Vast-10M-Flash’s 1M-token BEAM score retained at 10M tokens
+10.68%higher BEAM score than Claude Fable 5.1 at one million tokens

Extending Context While Boosting Intelligence

VSA improved every base model at every context lengthFig. 03
BEAM score with and without VSA at 100K, 500K and 1M tokens, for all three Vast-10M models 0 10 20 30 40 50 60 52.66 +8.95 100K 50.22 +5.50 500K 48.73 +9.04 1M DeepSeek V4.0 Flash → Vast-10M-Flash 49.02 +9.52 100K 44.75 +4.26 500K 45.67 +5.61 1M GLM-5.2 → Vast-10M-Medium 43.93 +1.66 100K 44.50 +3.45 500K 43.76 +5.74 1M DeepSeek V4.0 Pro → Vast-10M-Pro Base model Gain from VSA
BEAM scores for each model with and without VSA.

There have previously been two main approaches to extending the context windows of LLMs. One approach was to cram more tokens into transformers by making attention cheaper but less accurate. The other was to abandon transformers in favor of cheaper attention-free architectures like state-space models or other linear RNNs.

These strategies have the same flaw: they purchase more context at the cost of less intelligence. Vast-10M eliminates this tradeoff. The Vast-10M models are all transformers, but they use a new kind of scalable attention (VSA) that improves reasoning at shorter contexts rather than degrading it. This improvement is clear in the BEAM benchmark, where VSA improves scores at all context lengths, including the shortest 100K tier.

Adding VSA beat a new model generationFig. 04
BEAM scores: Vast-10M-Flash versus DeepSeek V4.1 Flash and DeepSeek V4.0 Flash at 100K, 500K and 1M 0 10 20 30 40 50 60 52.66 45.47 43.71 100K 50.22 44.94 44.72 500K 48.73 46.00 39.69 1M Vast-10M-Flash DeepSeek V4.1 Flash DeepSeek V4.0 Flash
Vast-10M-Flash is built on DeepSeek V4.0 Flash. It outscores DeepSeek's successor, V4.1 Flash, at 100K, 500K and 1M tokens.

The size of the improvement is underscored by the fact that Vast-10M-Flash dominates DeepSeek V4.1 Flash, even though Vast-10M was based on the older DeepSeek V4.0 Flash. The addition of VSA to V4.0 Flash yielded a larger boost on BEAM than DeepSeek achieved by training a new model.

World-class reasoningFig. 05
BEAM score at 1M tokens: Vast-10M-Flash 48.73, GPT-6 Astra 49.75, Claude Fable 5.1 44.03 0 10 20 30 40 50 60 Vast-10M-Flash 48.73 GPT-6 Astra 49.75 Claude Fable 5.1 44.03 BEAM score at the 1M-token tier (0–100)
BEAM, one-million-token tier. Vast-10M-Flash beats Claude Fable 5.1 by 4.70 points and reaches 97.95% of GPT-6 Astra's score.

Vast-10M-Flash is so capable that it can compete with the best closed models in the world inside their advertised context windows. Vast-10M-Flash beats Fable 5.1 on the one-million-token tier of BEAM, and it achieves parity with GPT-6 Astra.

This does not mean that Vast-10M is universally more capable than Fable or Astra. Those models possess more raw intelligence than Vast-10M for the absolute hardest tasks, such as frontier mathematics. But Vast-10M can outperform them on consumer and enterprise workloads that require a fusion of high-precision recall and reasoning.

A New Scaling Dimension

Vast-10M's native context window is 10× larger than Fable 5.1's and 9.5× larger than GPT-6 Astra's. This provides unprecedented capabilities for processing large quantities of data. With Vast-10M, the entire U.S. tax code fits into context. So do earnings calls for the entire S&P 500, the transcript of a multi-month trial, or a full year of the New England Journal of Medicine.

A larger context window also provides a major advantage when searching for cybersecurity vulnerabilities. Instead of looking at pieces of a repo in isolation, Vast-10M can hunt bugs caused by the interplay between distant lines of code. It can analyze the entire React codebase. Or every line of SQLite. Or the full TypeScript compiler. Each of those fits into the model's native context window.

Native Context vs. External Memory

A natural question is why the world needs models with ten million tokens of native context, when an entire ecosystem of external retrieval and memory systems has cropped up.

More context means retrieval needs less precisionFig. 06
With 1,000-token chunks, a 1M-token context holds the top 1,000 retrieved chunks and a 10M-token context holds the top 10,000. A needed chunk ranked 3,500th by the retriever is missed at 1M and included at 10M. 1M-token context Top 1,000 chunks fit Missed 10M-token context Top 10,000 chunks fit In context

This is a false dichotomy: mega-context models like Vast-10M can be combined with external retrieval systems to make those systems more powerful. When native context grows 10×, ten times as many chunks can be injected into that context by a RAG system, so the precision required to surface a particular chunk drops by an order of magnitude. Expanding the target makes it easier to hit.

Native context also has many advantages over external retrieval. The chunks of text that an external retrieval system fetches are often poor predictions of what a native model will find valuable for accomplishing a task. This difference can be obscured on retrieval benchmarks that favor RAG systems, but it shows up in real-world usage, where RAG frequently produces disjointed outputs.

Vast-10M also offers far greater flexibility than RAG systems. Retrieval pipelines typically must be tuned for a particular dataset, with custom ontologies or chunking strategies that cannot be transferred between tasks. A Vast-10M endpoint processes ten million tokens out of the box, without any tuning or configuration. Tasks that used to need a custom retrieval pipeline now take a single API call.

Mega Context as Continual Learning

Larger context has immediate benefits for today's workloads, but our lab was founded on the belief that expanding context windows is also an overlooked path to building superintelligence. We created Voltropy Scalable Attention (VSA) to unblock that path.

The greatest shortcoming of current models is that they cannot learn effectively from experience once they are deployed. Anthropic, OpenAI, and others have tried to solve this problem by pursuing models that can retrain their weights at inference time. So far, that technique has not worked well enough to be broadly deployed in the economy.

At Voltropy, we believe there is a faster path to continual learning. Transformers are already remarkably good at learning inside their context windows. They can not only remember facts but pick up new skills and even new languages.

If VSA can make their context windows large enough, the models will be able to learn continuously once deployed, without needing to retrain their weights.

Even the partial realization of this vision, such as VSA models capable of learning on the job for a few days, will allow AI agents to enter sectors of the economy where they have not yet made a dent.

A Safer Kind of Superintelligence

We plan to push VSA even further, creating models capable of processing trillions of tokens inside native context. We believe this is a capital-efficient path to training the smartest, safest models the world has ever seen.

Separating what a model knows from what it can doFig. 07
Today, a transformer stores skills and knowledge together in its weights. In a VSA transformer, skills stay in the weights and knowledge is stored in context. Today’s LLMs Weights Future VSA transformers Weights Context Skills Knowledge
Today’s transformers store skills and knowledge together in their weights. VSA may allow us to create transformers that store their knowledge in context, where it can be edited, upgraded, and deleted.

Today, transformers store their knowledge and skills together inside their weights. It is our goal to use VSA to create a new kind of transformer: one whose neurons develop deep general intelligence, but whose knowledge of the world is stored separately in a structured, context-like format where it can be edited, upgraded, and deleted on demand.

The result will be transformers that behave more like traditional computers (what researchers call von Neumann machines) and less like human brains. This has a variety of advantages over the existing biologically inspired paradigm.

It improves training efficiency, because backpropagation can be targeted to skill development rather than memorization. It slashes serving costs, because each model can be customized to store only the specific knowledge needed to perform its job, not everything it memorizes from the internet. And it provides new safety guarantees, because humans can monitor and control what models know rather than having that information hidden in their weights.

Once again, even the partial realization of this vision would have a transformative effect. Knowledge and skills may be too tightly coupled to ever be completely disaggregated. But the more that we can shrink the neural surface area by shifting data from parametric storage to context, the more realistic it becomes to use techniques like mechanistic interpretability to design provably safe LLMs.

From the user perspective, the result of this architecture will be something new: an LLM that is not just smart, but dependable. One that remembers everything it has ever read, and everything it has ever been told, unless you tell it to forget. An AI that is more responsible than a human employee or assistant, not less.

Get Early Access to Vast-10M

Vast-10M is an intermediate step towards the kinds of systems we hope to build. It retrofits VSA to base models that were trained with different forms of attention.

Yet even this imperfect combination is already enough to deliver an order-of-magnitude improvement in long-context reasoning. To try ten million tokens of native context yourself, sign up here for early access to Vast-10M-Flash.

Get early access → Technical report ↗