Research Papers

01

Vast-10M: Technical Report

New · September 29, 2026

We introduce Vast-10M, a family of LLMs with ten million tokens of native context. The key enabling technology is a new form of attention, Voltropy Scalable Attention (VSA), which can be retrofitted to existing transformer models.

Prior methods for extending context beyond a million tokens significantly degraded model intelligence. Our results indicate that VSA not only eliminates this tradeoff but improves reasoning quality at shorter contexts. Every VSA-augmented model outperforms its corresponding base model on BEAM at 100K, 500K, and 1M tokens. The flagship model, Vast-10M-Flash, achieves a higher score at 10M tokens than its base model achieves at 1M tokens. At 1M tokens, it beats Anthropic's Fable 5.1 and is on par with OpenAI's GPT-6 Astra.

02

Lossless Context Management

February 14, 2026

We introduce Lossless Context Management (LCM), a deterministic architecture for LLM memory that outperforms Claude Code on long-context tasks. When benchmarked using Opus 4.6, our LCM-augmented coding agent, Volt, achieves higher scores than Claude Code on the OOLONG longcontext eval, including at every context length between 32K and 1M tokens.

LCM may be considered both a vindication and extension of the recursive paradigm pioneered by Recursive Language Models (RLMs). Our results demonstrate that recursive context manipulation can outperform not just conventional LLMs, but frontier coding agents with native file-system access.retrievability of all prior state.

LCM departs from RLM by decomposing symbolic recursion into two deterministic, enginemanaged mechanisms: recursive context compression, in which a hierarchical summary DAG automatically compacts older messages while retaining lossless pointers to every original; and recursive task partitioning, in which engine-managed parallel primitives like LLM-Map replace model-written loops. This trade-off, analogous to the move from GOTO to structured control flow in programming language design, sacrifices maximal flexibility for termination guarantees, zero-cost continuity on short tasks, and lossless retrievability of all prior state.

Read more →