Vast-10M: Technical Report
We introduce Vast-10M, a family of LLMs with ten million tokens of native context. The key enabling technology is a new form of attention, Voltropy Scalable Attention (VSA), which can be retrofitted to existing transformer models.
Prior methods for extending context beyond a million tokens significantly degraded model intelligence. Our results indicate that VSA not only eliminates this tradeoff but improves reasoning quality at shorter contexts. Every VSA-augmented model outperforms its corresponding base model on BEAM at 100K, 500K, and 1M tokens. The flagship model, Vast-10M-Flash, achieves a higher score at 10M tokens than its base model achieves at 1M tokens. At 1M tokens, it beats Anthropic's Fable 5.1 and is on par with OpenAI's GPT-6 Astra.