about
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters (arxiv.org)
1 point by onurkanbkrc 234 days ago | hide | past | pdf | 1 comment on HN

In plain words: Built a huge AI model that wakes up only 11 billion of its 196 billion parts per question, so it runs fast and cheap while still reasoning sharply. It matched top models on math, coding, and web-browsing tasks.

Abstract

We introduce Step 3.5 Flash, a sparse Mixture-of-Experts (MoE) model that bridges frontier-level agentic intelligence and computational efficiency. We focus on what matters most when building agents: sharp reasoning and fast, reliable execution. Step 3.5 Flash pairs a 196B-parameter foundation with 11B active parameters for efficient inference. It is optimized with interleaved 3:1 sliding-window/full attention and Multi-Token Prediction (MTP-3) to reduce the latency and cost of multi-round agentic interactions. To reach frontier-level intelligence, we design a scalable reinforcement learning framework that combines verifiable signals with preference feedback, while remaining stable under large-scale off-policy training, enabling consistent self-improvement across mathematics, code, and tool use. Step 3.5 Flash demonstrates strong performance across agent, coding, and math tasks, achieving 85.4% on IMO-AnswerBench, 86.4% on LiveCodeBench-v6 (2024.08-2025.05), 88.2% on tau2-Bench, 69.0% on BrowseComp (with context management), and 51.0% on Terminal-Bench 2.0, comparable to frontier models such as GPT-5.2 xHigh and Gemini 3.0 Pro. By redefining the efficiency frontier, Step 3.5 Flash provides a high-density foundation for deploying sophisticated agents in real-world industrial environments.

Ailin Huang, Ang Li, Aobo Kong, Bin Wang, Binxing Jiao, Bo Dong, Bojun Wang, Boyu Chen, Brian Li, Buyun Ma, Chang Su, Changxin Miao, et al.
arXiv:2602.10604 · cs.CL, cs.AI · submitted Feb 11, 2026 · updated Feb 23, 2026
abstract · pdf · html · Technical report for Step 3.5 Flash

add comment on HN

In the process of downloading this to run on a DGX Spark which is explicitly mentioned in their documentation. Never heard of it before but sounds good, too good to be true?