about
K2-think: A parameter-efficient reasoning system (arxiv.org)
52 points by mgl on Sep 12, 2025 | hide | past | pdf | 7 comments on HN

In plain words: A 32-billion-parameter model is trained on long step-by-step reasoning examples and rewarded for correct answers, then sped up with faster decoding hardware. It matches or beats models nearly four times its size on math, while generating over 2,000 tokens per second.

Abstract · K2-Think: A Parameter-Efficient Reasoning System

K2-Think is a reasoning system that achieves state-of-the-art performance with a 32B parameter model, matching or surpassing much larger models like GPT-OSS 120B and DeepSeek v3.1. Built on the Qwen2.5 base model, our system shows that smaller models can compete at the highest levels by combining advanced post-training and test-time computation techniques. The approach is based on six key technical pillars: Long Chain-of-thought Supervised Finetuning, Reinforcement Learning with Verifiable Rewards (RLVR), Agentic planning prior to reasoning, Test-time Scaling, Speculative Decoding, and Inference-optimized Hardware, all using publicly available open-source datasets. K2-Think excels in mathematical reasoning, achieving state-of-the-art scores on public benchmarks for open-source models, while also performing strongly in other areas such as Code and Science. Our results confirm that a more parameter-efficient model like K2-Think 32B can compete with state-of-the-art systems through an integrated post-training recipe that includes long chain-of-thought training and strategic inference-time enhancements, making open-source reasoning systems more accessible and affordable. K2-Think is freely available at k2think.ai, offering best-in-class inference speeds of over 2,000 tokens per second per request via the Cerebras Wafer-Scale Engine.

Zhoujun Cheng, Richard Fan, Shibo Hao, Taylor W. Killian, Haonan Li, Suqi Sun, Hector Ren, Alexander Moreno, Daqian Zhang, Tianjun Zhong, Yuxin Xiong, Yuanzhe Hu, et al.
arXiv:2509.07604 · cs.LG · submitted Sep 9, 2025 · updated Sep 15, 2025
abstract · pdf · html · To access the K2-Think reasoning system, please visit www.k2think.ai

add comment on HN

Debunking the Claims of K2-Think https://www.sri.inf.ethz.ch/blog/k2think
Thanks for sharing that post here!
Can reasoners be optimizers?

Like does reasoning find a gradient to optimize a solution? Or are they just trying to expand state until finding what the LLMs world knowledge would say is highest probability?

For example, I can imagine an LLM reasoner might run out of state trying to perfectly solve for 50 intricate unit tests. Because it ping pongs between solving one case, then another, playing whack-a-mole and not converging.

Maybe there's an "oh duh" answer to this, but where I struggle with the limits of agentic work vs traditional ML.

They can be, in the same way humans can be optimizers.

In most cases, there's no explicit descent - and if any descent-like process happens at all, it's not exactly exposed or expressed as hard logic.

If you want it to happen consistently, you add scaffolding and get something like AlphaEvolve at home.

Does it ever respond when you ask it something? I started a query at https://www.k2think.ai/guest thirteen minutes ago and haven't gotten an answer yet.
So currently what are the best OSS reasoning models? (and how much compute the needed)