In plain words: Instead of teaching reasoning by trial and error or copying a teacher, this approach works backward from good answers to find the step-by-step thinking that produced them. Trained on 20,000 traces, it beats open-source rivals and matches or even beats GPT-4o and Claude 3.5.
Abstract
While the ``deep reasoning'' paradigm has spurred significant advances in verifiable domains like mathematics, its application to open-ended, creative generation remains a critical challenge. The two dominant methods for instilling reasoning -- reinforcement learning (RL) and instruction distillation -- falter in this area; RL struggles with the absence of clear reward signals and high-quality reward models, while distillation is prohibitively expensive and capped by the teacher model's capabilities. To overcome these limitations, we introduce REverse-Engineered Reasoning (REER), a new paradigm that fundamentally shifts the approach. Instead of building a reasoning process ``forwards'' through trial-and-error or imitation, REER works ``backwards'' from known-good solutions to computationally discover the latent, step-by-step deep reasoning process that could have produced them. Using this scalable, gradient-free approach, we curate and open-source DeepWriting-20K, a large-scale dataset of 20,000 deep reasoning trajectories for open-ended tasks. Our model, DeepWriter-8B, trained on this data, not only surpasses strong open-source baselines but also achieves performance competitive with, and at times superior to, leading proprietary models like GPT-4o and Claude 3.5.
Haozhe Wang, Haoran Que, Qixin Xu, Minghao Liu, Wangchunshu Zhou, Jiazhan Feng, Wanjun Zhong, Wei Ye, Tong Yang, Wenhao Huang, Ge Zhang, Fangzhen Lin
arXiv:2509.06160 · cs.AI, cs.CL · submitted Sep 7, 2025
abstract · pdf · html · Preprint
> Using this scalable, gradient-free approach, we curate and open-source DeepWriting-20K, a large-scale dataset of 20,000 deep reasoning trajectories for open-ended tasks. Our model, DeepWriter-8B, trained on this data, not only surpasses strong open-source baselines but also achieves performance competitive with, and at times superior to, leading proprietary models like GPT-4o and Claude 3.5
Could be a game changer. Also and surely if that 8B in "DeepWriter-8B" indicates the NN size (edit: and it certainly should, since DeepWriter-8B is a fine-tuning of Qwen3-8b), and the results are comparable to much bigger models...
Any chance we will be able to try that DeepWriter?
--
Edit: the catch seems to be their «goal is to instill deep reasoning in LLMs for open-ended tasks». If you implement reasoning, the goal is to achieve results that are actually better, getting to "right answers", in problems that while complex are territory for more-optimal and less-optimal solutions.
The question is, does "REverse-Engineered Reasoning" also enhance solution-oriented reasoning, and thought-perfecting reasoning? This is what matters.