about
Bolt: Bootstrap long chain-of-thought in LLMs without distillation [pdf] (arxiv.org)
15 points by TaurenHunter on Feb 8, 2025 | hide | past | pdf | 5 comments on HN

In plain words: Instead of copying long reasoning from models like o1, this method makes an instruct model write its own long reasoning from just 10 examples, then trains on it. It worked well on math, logic, and chat tests at three model sizes, needing no stronger teacher.

Abstract · BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation

Large language models (LLMs), such as o1 from OpenAI, have demonstrated remarkable reasoning capabilities. o1 generates a long chain-of-thought (LongCoT) before answering a question. LongCoT allows LLMs to analyze problems, devise plans, reflect, and backtrack effectively. These actions empower LLM to solve complex problems. After the release of o1, many teams have attempted to replicate its LongCoT and reasoning capabilities. In terms of methods, they primarily rely on knowledge distillation with data from existing models with LongCoT capacities (e.g., OpenAI-o1, Qwen-QwQ, DeepSeek-R1-Preview), leaving significant uncertainties on systematically developing such reasoning abilities. In terms of data domains, these works focus narrowly on math while a few others include coding, limiting their generalizability. This paper introduces a novel approach to enable LLM's LongCoT capacity without distillation from o1-like models or expensive human annotations, where we bootstrap LongCoT (BOLT) from a standard instruct model. BOLT involves three stages: 1) LongCoT data bootstrapping with in-context learning on a standard instruct model; 2) LongCoT supervised finetuning; 3) online training to further refine LongCoT capacities. In BOLT, only a few in-context examples need to be constructed during the bootstrapping stage; in our experiments, we created 10 examples, demonstrating the feasibility of this approach. We use Llama-3.1-70B-Instruct to bootstrap LongCoT and apply our method to various model scales (7B, 8B, 70B). We achieve impressive performance on a variety of benchmarks, Arena-Hard, MT-Bench, WildBench, ZebraLogic, MATH500, which evaluate diverse task-solving and reasoning capabilities.

Bo Pang, Hanze Dong, Jiacheng Xu, Silvio Savarese, Yingbo Zhou, Caiming Xiong
arXiv:2502.03860 · cs.CL · submitted Feb 6, 2025
abstract · pdf · html · 36 pages

add comment on HN

Can someone explain what distillation is exactly? I keep seeing people posting comments here and elsewhere about how DeepSeek “distilled” OpenAI outputs to train their new model. How could that even work - wouldn’t you need to ask millions of questions to get enough data to be able to train a whole another LLM? Or am I just uneducated about this topic?
First, I don’t believe there has been a shred of evidence that DeepSeek distilled OpenAI’s model.

As to distillation: people use the term kind of imprecisely. The most powerful form of distillation is one in which you train a smaller “student” model on a large amount of predictions from a larger, more powerful “teacher” model that uses the same tokenization scheme. You train the smaller model to output not just the same tokens as the teacher, but rather on the full probability distribution predicted by that model for each token. It’s a very dense transfer of knowledge. This is best done using the same training data that the teacher model was trained on, since you are asking the student to learn the teacher’s distribution.

The more common version is to take an existing model and fine-tune it on a small number of text outputs from a larger model. No probability distributions over tokens - just the text itself. This doesn’t significantly alter the smaller model’s knowledge or capabilities, but rather, causes it to imitate the other model’s style. But it can unlock capabilities that were in the small model to begin with, but were not as accessible in its base form. For this type of distillation, it has been shown that a very small number of training examples are required if they are carefully selected.

https://huggingface.co/docs/trl/main/en/gkd_trainer

Take a set of prompts, run it through the large model and small model, calculate KL divergence between all the logits, then update the small model to minimize the loss.

This gives the smaller model a much higher density signal than SFT which is usually cross entropy against the single correct logit.

Edit: note that the term is abused quite often and SFT on just the traces from the big model is also sometimes referred to as distillation (Eg the Deepseek smaller models). IMO that is incorrect.

in less technical terms: just run some prompts through original model, and fine tune smaller to respond in the same way
I believe the number needed is not millions but more like thousands or even hundreds depending on the size of the model you are distilling into.