about
Ministral 3 – pruning via Cascade Distillation (arxiv.org)
5 points by everlier 263 days ago | hide | past | pdf | discuss on HN

In plain words: Instead of training small models from scratch, a larger model is repeatedly trimmed and retrained to copy its answers, shrinking it step by step. This yields small models in three sizes for limited devices, each in base, instruction, and reasoning versions with image understanding.

Abstract · Ministral 3

We introduce the Ministral 3 series, a family of parameter-efficient dense language models designed for compute and memory constrained applications, available in three model sizes: 3B, 8B, and 14B parameters. For each model size, we release three variants: a pretrained base model for general-purpose use, an instruction finetuned, and a reasoning model for complex problem-solving. In addition, we present our recipe to derive the Ministral 3 models through Cascade Distillation, an iterative pruning and continued training with distillation technique. Each model comes with image understanding capabilities, all under the Apache 2.0 license.

Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian, Victor Jouault, Abhinav Rastogi, Adrien Sadé, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, Alexandre Sablayrolles, Amélie Héliou, et al.
arXiv:2601.08584 · cs.CL · submitted Jan 13, 2026
abstract · pdf · html · Release page: https://mistral.ai/news/mistral-3 ; Models available at https://huggingface.co/collections/mistralai/ministral-3

add comment on HN