about
Mercury: Ultra-Fast Language Models Based on Diffusion (arxiv.org)
10 points by simonpure on Jun 29, 2025 | hide | past | pdf | 2 comments on HN

In plain words: Instead of writing one word at a time, Mercury generates many words at once using a diffusion process, like refining a rough draft, for coding tasks. It runs up to 10 times faster than the quickest top models while keeping similar quality.

Abstract

We present Mercury, a new generation of commercial-scale large language models (LLMs) based on diffusion. These models are parameterized via the Transformer architecture and trained to predict multiple tokens in parallel. In this report, we detail Mercury Coder, our first set of diffusion LLMs designed for coding applications. Currently, Mercury Coder comes in two sizes: Mini and Small. These models set a new state-of-the-art on the speed-quality frontier. Based on independent evaluations conducted by Artificial Analysis, Mercury Coder Mini and Mercury Coder Small achieve state-of-the-art throughputs of 1109 tokens/sec and 737 tokens/sec, respectively, on NVIDIA H100 GPUs and outperform speed-optimized frontier models by up to 10x on average while maintaining comparable quality. We discuss additional results on a variety of code benchmarks spanning multiple languages and use-cases as well as real-world validation by developers on Copilot Arena, where the model currently ranks second on quality and is the fastest model overall. We also release a public API at https://platform.inceptionlabs.ai/ and free playground at https://chat.inceptionlabs.ai

Inception Labs, Samar Khanna, Siddhant Kharbanda, Shufan Li, Harshit Varma, Eric Wang, Sawyer Birnbaum, Ziyang Luo, Yanis Miraoui, Akash Palrecha, Stefano Ermon, Aditya Grover, et al.
arXiv:2506.17298 · cs.CL, cs.AI, cs.LG · submitted Jun 17, 2025
abstract · pdf · html · 15 pages; equal core, cross-function, senior authors listed alphabetically

add comment on HN
Also discussed: Jul 2025 (576 points, 242 comments)

Mercury was released at least 59 days ago and precedes Gemini Diffusion: https://news.ycombinator.com/item?id=43851099

You're right though that they weren't the first, e.g. here's a survey paper on text diffusion models from 2023: https://arxiv.org/abs/2303.06574