about
Discrete Tilt Matching (arxiv.org)
3 points by PaulHoule 120 days ago | hide | past | pdf | discuss on HN

In plain words: It fine-tunes blank-filling language models by nudging each fill-in choice toward higher-reward answers, using a simple weighted loss instead of the usual sequence-level probabilities that are too hard to compute. Training this way improved Sudoku and Countdown while staying competitive on two math benchmarks.

Abstract

Masked diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. While reinforcement learning (RL) methods have recently been adapted to dLLM fine-tuning, their objectives typically depend on sequence-level marginal likelihoods, which are intractable for masked diffusion models. To address this, we derive Discrete Tilt Matching (DTM), a likelihood-free method that recasts dLLM fine-tuning as state-level matching of local unmasking posteriors under reward tilting. DTM takes the form of a weighted cross-entropy objective with explicit minimizer, and admits control variates that improve training stability. On a synthetic maze-planning task, we analyze how DTM's annealing schedule and control variates affect training stability and prevent mode collapse. At scale, fine-tuning LLaDA-8B-Instruct with DTM yields strong gains on Sudoku and Countdown while remaining competitive on MATH500 and GSM8K.

Yuyuan Chen, Shiyi Wang, Peter Potaptchik, Jaeyeon Kim, Michael S. Albergo
arXiv:2604.18739 · cs.LG, stat.ML · submitted Apr 20, 2026 · updated May 19, 2026
abstract · pdf · html

add comment on HN