about
Cascade Reward Sampling for Efficient Decoding-Time Alignment (arxiv.org)
3 points by Garcia98 on Jul 11, 2024 | hide | past | pdf | discuss on HN

In plain words: Instead of writing a whole answer and then scoring it, this method cuts text into chunks where the scorer is unsure and rejects bad chunks early. It cut decoding time by about 70% while matching or beating full-answer scoring on quality, safety, and usefulness.

Abstract

Aligning large language models (LLMs) with human preferences is essential for their applications. Recently, decoding-time alignment has emerged as an effective plug-and-play technique that avoids fine-tuning model parameters. This approach retains the general utility of pretrained LLMs but often suffers from significant inefficiencies during decoding, primarily due to wasted token generation and excessive reward evaluations. To address these challenges, we introduce Cascade Reward Sampling (CARDS) to resolve both efficiency bottlenecks in decoding-time alignment. Specifically, we develop a segment-level rejection sampling algorithm that minimizes redundant computations of both LLMs and reward models (RMs). Central to CARDS is an uncertainty-based segmentation mechanism, which ensures the accuracy of RMs evaluations on incomplete segments. Furthermore, we provide a detailed analysis of reward scores on segments to elucidate the improved alignment performance. Experimental results demonstrate that CARDS significantly improves decoding efficiency, alignment quality, and general utility compared to existing decoding-time alignment methods, achieving approximately a 70% reduction in decoding time and over 90% win-ties in utility and safety benchmarks.

Bolian Li, Yifan Wang, Anamika Lochab, Ananth Grama, Ruqi Zhang
arXiv:2406.16306 · cs.CL, cs.LG, stat.ML · submitted Jun 24, 2024 · updated Aug 3, 2025
abstract · pdf · html

add comment on HN