about
Deep Delta Learning (arxiv.org)
1 point by mfiguiere 268 days ago | hide | past | pdf | discuss on HN

In plain words: Instead of just adding to a transformer's running state, each layer reads one value, compares it to a target, and writes a gated correction that can overwrite it. Models trained this way beat the usual additive update, the best raising one-shot accuracy by 1.18 points.

Abstract

Transformer residual streams are updated by addition. A sufficiently expressive residual block can represent content replacement, but the residual update itself has no operation that reads, compares, and replaces content. We introduce Deep Delta Learning (DDL), which applies the delta rule over network depth. Each layer reads the residual state along a learned direction, compares the readout with a learned target, and writes a gated rank-1 correction back along the same direction. A closed gate gives the identity map, and a unit gate overwrites the selected readout with the target. DDL works with the usual vector state or with an expanded state that stores several value channels, while attention and MLP blocks keep the original model width. We pretrain decoder-only models with approximately GPT-2 small and medium sizes on the FineWeb-Edu dataset. At both scales, every DDL variant has lower validation loss and higher average one-shot accuracy than the additive baseline, and the best expanded variants raise that average by 0.91 and 1.18 points.

Yifan Zhang, Yifeng Liu, Mengdi Wang, Quanquan Gu
arXiv:2601.00417 · cs.LG, cs.AI, cs.CL, cs.CV · submitted Jan 1, 2026 · updated Sep 25, 2026
abstract · pdf · html · Project Page: https://github.com/yifanzhang-pro/deep-delta-learning

add comment on HN