In plain words: A network repeats one layer until its output settles, but runs it backwards too, so training gets exact gradients instead of rough estimates. This dropped stabilizing tricks and cut extra calculations, while beating similar repeated-layer and fixed-depth models on language and image tasks.
Abstract
Deep Equilibrium Models (DEQs) are an interesting class of implicit model where the model output is implicitly defined as the fixed point of a learned function. These models have been shown to outperform explicit (fixed-depth) models in large-scale tasks by trading many deep layers for a single layer that is iterated many times. However, gradient calculation through DEQs is approximate. This often leads to unstable training dynamics and requires regularisation or many function evaluations to fix. Here, we introduce Reversible Deep Equilibrium Models (RevDEQs) that allow for exact gradient calculation, no regularisation and far fewer function evaluations than DEQs. We show that RevDEQs significantly improve performance on language modelling and image classification tasks against comparable implicit and explicit models.
Sam McCallum, Kamran Arora, James Foster
arXiv:2509.12917 · cs.LG, stat.ML · submitted Sep 16, 2025 · updated Dec 3, 2025
abstract · pdf · html
They use the trick from Feistel ciphers for reversibility.