about
Matching Principle: Adversarial, augmentation, etc. are estimators of one matrix (arxiv.org)
2 points by visark_42 135 days ago | hide | past | pdf | discuss on HN

In plain words: Training usually ignores how much inner signals wobble, so harmless noise can throw the model off; this adds a penalty for that wobble in directions noise hits. An evenly spread penalty beat plain training across seven domains, and matching the real noise directions worked best.

Abstract · The Matching Principle: When Does a Training Penalty Cover Deployment Shift?

Ordinary training optimises the task loss and then stops. It never pays for internal representation energy: Jacobians can stay large in directions that never helped the label, so even small label-preserving noise throws the model off---a design gap that classical noise-injection theory fixes at second order, but only when applied as default regularisation, which current practice does not do. We make that precise with a Matching Principle: name deployment directions (Sigma_task) and the training penalty Sigma', and ask whether the second covers the first. The no-thinking default is even-spread / isotropic penalty (Sigma' proportional to I)---classical Gaussian / Tikhonov at second order: no axis estimate, no architecture change, and---in a simple linear ridge model---strictly less deployment drift than task-only training, with no coverage miss by construction. When axes are known, matching is sharper; when they are missed, a residual floor remains. Across seven domains a named second-moment penalty beats unregularised training; a controlled illustration recovers match > even-spread > wrong-axis when axes are forced. The ridge theorems are proved; deep nets remain experiments under a specified perturbation. Design rule: fix internal energy by default (even-spread); match when axes are known; treat losses that control representation sensitivity as first-class design.

Vishal Rajput
arXiv:2605.22800 · cs.LG, cs.AI, stat.ML · submitted May 21, 2026 · updated Aug 10, 2026
abstract · pdf · html · 51 pages. Journal-aligned revision of this preprint for JMLR. Title and abstract updated to the even-spread coverage framing. Same author, Matching Principle, and experimental lineage; not a new paper. Under submission at JMLR. Companion: arXiv:2604.21395

add comment on HN