about
Source-Optimal Training Is Transfer-Suboptimal (arxiv.org)
1 point by ceh123 325 days ago | hide | past | pdf | 1 comment on HN

In plain words: Training a source model to fit its own task as well as possible isn't best for transfer: in a linear case, the ideal penalty for transfer differs from the source-best choice. Partly aligned tasks favor stronger penalties for transfer; unusually good alignment favors weaker.

Abstract · Source-Optimal Training is Transfer-Suboptimal

We prove that training a source model optimally for its own task is generically suboptimal when the objective is downstream transfer. We study the source-side optimization problem in L2-SP ridge regression and show a fundamental mismatch between the source-optimal and transfer-optimal source regularization: outside of a measure-zero set, $τ_0^* \neq τ_S^*$. We characterize the transfer-optimal source penalty $τ_0^*$ as a function of task alignment and identify an alignment-dependent reversal: with imperfect alignment ($0<ρ<1$), transfer benefits from stronger source regularization, while in super-aligned regimes ($ρ>1$), transfer benefits from weaker regularization. Additionally, in isotropic settings, the decision of whether transfer helps is independent of the target sample size and noise, depending only on task alignment and source characteristics. We verify the linear predictions in a synthetic ridge regression experiment, and we present experiments on MNIST, CIFAR-10, and 20 Newsgroups as evidence that the source-optimal versus transfer-optimal mismatch persists in standard nonlinear transfer learning pipelines.

C. Evans Hedges
arXiv:2511.08401 · stat.ML, cs.LG, math.ST · submitted Nov 11, 2025 · updated Jan 6, 2026
abstract · pdf · html

add comment on HN

This paper is a theoretical analysis showing that the ridge regularization that optimizes the source task almost never optimizes transfer performance. Interestingly, in high SNR regimes (low noise) the optimal regularization for pre-training is higher than the task specific optimal regularization, and in low SNR regimes (high noise) it’s better to regularize less than you would if you were just optimizing for that task.

Although the proofs are in the world of (L2-SP) ridge regression, experiments were run using an MLP on MNIST and CNN on CIFAR-10 and suggest the SNR-regularization relationship persists in non-linear networks.