about
Why not to use the Gaussian kernel (arxiv.org)
2 points by E-Reverance 36 days ago | hide | past | pdf | discuss on HN

In plain words: The Gaussian kernel, a popular similarity measure in regression, is so unnaturally smooth that its uncertainty estimates come out far too small and its calculations become unstable. This brittleness is almost inevitable, so this kernel and any equally smooth one should never be the default.

Abstract

Kernels measure similarity or correlation in tasks such as regression and classification. The Gaussian kernel, other names of which include squared exponential and radial basis function kernel, is one of the most popular in Gaussian process regression. We argue that the Gaussian kernel is best avoided and should never be used as a default. The argument rests on two results demonstrating that the Gaussian kernel is extremely brittle. First, the Gaussian kernel gives rise to a conditional variance that is unrealistically small. If the variance is used to quantify predictive uncertainty, catastrophic overconfidence is almost inevitable. Second, a small variance goes hand in hand with numerical ill-conditioning, so that to use the Gaussian kernel in practice requires tricks such as nugget terms that effectively modify the underlying regression or classification model. These problems are caused by the unnatural smoothness of the Gaussian kernel, a fact we are far from the first to take notice of. The problem is not the Gaussian form itself but the analyticity of the kernel: Our argument is more broadly that analytic kernels are best avoided. For stationary kernels analyticity is essentially equivalent to an exponential decay of the spectral density.

Toni Karvonen, Chris J. Oates
arXiv:2608.26974 · stat.ML, cs.LG, math.NA, math.ST · submitted Aug 27, 2026
abstract · pdf

add comment on HN