about
Improving KAN with CDF normalization to quantiles (arxiv.org)
1 point by jarekd on Jul 23, 2025 | hide | past | pdf | 2 comments on HN

In plain words: Instead of the usual subtract-the-mean, divide-by-the-spread rescaling, each input is mapped to the share of data below it, spreading values evenly from 0 to 1. Just swapping in this rescaling improved a curve-learning network's predictions over its standard version.

Abstract

Data normalization is crucial in machine learning, usually performed by subtracting the mean and dividing by standard deviation, or by rescaling to a fixed range. In copula theory, popular in finance, there is used normalization to approximately quantiles by transforming x to CDF(x) with estimated CDF (cumulative distribution function) to nearly uniform distribution in [0,1], allowing for simpler representations which are less likely to overfit. It seems nearly unknown in machine learning, therefore, we would like to present some its advantages on example of recently popular Kolmogorov-Arnold Networks (KANs), improving predictions from Legendre-KAN by just switching rescaling to CDF normalization. Additionally, in HCR interpretation, weights of such neurons are mixed moments providing local joint distribution models, allow to propagate also probability distributions, and change propagation direction.

Jakub Strawa, Jarek Duda
arXiv:2507.13393 · cs.LG · submitted Jul 16, 2025
abstract · pdf · html · 7 pages, 9 figures

add comment on HN

In ML there is usually used normalization by subtracting the mean and dividing by standard deviation - I haven't seen by CDF in ML (?, they are popular in finance for copulas: https://en.wikipedia.org/wiki/Copula_(statistics) ), which provides more uniform distributions, allowing for better description with smaller models, what seems beneficial for generalization (e.g. description with low degree polynomials in this arXiv).

For which tasks CDF/EDF normalization could be beneficial in ML? Any reasons it seems unknown in ML?

Any other interesting nonstandard normalizations?