In plain words: Instead of the usual subtract-the-mean, divide-by-the-spread rescaling, each input is mapped to the share of data below it, spreading values evenly from 0 to 1. Just swapping in this rescaling improved a curve-learning network's predictions over its standard version.
Abstract
Data normalization is crucial in machine learning, usually performed by subtracting the mean and dividing by standard deviation, or by rescaling to a fixed range. In copula theory, popular in finance, there is used normalization to approximately quantiles by transforming x to CDF(x) with estimated CDF (cumulative distribution function) to nearly uniform distribution in [0,1], allowing for simpler representations which are less likely to overfit. It seems nearly unknown in machine learning, therefore, we would like to present some its advantages on example of recently popular Kolmogorov-Arnold Networks (KANs), improving predictions from Legendre-KAN by just switching rescaling to CDF normalization. Additionally, in HCR interpretation, weights of such neurons are mixed moments providing local joint distribution models, allow to propagate also probability distributions, and change propagation direction.
Jakub Strawa, Jarek Duda
arXiv:2507.13393 · cs.LG · submitted Jul 16, 2025
abstract · pdf · html · 7 pages, 9 figures
For which tasks CDF/EDF normalization could be beneficial in ML? Any reasons it seems unknown in ML?
Any other interesting nonstandard normalizations?