In plain words: Instead of fixed activation functions on neurons with plain weighted links, this network puts a small learnable curve on every connection. Much smaller versions match or beat much larger standard networks at fitting data and solving equations, and are easier to visualize and understand.
Abstract · KAN: Kolmogorov-Arnold Networks
Inspired by the Kolmogorov-Arnold representation theorem, we propose Kolmogorov-Arnold Networks (KANs) as promising alternatives to Multi-Layer Perceptrons (MLPs). While MLPs have fixed activation functions on nodes ("neurons"), KANs have learnable activation functions on edges ("weights"). KANs have no linear weights at all -- every weight parameter is replaced by a univariate function parametrized as a spline. We show that this seemingly simple change makes KANs outperform MLPs in terms of accuracy and interpretability. For accuracy, much smaller KANs can achieve comparable or better accuracy than much larger MLPs in data fitting and PDE solving. Theoretically and empirically, KANs possess faster neural scaling laws than MLPs. For interpretability, KANs can be intuitively visualized and can easily interact with human users. Through two examples in mathematics and physics, KANs are shown to be useful collaborators helping scientists (re)discover mathematical and physical laws. In summary, KANs are promising alternatives for MLPs, opening opportunities for further improving today's deep learning models which rely heavily on MLPs.
Ziming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle, James Halverson, Marin Soljačić, Thomas Y. Hou, Max Tegmark
arXiv:2404.19756 · cs.LG, cond-mat.dis-nn, cs.AI, stat.ML · submitted Apr 30, 2024 · updated Feb 9, 2025
abstract · pdf · html · Accepted by International Conference on Learning Representations (ICLR) 2025 (conference version: https://openreview.net/forum?id=Ozo7qJ5vZi). Codes are available at https://github.com/KindXiaoming/pykan
Overall the idea, if I understand it, is that rather than using linear activations between layers in a deep learning setup (so called MLP - Multi-Layer Perceptrons), you use functions ("splines") which can be learned using some sort of backprop at the nodes.
The idea (well, one of many) is that using more complex functions and making them learnable during training means you can encapsulate more complex functions in a lower dimensionality matrix.
Upsides: they claim smaller number of training steps and lower dimensionality for a number of toy problems. Also, since you're training functions and these functions might have periodicity to them, and that periodicity can be adjusted during training, and that periodicity might tie to sets of data with varying states, they claim you can update parts of the periodic functions for specific information without impacting the whole function, and therefore this is a better conceptual architecture for dealing with special cases, "forgetting" and so on.
Downsides: they mention it's much slower to train than an MLP architecture, although they claim they didn't try very hard to optimize.
My totally uninformed complaints: The networks are trained on REALLY simple problems, like approximating x*y. I mean this is a fairly beautifully written and illustrated 40+ page paper, with really quite a lot of LaTeX, and I (mostly) read the whole thing. It feels hard to read it without at least seeing, like, some MNIST training results.
Upshot - I'm not sure that we'll all be training Kan networks in 2024.