about
Spherical CNNs (2018) (arxiv.org)
22 points by rkp8000 on Jun 16, 2025 | hide | past | pdf | 3 comments on HN

In plain words: Flattening a sphere onto a flat image stretches parts unevenly, so ordinary image filters break. This network's filters slide over the sphere and still work when it rotates, computed with a Fourier trick, and it handled 3D object recognition and molecular energy prediction.

Abstract · Spherical CNNs

Convolutional Neural Networks (CNNs) have become the method of choice for learning problems involving 2D planar images. However, a number of problems of recent interest have created a demand for models that can analyze spherical images. Examples include omnidirectional vision for drones, robots, and autonomous cars, molecular regression problems, and global weather and climate modelling. A naive application of convolutional networks to a planar projection of the spherical signal is destined to fail, because the space-varying distortions introduced by such a projection will make translational weight sharing ineffective. In this paper we introduce the building blocks for constructing spherical CNNs. We propose a definition for the spherical cross-correlation that is both expressive and rotation-equivariant. The spherical correlation satisfies a generalized Fourier theorem, which allows us to compute it efficiently using a generalized (non-commutative) Fast Fourier Transform (FFT) algorithm. We demonstrate the computational efficiency, numerical accuracy, and effectiveness of spherical CNNs applied to 3D model recognition and atomization energy regression.

Taco S. Cohen, Mario Geiger, Jonas Koehler, Max Welling
arXiv:1801.10130 · cs.LG, stat.ML · submitted Jan 30, 2018 · updated Feb 25, 2018
abstract · pdf · html · Proceedings of the 6th International Conference on Learning Representations (ICLR), 2018

add comment on HN

Great subject, thanks. I recently built SpinStep[0], a tool for visualizing and stepping through SCNN computations.

It lets you upload a model, then see—layer by layer—how inputs are transformed, which kernels activate, and how feature maps evolve. It's a hands-on exploration of what’s actually happening under the hood in Spherical CNNs.

For anyone who's been frustrated by the opaque "black‑box" nature of CNNs, SpinStep might be a fun way to poke around and build intuition.

[0] https://github.com/VoxleOne/SpinStep/blob/main/docs/index.md

A notable and interesting point of this article is that convolutions and correlations (convolutions without flipping the filter) are quite a bit more subtle on the sphere than on Cartesian spaces. For a convolution between a function and a filter on R^N you just "slide" the filter around, integrating at each shift, which produces another function on R^N. On a sphere, however, there is not a clear cut way to slide a filter around a sphere. For instance, there are multiple ways to slide a filter centered at the north pole to the south pole, which will result in different filter orientations.

More generally, the space of rotations, which is the argument of the convolution (analogous to the shift amount being the argument of a standard convolution), is 3D (3 Euler angles), whereas the space of points on the sphere is 2D (polar and azimuthal angles). Thus, whereas convolution over R^N returns a function over R^N, convolution over the sphere actually returns a function over the 3D rotation group SO(3). This has interesting consequences for e.g. the convolution theorem on the sphere, which is not as clear cut as simply rewriting the standard convolution theorem in spherical terms.

also relevant to 3d modeling of molecules using in drug discovering modeling -- and a subset of these authors have published along those lines more recently -- for e.g. https://arxiv.org/abs/2104.13478