about
Tackling the curse of dimensionality with physics-informed neural networks (arxiv.org)
77 points by jhoho on Sep 19, 2023 | hide | past | pdf | 17 comments on HN

In plain words: Instead of checking every dimension while training a neural net to solve physics equations without a grid, this splits the training signal into per-dimension pieces, randomly using a few each step. It solved nonlinear equations with 100,000 dimensions in 12 hours on one GPU.

Abstract · Tackling the Curse of Dimensionality with Physics-Informed Neural Networks

The curse-of-dimensionality taxes computational resources heavily with exponentially increasing computational cost as the dimension increases. This poses great challenges in solving high-dimensional PDEs, as Richard E. Bellman first pointed out over 60 years ago. While there has been some recent success in solving numerically partial differential equations (PDEs) in high dimensions, such computations are prohibitively expensive, and true scaling of general nonlinear PDEs to high dimensions has never been achieved. We develop a new method of scaling up physics-informed neural networks (PINNs) to solve arbitrary high-dimensional PDEs. The new method, called Stochastic Dimension Gradient Descent (SDGD), decomposes a gradient of PDEs into pieces corresponding to different dimensions and randomly samples a subset of these dimensional pieces in each iteration of training PINNs. We prove theoretically the convergence and other desired properties of the proposed method. We demonstrate in various diverse tests that the proposed method can solve many notoriously hard high-dimensional PDEs, including the Hamilton-Jacobi-Bellman (HJB) and the Schrödinger equations in tens of thousands of dimensions very fast on a single GPU using the PINNs mesh-free approach. Notably, we solve nonlinear PDEs with nontrivial, anisotropic, and inseparable solutions in 100,000 effective dimensions in 12 hours on a single GPU using SDGD with PINNs. Since SDGD is a general training methodology of PINNs, it can be applied to any current and future variants of PINNs to scale them up for arbitrary high-dimensional PDEs.

Zheyuan Hu, Khemraj Shukla, George Em Karniadakis, Kenji Kawaguchi
arXiv:2307.12306 · cs.LG, cs.AI, math.DS, math.NA, stat.ML · submitted Jul 23, 2023 · updated May 17, 2024
abstract · pdf · html · Accepted by Neural Networks. Code is available at https://github.com/zheyuanhu01/SDGD_PINN

add comment on HN

> For instance, we solve nontrivial nonlinear PDEs (one HJB equation and one Black-Scholes equation) in 100,000 dimensions in 6 hours on a single GPU using SDGD with PINNs

100,000 dimensions? I thought there were like... 11 tops? https://imagine.gsfc.nasa.gov/science/questions/superstring....

(edit: oops, I misunderstood the context of dimensions here. my bad. thanks)

The way ML people use "dimension", each free parameter is an extra dimension. A high-resolution 2D image is considered to have millions of dimensions - one for each pixel.
That’s not just ML, that’s linear algebra. Moreover, across fields in math, the dimension of an object is defined differently, but it’s always a very fundamental property of an object that captures loosely how complex an object is. Often, that complexity is related to how many numbers you need to describe said object.
These are all expressions of the same concept. The confusion is that physicists are describing the dimensionality of specific systems (space, spacetime, superstring theory, supergravity) - that doesn't mean that this limits the dimensionality of other systems (which is often where confusion lies between laymen).
And each pixel is what, a RGB-vector?
We usually do HxWxC, for height, width, and channels, so each pixel is addressed via the two first dims of the input, and then it has 3 channels. Of course, you can transpose the tensor to CxHxW or CxWxH. Different ordering behaves differently with respect to memory locality.
In the context of neural networks, kind of. It’s 3 numbers on the input layer, but that could influence N parameters in later layers.
Yes, each pixel is a vector of the rasterized vector space.
Thanks for clarifying the dimensions.
Physics uses vectors, which are multi-dimensional and can represent a lot more than 11 dimensions... Machine learning typically uses feature vectors, which are basically just lists of numerical properties.
To add to the other answers: dimensions also appear in higher-order PDE: you typically go from an nth order 1-dimensional equation (an equation mentioning the n-th derivative of your unknown function) to an n-dimensional 1st order equation. There is a general pattern of adding parameters to have several simple problems instead of one complicated.
Iiii do not think you should be able to solve the Schrodinger equation with thousands of dimensions in general on a non-quantum computer, what with that being a quantum-mechanical equation some of whose solutions would reflect quantum-hard problems?
What? The Schrodinger equation is just a linear differential equation. I can solve it here right now by hand if the Hamiltonian is time-independent, even for millions of dimensions.

|psi(t)> = exp(-iHt/hbar) |psi(t=0)>

No offence but this sounds like you've never actually studied any quantum mechanics. The difficulty of solving such equations is purely just the difficulty of solving any complex set of PDEs, as the article talks about. There is nothing about the Schrodinger equation that means it's not computable on a classical computer (EDIT nor would it be easier on a quantum computer).

Calculating the mean field of a thousand electrons is easy for a classical computer. Calculating the exchange and correlation energies of a thousand interacting electrons is not. Quantum computers would have an enormous advantage there.
The time evolution or trotter evolution has some advantages on a quantum computer
There of course may be advantages to specific calculations on quantum computers but that isn't because the Schroedinger equation itself is somehow "a quantum equation"
The example they try it on is a quantum harmonic oscillator, i.e. the potential V(x) is just the squared norm of x. That's analytically solvable. Also,

> in the case of the Schro ̈dinger equation, due to the separable nature of its network structure, we employ a separate neural network for each dimension.

which doesn't sound very general. Admittedly in my quick skim of this paper I didn't follow much.