about
Bridging empirical-theoretical gap in neural network formal language learning (arxiv.org)
68 points by puttycat on Feb 22, 2024 | hide | past | pdf | 31 comments on HN

In plain words: Training a network on one simple formal language, the study checks whether the mathematically correct rule is the best answer under usual training goals. It is not, even with weight penalties or dropout, but rewarding the shortest description of the data makes it the optimum.

Abstract · Bridging the Empirical-Theoretical Gap in Neural Network Formal Language Learning Using Minimum Description Length

Neural networks offer good approximation to many tasks but consistently fail to reach perfect generalization, even when theoretical work shows that such perfect solutions can be expressed by certain architectures. Using the task of formal language learning, we focus on one simple formal language and show that the theoretically correct solution is in fact not an optimum of commonly used objectives -- even with regularization techniques that according to common wisdom should lead to simple weights and good generalization (L1, L2) or other meta-heuristics (early-stopping, dropout). On the other hand, replacing standard targets with the Minimum Description Length objective (MDL) results in the correct solution being an optimum.

Nur Lan, Emmanuel Chemla, Roni Katzir
arXiv:2402.10013 · cs.CL, cs.FL · submitted Feb 15, 2024 · updated Jun 6, 2024
abstract · pdf · html · 9 pages, 5 figures, 3 appendix pages

add comment on HN
Also discussed: Feb 2024 (1 point, 0 comments)

To summarize the article: using backpropagation to train a small LSTM to accept strings of the form (a^n b^n), regularization with both L1 and L2 losses result in loss spaces where the loss-optimal solution is not the generally-correct one. The resulting networks fail out-of-sample.

The LSTM architecture does permit a correct, 'general' solution to exist, and the authors show that the general solution is an optimum when using a minimum-description-length error function. The MDL error function is effectively entropy(weights) + relative entropy(training data | model).

The researcher's problem is that entropy(weights) is not well-defined. To prevent the network from "smuggling" information through highly precise but otherwise random-looking weights, they define an entropy (coding length) that rewards simple rational fractions (1/2, 3/4, 7/11) and penalizes complex ones (715/937). With this loss function, their hand-crafted optimal solution also lies at a loss-optimum.

Unfortunately, the resulting loss space is non-differentiable, making it useless for any gradient-descent-like training.

To muse about this, the authors' problem of weight entropy here is similar to the problem faced by autoencoders, whereby they want to have a minimally-structured latent space to avoid the same information-smuggling problem. Perhaps we could evolve a differentiable version of the model-entropy loss function by treating the model weights as random variables, drawn from a distribution of learned mean and variance with the same 'reparameterization trick' that makes variational autoencoders work. The model entropy loss term is then the same as the latent-space regularization term in VAE training.

>The LSTM architecture does permit a correct, 'general' solution to exist, and the authors show that the general solution is an optimum when using a minimum-description-length error function. The MDL error function is effectively entropy(weights) + relative entropy(training data | model).

>The researcher's problem is that entropy(weights) is not well-defined.

I never thought I would say this, but it sounds like a job for Bayesian neural nets. The prior or posterior entropy of the weights should be well-defined in that setting.

If you are interested in this field (exploring the limits of neural models using formal language theory), I help run a weekly seminar on it, Formal Languages and Neural Networks:

https://flann.super.site/

We have had many great speakers (most of them are recorded and available on YouTube) and have a welcoming Discord.

The RNN in question is over-parameterized, so most loss functions wouldn't be expected to coincide with MDL (unless maybe you added in heavier regularization than just L1/L2). There is no incentive towards simplicity. But it is also a single-task network, and it is increasingly rare to train any of those. A multi-task network must share its weights across all the tasks, and so the more tasks it does for a particular parameterization, the more incentive it has towards simplicity in order to do the tasks at all. So I wonder if scaling up task diversity would increasingly approximate MDL and presumably help it learn the true algorithmic cores?
And yet, ChatGPT can generate these strings. Somehow despite using the wrong loss function it still seems to work by simply absorbing more training data.

https://chat.openai.com/share/82509815-d418-43bb-95a3-348bd5...

It can also recognize them, albeit it tends to cheat by shelling out to python (which makes sense, since it tends to lose count on large strings just like a human...)

https://chat.openai.com/share/b106ca5f-409a-43db-bc02-21da86...

ChatGPT incorrectly put a space between the as and bs, but if we let that slide, there's still the issue that the best trained model in the article got 77.3% of the first 1500 strings correct, i.e. even if ChatGPT performed exactly the same, you'd expect it to get a single example correct more often than not.
The difference is that you told it the language you wanted to recognize. In the paper, they are trying to learn the language from example strings alone.
Arguably the skill to generate a program to do this is a higher-order skill and much more impressive.
The issue is that language comes after we form the concepts mentally that the words refer to from our space-time experience. So the language itself is just the token used to symbolize a concept or entity etc that is mentally modeled by the speaker. So the tokens themselves give you no actual knowledge of the world without being able to convert those tokens into the mental objects they represent and then think about them. This is why LLM tech will never improve as its just statistics on serialized symbols, rather than about the world model those symbols represent.
This isn't what the paper gives evidence for. They found that if you use a different function in the training phase, the networks learn formal languages perfectly well. It's about the capabilities of a particular training process, not about the capabilities of the neural networks themselves.
I'm not sure what you mean by the system learning "perfectly well" as it literally knows nothing about what those words mean. This is why LLMs can give statistical word results that can be true of false or in between and those systems have no way of knowing which status the result has as the system has no understanding - that is, it does not model what the words mean as we humans do to understand if something is true or false etc.
> as it literally knows nothing about what those words mean

Speculation. We don't know what it means, mechanistically, to know what something means. Knowing what something means could consist of having a model of how a word relates to other things, in which case the system does indeed know something about what those words mean. It doesn't have precisely the same meaning as humans, because humans also relate words to other sense data, but there's still meaning.

Sure we do. When you read a sentence, and are able to model that in your mind and understand the space-time being referred to, you understand it. LLM is like a calculator in the sense that given an input it gives an output. There is no modeling of space-time in the mind to know if the sentence makes sense etc. That's why if you read a language you don't understand, you get nothing as you can't link those words with the concepts in your mind to understand it. For the LLM, all languages are like that as it understands nothing.
> For the LLM, all languages are like that as it understands nothing.

Again, that's too far. An LLM will not assert that "married men are bachelors" because it's been trained that "bachelor" is not associated with "married". Even if it doesn't understand what a bachelor is or marriage is, it understands that these two concepts are not positively associated with each other, and in fact, are negatively correlated. Logical associations like this are semantic content, which thus exhibits some degree of understanding. That's why LLMs can make sense, even when they're lacking the full understanding of a human.

I'm not arguing that they are not fantastic statistics machines, but there is nothing like semantics or "two concepts" for LLM, its just word strings that statistically occur together, etc. This is precisely why they are so convincingly wrong, wright, or in between, and the technology has no way to ever "know" the difference.
Again, we don't have a precise understanding of what distinguishes syntax and semantics, or whether there's any difference at all. You're just asserting that logical relationships are not semantics, but that's not a proof, nor are fallacious thought experiments like Searle's Chinese Room legit proofs that this distinction is meaningful.

At the end of the day, semantics are about making meaningful distinctions. If all I gave you was a formula that related entities X and Y, but you didn't specifically know what X and Y represented, it is simply not correct to say that you don't know anything about X and Y despite not knowing what they are.

These topics are awesome- they go to the heart of what it means to be human. With that, I'm not so sure your premise of "semantics are meaningful distinctions."

In my mental model of an apple, I envision an apple against the backdrop of a void. The only existing relationship is "apple" to "not apple." This is a meaningful property, but I don't think it means I have any clue about what an apple is.

I wonder about asking chatgpt "what is two plus two" a billion times or whatever. Given errors in the training data, and the negative correlation to answers that are not "four" is only a probability, will it tell me two plus two equals five? My gut says yes. And yes, humans could misspeak, but we can also self correct.

"Apple-not apple" also makes me doubt adding space-time to the mix is enough for actual understanding.

> With that, I'm not so sure your premise of "semantics are meaningful distinctions."

That's a simplistic definition for introductory purposes. I think a more precise definition is that syntax consists of local properties of X, and semantics are broader, possibly global properties of X, eg. how X is related to Y, Z, A, B, etc.

For any real world object, semantic properties of an apple pull in all sorts of sophisticated relationships depending on the perspective, eg. redness, taste, crunch if you're hungry, genes if you're a biologist, particles if you're a physicist, etc. But in the end, we're all just locally acting particles and fields (syntax) lacking any semantic properties, and so semantics must be emerge from syntax. It's the only thing that makes sense in a mechanistic universe.

It seems clear that LLMs do build a model relating X, Y, Z, A, B, etc., it's simply lacking some relations, namely, the sensory relations that humans have. That's a lot of missing data, which is one of the reasons why LLMs make mistakes.

> My gut says yes. And yes, humans could misspeak, but we can also self correct.

Yes, LLMs don't typically have review-loops, where they check what they were about to say before spitting out an answer. But if you explicitly ask it to review its answer for any errors, it's accuracy goes up noticeably. There are plenty of follow-up papers that add this sort of review automatically and show considerable benefits.

There's still probably still some form of generalization that's missing to make best use of training data, but I still think it's wrong to say that LLMs lack any understanding. I think this is a good analysis of the situation:

https://reddit.com/r/naturalism/comments/1236vzf/on_large_la...

Interesting points, and thanks for the link!
The words in the article look like #aaaaaaabbbbbbb#, they do not mean anything.
You appear to know literally nothing about what's in the article.
If you could, would you mind pointing us to some foundational theory on this that actually proves (preferably with analytical expressions or some other form of mathematics) anything around these statements?

It sounds like you read some of Emily Bender, et al's works and then simply regurgitated them. Unfortunately, everyone in this camp simply has large walls of text for their arguments, and basically no data, proofs, or any evidence whatsoever. The 'evidence' around "system X doesn't understand, like I do" amount to 'I don't like this' and 'trust me bro'.

In formal lingiistics, Chomsky has a point against ML for LMs, which is based around the lack of interpretability. His works did push our understanding forward - because they included formal mathematics in proving his statements. However, they tended to be more specific about what limits exist around expressivity, rather than your issues laid out here.

There are strong arguments to say that we do learn and represent language in a similar way to these NNs, at least at some level. 1) the distributional hypothesis and 'words defined by company they keep' is intuitive for how we learn language, and has provided enormous advances in NLP starting with w2v. 2) Dozens of academic papers from around the globe in different labs all reproducing the result in various ways that the distributional hypothesis word vectors (either contextual as in Bert/Gpt, or not as in gensim) map to brain activations by fMRI. This result also tracks in more resolute experiments in visual modalities for more invasive experiments in non-human primates.

I could go on providing data and experiments, but it would be nice if the opposing side to this would provide any at all that were convincing.

My basis would be the conceptual AI on top of space-time my company has built: https://graphmetrix.com/

I started the epistemological research in the late 90's for understanding how human concepts work and started building it 5 years ago. (In Common Lisp)

Putting aside the fact that everything you said is wrong or unsubstantiated, LLMs don’t work like humans because they can learn languages/grammar that humans cannot.
You haven't provided any counter argument, and there are many articles with data backing this up.

Unfortunately I am not aware of any articles that can show convincing data that "LLMs don't learn like humans". I'm not really even sure what that means precisely.

Your statement could be understood to claim that language models cannot predict brain activations or vice versa? Predictive power is the foundation of science, so it's one of the better tests we have for these types of problems. To that point, there is substantial evidence of similarity and predictive power, as noted below.

Perhaps you mean something else? Surely you're point is not that GPUs simply aren't human. Of course - there are differences at various levels. So at some level, everyone recognizes they're not identical. Perhaps it's that attention architectures are not reproduced within neural microarchitectures in tbe neocortex? I haven't seen any studies on this within the relevant cortical areas to address this, though microarchitecture matching may not be required for output distribution matching. The interesting endeavour for science to show right now is how these systems are similar to our processing mechanisms. The ways in which they're not similar typically tend to be uninteresting, or obvious.

The fact that we can predict brain activity from NN model activations raises many more opportunities in science, and the predictive power has gotten far better with these advances,.

Here are a small set of articles from MIT, deepmind, etc. for some of the points I made:

  -  Tuckute, Greta, Aalok Sathe, Shashank Srikant, Maya Taliaferro, Mingye Wang, Martin Schrimpf, Kendrick Kay, and Evelina Fedorenko. "Driving and suppressing the human language network using large language models." Nature Human Behaviour (2024): 1-18.

  - Schrimpf, Martin, Jonas Kubilius, Ha Hong, Najib J. Majaj, Rishi Rajalingham, Elias B. Issa, Kohitij Kar et al. "Brain-score: Which artificial neural network for object recognition is most brain-like?." BioRxiv (2018): 407007.

  - Pereira, Francisco, Bin Lou, Brianna Pritchett, Samuel Ritter, Samuel J. Gershman, Nancy Kanwisher, Matthew Botvinick, and Evelina Fedorenko. "Toward a universal decoder of linguistic meaning from brain activation." Nature communications 9, no. 1 (2018): 963.

  - Arend, Luke, Yena Han, Martin Schrimpf, Pouya Bashivan, Kohitij Kar, Tomaso Poggio, James J. DiCarlo, and Xavier Boix. Single units in a deep neural network functionally correspond with neurons in the brain: preliminary results. Center for Brains, Minds and Machines (CBMM), 2018.
I did provide a counter argument and referenced evidence "backing it up": Namely LLMs can learn languages that are not human-compatible (Moro, "Secrets of Words").

I seriously doubt you have read, never mind understood, any of the papers you cite; if you had we could discuss what's wrong with them. Without even having read them there are obvious problems like the fact that you can't measure a large number of individual neuron activation in the brain and that different deep learning networks have substantially different activations so they cannot be meaningfully be similar to the brain if they aren’t even similar to one another (or that the similarity is so nebulous as to be meaningless).

But this is moot because as I said at the outset there are fundamental differences at the functional level which make existing LLMs (and deep learning) unlike brains.

I think we may agree more than you realize, but perhaps we aren't clear about what we mean. As I said pretty early on - they are similar at some level. To make an analogy, that could be as simple as a submarine (w2v) and a fish (authentic language) now 'uses fins' to move (bert). They're now more similar, but not the same. Of course there are differences, but we can learn by studying what structure has been added and how in order to work towards better theory (as well as looking at what differences there are). I have read all of these studies, as I am an academic in related fields, and you are correct that the resolution is not as good as we would hope (it's better in the ventral visual stream in non human primates, which is one reason we started there); however it doesn't negate the importance of the predictive power here. We can make very useful tools out of these models which can improve the lives of non-verbal people for example.

Another interesting set of works going in recently is in testing out the poverty of stimulus. There is now evidence from many labs that LMs learn with the same amount of data as humans, so the sample efficiency arguments and the 'poverty of stimulus' aren't really as successful as arguments for training difference either.

To the point of humans not learning certain languages, I haven't seen it proven that humans are absolutely incapable of learning certain languages. I do know that children ignore much of sequence structure and create their own internal structures, but this does not negate learning languages which have some arbitrary set of rules. I'm aware that models learn DNA more easily _practically_, but this is more of a statement about human experience and various other environmental factors - not a mathematical proof that it's impossible for humans to learn. Furthermore, that is a nice property of these systems. If they are capable of modeling many different phenomenon, they are useful for many things. If they can settle on a model configuration that linearly maps to a similar activation space with similar topological properties to a patient, that's also very useful. In other words, If the topological properties of the activation space are more constrained in the cortex vs more flexible NNs, but the NNs can fit to the constrained space, that doesn't seem like an insurmountable problem in studying either trained system.

Ultimately, I think we can agree that there are differences, but there are also striking similarities that can be very useful for improving science and medicine going forward, and perhaps (with _substantial_ effort in linguistics and mechanistic interpretability fields) we may be able to improve some of our understanding of linguistics and/or neuroscience (or perhaps not, but it's at least a promising potential lead).

> I haven't seen it proven that humans are absolutely incapable of learning certain languages.

I cited Moro who showed it in a series of experiments.

> they are useful many things

The fact that they are much more general than humans makes them not restrictive enough to be models of the human language faculty.

[edit: title changed]
> However, replacing standard targets with the Minimum Description Length objective (MDL) results in the correct solution being an optimum.

The paper is about how the "default" training doesn't work, not how neural networks "can't learn" something.

Indeed. I read it as what loss function works better for this use case.