about
Large language models struggle to learn long-tail knowledge (arxiv.org)
27 points by FiberBundle on Nov 16, 2022 | hide | past | pdf | 2 comments on HN

In plain words: They counted how many training documents mention the same things as a question's answer, then checked how often models got it right. Accuracy rises with how often a fact appears; rare facts need models many orders of magnitude bigger, though looking up documents helps.

Abstract · Large Language Models Struggle to Learn Long-Tail Knowledge

The Internet contains a wealth of knowledge -- from the birthdays of historical figures to tutorials on how to code -- all of which may be learned by language models. However, while certain pieces of information are ubiquitous on the web, others appear extremely rarely. In this paper, we study the relationship between the knowledge memorized by large language models and the information in pre-training datasets scraped from the web. In particular, we show that a language model's ability to answer a fact-based question relates to how many documents associated with that question were seen during pre-training. We identify these relevant documents by entity linking pre-training datasets and counting documents that contain the same entities as a given question-answer pair. Our results demonstrate strong correlational and causal relationships between accuracy and relevant document count for numerous question answering datasets (e.g., TriviaQA), pre-training corpora (e.g., ROOTS), and model sizes (e.g., 176B parameters). Moreover, while larger models are better at learning long-tail knowledge, we estimate that today's models must be scaled by many orders of magnitude to reach competitive QA performance on questions with little support in the pre-training data. Finally, we show that retrieval-augmentation can reduce the dependence on relevant pre-training information, presenting a promising approach for capturing the long-tail.

Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, Colin Raffel
arXiv:2211.08411 · cs.CL, cs.LG · submitted Nov 15, 2022 · updated Jul 27, 2023
abstract · pdf · html · ICML 2023 Camera Ready Version

add comment on HN
Also discussed: Dec 2025 (1 point, 0 comments)

I so cannot wait to use this LLM approach for a simple "50001-001" error message.

I blame app designers (looking at you, HTML, web browsers, CSS, C, IBM, RedHat, Aw shuck, nearly everybody and everything) for adopting such terseness of error codes when troubleshooting servers.

Seems like Google Search has given up on this much needed long-tail search results nowadays (presumably in favor of what it is to me are useless ads).

If you had to make me use ‘strace’ to solve your app error, you failed in error checking.

Seems like self driving cars have the same problem, all the rare edge cases that it is almost impossible to collect training data for. Maybe for cars this is a place where a voice interface would be useful. Instead of having the "driver" of the car be ready to take the wheel, have the system respond to voice commands in order to aid decisions. "Go left", "stop", "take the middle fork" "be careful here" could give the AI more data with which to make the correct decision and handle unusual cases. Maybe enough so to make self driving practical. A simple Green/Yellow/Red dashboard indicator showing an estimate of confidence could tell the driver when a voice command would be helpful. Instrumenting roads with sensors and navigation guidance would also help a lot. Self driving does NOT have to be totally autonomous, you can "cheat". A similar voice UI might have some use in large language models used in practical applications. Essentially a hybrid system, AI plus voice UI plus a few well thought out heuristics.