about
Challenges with unsupervised LLM knowledge discovery (arxiv.org)
2 points by westurner on Dec 24, 2023 | hide | past | pdf | 3 comments on HN

In plain words: They checked a popular trick for pulling hidden knowledge out of an AI's internal signals, which hunts for patterns that stay consistent across questions. It actually grabs whatever feature stands out most, so answers often track some other trait instead of what the AI knows.

Abstract

We show that existing unsupervised methods on large language model (LLM) activations do not discover knowledge -- instead they seem to discover whatever feature of the activations is most prominent. The idea behind unsupervised knowledge elicitation is that knowledge satisfies a consistency structure, which can be used to discover knowledge. We first prove theoretically that arbitrary features (not just knowledge) satisfy the consistency structure of a particular leading unsupervised knowledge-elicitation method, contrast-consistent search (Burns et al. - arXiv:2212.03827). We then present a series of experiments showing settings in which unsupervised methods result in classifiers that do not predict knowledge, but instead predict a different prominent feature. We conclude that existing unsupervised methods for discovering latent knowledge are insufficient, and we contribute sanity checks to apply to evaluating future knowledge elicitation methods. Conceptually, we hypothesise that the identification issues explored here, e.g. distinguishing a model's knowledge from that of a simulated character's, will persist for future unsupervised methods.

Sebastian Farquhar, Vikrant Varma, Zachary Kenton, Johannes Gasteiger, Vladimir Mikulik, Rohin Shah
arXiv:2312.10029 · cs.LG, cs.AI · submitted Dec 15, 2023 · updated Dec 18, 2023
abstract · pdf · html · 12 pages (38 including references and appendices). First three authors equal contribution, randomised order

add comment on HN

"Challenges with unsupervised LLM knowledge discovery" (2023) https://arxiv.org/abs/2312.10029 :

> Abstract: We show that existing unsupervised methods on large language model (LLM) activations do not discover knowledge -- instead they seem to discover whatever feature of the activations is most prominent.

- NewsArticle about / that repeats a ScholarlyArtice: "A New Research from Google DeepMind Challenges the Effectiveness of Unsupervised Machine Learning Methods in Knowledge Elicitation from Large Language Models" (2023) https://www.marktechpost.com/2023/12/20/a-new-research-from-...
Do LLMs converge upon the same answer when asked the same question multiple times?

If they do not converge, are any of the replies true?

Or do they have high attention values?