about
Choosing the "Brain" for your AI-powered app – My new method, feedback requested (arxiv.org)
2 points by robert_sim on Aug 29, 2024 | hide | past | pdf | 1 comment on HN

In plain words: Instead of writing quiz questions and grading answers, this asks a chatbot the same opinion question many times and measures how much its answers vary; less variety means it knows the topic better. That spread matched real quiz rankings 74% to 89% of the time.

Abstract · No Dataset Needed for Downstream Knowledge Benchmarking: Response Dispersion Inversely Correlates with Accuracy on Domain-specific QA

This research seeks to obviate the need for creating QA datasets and grading (chatbot) LLM responses when comparing LLMs' knowledge in specific topic domains. This is done in an entirely end-user centric way without need for access to any inner workings of the LLM, so long as it can be prompted and given a random seed to create different generations to the same prompt. The paper does this by, for a given topic domain, defining the "response dispersion" of an LLM by repeatedly asking an LLM the same opinion question about that topic domain. Namely, the response dispersion is the count of singular values needed to explain 95% of the variance in the embedding matrix of the LLM's responses. It is found that the response dispersion is inversely correlated with accuracy on relevant QA evaluations (average spearman rank correlation stronger than -.59). A use-case analysis shows that when comparing two different LLMs on the same topic domain, comparing their response dispersion is a suitable replacement for comparing their QA accuracy between 74% and 89% of the time, the range depending on certain reasonable accuracy-difference tolerances that may be acceptable to an end-user in exchange for the labor being saved using response dispersion instead of QA accuracy for comparison. Two response embeddings are studied for creating the embedding matrix in this study, one is from OpenAI's APIs and one is a novel embedding, here named reference sentence similarity embeddings, that can be computed locally and performs very nearly as well in calculating response dispersion. Also in this research, a pre-existing dataset called the IRC-Wiki Trivia dataset, originally developed for trivia games, has been re-purposed, curated, and the curation, called IRC-WikiTriviaQA, is made available for the purpose of this research.

Robert L Simione
arXiv:2408.13624 · cs.CL, cs.AI · submitted Aug 24, 2024
abstract · pdf · html · 16 pages, 3 tables, 1 figure

add comment on HN

When building some app or pipeline that has an LLM component at some point (what I mean by "the brain"), there's a question of which model should you choose? Many models nowadays can basically be drop-in replacements for each other. All else being equal, it's better to choose one that's more "knowledgeable" about your application's domain area than not. "Knowledge" of LLMs is usually measured, kind of like giving students a test in school, with a question-answer benchmark dataset where someone makes questions and an answer key. The LLMs are asked the questions, someone has to grade the answers, and whichever LLM scores higher on this test (the QA benchmark) "knows more" about that domain.

HOWEVER, creating the QA benchmark is labor intensive! So what I have created is a procedure you can use that runs programmatically without needing a QA benchmark set at all. Effectively what the procedure is is to probe the LLM with an opinion question in the desired domain many times, and see how consistent or inconsistent its responses are. Fracturing of the answers is quantified by "response dispersion", and the preprint I linked to shows a strong inverse correlation between response dispersion and accuracy on QA benchmark datasets. The point of course is not for people to do the comparison themselves, but to just use response dispersion to get a similar result as they would have if they had instead gone through the entire process of QA benchmark tests.

I'm posting it here because I am requesting constructive criticism on my preprint before submitting it to a journal. The paper itself is geared a little more towards the NLP research community than the average developer, however one of the main products of the paper itself is meant to benefit the average developer who is building AI-powered applications (by which I mean they drop-in an LLM at some point in their pipeline) and wishes for a quick and cheap (nearly-free) way to compare LLMs for his or her application domain.

I will reply to every response here, my responses will be early drafts of improvements I wish to make to the paper, so please criticize them as well. Thank you!