about
Exploring the Latest LLMs for Leaderboard Extraction (arxiv.org)
1 point by PaulHoule on Jun 18, 2024 | hide | past | pdf | discuss on HN

In plain words: Four AI language models were tested on pulling task, dataset, metric, and score details from research papers, using the setup, the results sections, or the whole text. Each model and text slice showed its own strengths and limits, so none wins everywhere.

Abstract

The rapid advancements in Large Language Models (LLMs) have opened new avenues for automating complex tasks in AI research. This paper investigates the efficacy of different LLMs-Mistral 7B, Llama-2, GPT-4-Turbo and GPT-4.o in extracting leaderboard information from empirical AI research articles. We explore three types of contextual inputs to the models: DocTAET (Document Title, Abstract, Experimental Setup, and Tabular Information), DocREC (Results, Experiments, and Conclusions), and DocFULL (entire document). Our comprehensive study evaluates the performance of these models in generating (Task, Dataset, Metric, Score) quadruples from research papers. The findings reveal significant insights into the strengths and limitations of each model and context type, providing valuable guidance for future AI research automation efforts.

Salomon Kabongo, Jennifer D'Souza, Sören Auer
arXiv:2406.04383 · cs.CL, cs.AI · submitted Jun 6, 2024 · updated Jul 8, 2024
abstract · pdf · html

add comment on HN