about
Competitive debate evidence dataset for LLM persuasion (arxiv.org)
2 points by Der_Einzige on Dec 16, 2024 | hide | past | pdf | discuss on HN

In plain words: A free collection of 3.5 million debate evidence documents, with details on each, lets systems learn to find and summarize arguments from high school and college debates. Tests showed that teaching large language models on this material produces good argument summaries.

Abstract · OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset

We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community. This dataset includes over 3.5 million documents with rich metadata, making it one of the most extensive collections of debate evidence. OpenDebateEvidence captures the complexity of arguments in high school and college debates, providing valuable resources for training and evaluation. Our extensive experiments demonstrate the efficacy of fine-tuning state-of-the-art large language models for argumentative abstractive summarization across various methods, models, and datasets. By providing this comprehensive resource, we aim to advance computational argumentation and support practical applications for debaters, educators, and researchers. OpenDebateEvidence is publicly available to support further research and innovation in computational argumentation. Access it here: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized

Allen Roush, Yusuf Shabazz, Arvind Balaji, Peter Zhang, Stefano Mezza, Markus Zhang, Sanjay Basu, Sriram Vishwanath, Mehdi Fatemi, Ravid Shwartz-Ziv
arXiv:2406.14657 · cs.CL, cs.AI, cs.LG · submitted Jun 20, 2024 · updated Aug 2, 2026
abstract · pdf · html · Published to the 38th Conference on Neural Information Processing Systems (NeurIPS 2024) Track on Datasets and Benchmarks

add comment on HN