In plain words: Four AI models wrote questions from a passage, compared with human-written ones on six traits like type, length, and coverage. AI questions asked for longer descriptive answers and spread focus across the passage more evenly, while human questions favored certain spots.
Abstract · Can LLMs Ask Good Questions?
We evaluate questions generated by large language models (LLMs) from context, comparing them to human-authored questions across six dimensions: question type, question length, context coverage, answerability, uncommonness, and required answer length. Our study spans two open-source and two proprietary state-of-the-art models. Results reveal that LLM-generated questions tend to demand longer descriptive answers and exhibit more evenly distributed context focus, in contrast to the positional bias often seen in QA tasks. These findings provide insights into the distinctive characteristics of LLM-generated questions and inform future work on question quality and downstream applications.
Yueheng Zhang, Xiaoyuan Liu, Yiyou Sun, Atheer Alharbi, Hend Alzahrani, Tianneng Shi, Basel Alomair, Dawn Song
arXiv:2501.03491 · cs.CL, cs.AI · submitted Jan 7, 2025 · updated Jun 17, 2025
abstract · pdf · html