In plain words: A tool sends each paper to five separate expert-style reviewers, then a final stage argues their critiques together to name the core mechanism, hidden assumptions, and wider impact. Human researchers preferred its critique to human-written ones in 15 of 20 head-to-head comparisons.
Abstract · Can LLMs Perform Technical Comprehension of Computer Architecture Papers?
Can large language models perform technical comprehension of computer architecture papers--not summarization, but structured critique that names the core mechanism, surfaces buried assumptions, and connects a contribution beyond its own scope? We study Gauntlet, an open-source pipeline that analyzes a paper through five independent expert-persona reviewers and an adversarial synthesis stage. On 20 ISCA 2025 and HPCA 2026 papers, 10 researchers each wrote their own analyses and then judged, for papers other than their own, the human analysis against Gauntlet's. Across the 20 comparisons evaluators preferred Gauntlet in 15 (human in 4, one tie); its advantage is significant on per-analyst totals (two-sided Wilcoxon, p < 0.001) and largest on Critical Rigor. Where humans win, it is on trust and usefulness rather than depth: a confident wrong claim, a mechanism described but not taught, or unprioritized breadth. A 98-paper automated ablation shows the gain comes from the multi-agent structure: the pipeline beats the same model run as a single rich-persona agent on 96% of papers. We release all analyses, scores, and the rubric as a community resource.
Nishant Aggarwal, Aishwarya Lekshmi Chithra, Ayushi Dubal, Sreeraj Kannakarankodi, Ian McDougall, Adarsh Mittal, Vishnu Ramadas, Noah Scott, Ranganath Selagamsetty, Weichu Yang, Karthikeyan Sankaralingam
arXiv:2607.11859 · cs.CY, cs.AR, cs.MA · submitted Jul 13, 2026 · updated Sep 20, 2026
abstract · pdf · html · 4 pages, 1 figure