In plain words: A new collection pairs peer-review comments with the exact paper edits authors made in reply, labeled to show which edits respond to which comment. Ten models struggled to match edits to comments, and a top chatbot's rewrites copied feedback wording while missing technical detail.
Abstract · ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews
We introduce the task of automatically revising scientific papers based on peer feedback and release ARIES, a dataset of review comments and their corresponding paper edits. The data is drawn from real reviewer-author interactions from computer science, and we provide labels linking each reviewer comment to the specific paper edits made by the author in response. We automatically create a high-precision silver training set, as well as an expert-labeled test set that shows high inter-annotator agreement. In experiments with 10 models covering the state of the art, we find that they struggle even to identify which edits correspond to a comment -- especially when the relationship between the edit and the comment is indirect and requires reasoning to uncover. We also extensively analyze GPT-4's ability to generate edits given a comment and the original paper. We find that it often succeeds on a superficial level, but tends to rigidly follow the wording of the feedback rather than the underlying intent, and lacks technical details compared to human-written edits.
Mike D'Arcy, Alexis Ross, Erin Bransom, Bailey Kuehl, Jonathan Bragg, Tom Hope, Doug Downey
arXiv:2306.12587 · cs.CL · submitted Jun 21, 2023 · updated Aug 6, 2024
abstract · pdf · html · ACL 2024, 10 pages, 2 figures
Code and data: https://github.com/allenai/aries