In plain words: A team of AI agents turns loose research notes and sources into a submission-ready paper, writing the literature review and drawing plots and diagrams. In side-by-side human judging it beat other automatic writers, with a 50%-68% higher win rate on literature review quality.
Abstract · PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing
Synthesizing unstructured research materials into manuscripts is an essential yet under-explored challenge in AI-driven scientific discovery. Existing autonomous writers are rigidly coupled to specific experimental pipelines, and produce superficial literature reviews. We introduce PaperOrchestra, a multi-agent framework for automated AI research paper writing. It flexibly transforms unconstrained pre-writing materials into submission-ready LaTeX manuscripts, including comprehensive literature synthesis and generated visuals, such as plots and conceptual diagrams. To evaluate performance, we present PaperWritingBench, the first standardized benchmark of reverse-engineered raw materials from 200 top-tier AI conference papers, alongside a comprehensive suite of automated evaluators. In side-by-side human evaluations, PaperOrchestra significantly outperforms autonomous baselines, achieving an absolute win rate margin of 50%-68% in literature review quality, and 14%-38% in overall manuscript quality.
Yiwen Song, Yale Song, Tomas Pfister, Jinsung Yoon
arXiv:2604.05018 · cs.AI, cs.LG, cs.MA · submitted Apr 6, 2026
abstract · pdf · html · Project Page: https://yiwen-song.github.io/paper_orchestra/
It employs a specialized 5-aagents pipeline: Outline, Plotting/Lit Review, Section Writing, and Refinement. This setup greatly surpasses single-agent models in literature review quality and overall performance.
I created this repository to transform the paper’s prompts, schemas, and verification gates into a "skill pack" that any modern coding agent can use.
Repo: https://github.com/Ar9av/paper-orchestra
I am thinking of improving on it through: - optional semantic scholar support for verifying - an arxiv packager that strips comments and zips everything up for submission in one click. - human-in-the-loop checkpoints that pause the pipeline so you can approve the outline before it starts burning tokens