In plain words: They examined 33,000 pull requests from five AI coding agents on GitHub, comparing merged with rejected ones and reading 600 to find why they failed. Documentation and build updates merged most often, while bug fixes did worst; rejected changes were bigger and often failed checks.
Abstract · Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
AI coding agents are now submitting pull requests (PRs) to software projects, acting not just as assistants but as autonomous contributors. As these agentic contributions are rapidly increasing across real repositories, little is known about how they behave in practice and why many of them fail to be merged. In this paper, we conduct a large-scale study of 33k agent-authored PRs made by five coding agents across GitHub. (RQ1) We first quantitatively characterize merged and not-merged PRs along four broad dimensions: 1) merge outcomes across task types, 2) code changes, 3) CI build results, and 4) review dynamics. We observe that tasks related to documentation, CI, and build update achieve the highest merge success, whereas performance and bug-fix tasks perform the worst. Not-merged PRs tend to involve larger code changes, touch more files, and often do not pass the project's CI/CD pipeline validation. (RQ2) To further investigate why some agentic PRs are not merged, we qualitatively analyze 600 PRs to derive a hierarchical taxonomy of rejection patterns. This analysis complements the quantitative findings in RQ1 by uncovering rejection reasons not captured by quantitative metrics, including lack of meaningful reviewer engagement, duplicate PRs, unwanted feature implementations, and agent misalignment. Together, our findings highlight key socio-technical and human-AI collaboration factors that are critical to improving the success of future agentic workflows.
Ramtin Ehsani, Sakshi Pathak, Shriya Rawal, Abdullah Al Mujahid, Mia Mohammad Imran, Preetha Chatterjee
arXiv:2601.15195 · cs.SE, cs.AI · submitted Jan 21, 2026
abstract · pdf · html · Accepted at International Mining Software Repositories Conference (MSR 2026)
I always have to tell the agent to use functional iterators and itertools all the time but it still prefers to use primitive for loop and push into a mutable array. Not that it doesn't get the thing done but please when you can iter collect why can't you use them
It also lacks the ability to use high level data structure such as bit vector and matrices. I doubt they could use even harder stuff such as B+ tree, red black tree or Fibonacci heap...
In my experimentation in building vibewasm, a wasm engine using a binding that I wrote the low layer of sljit manually, and I instructed Opus 4.5 to build a wasm engine out of it, and take in designs from other wasm engine that is based on sljit but in C or C++...pwart and walrus to be specific
I have a specific case on the use of bit vector, at least for Opus, it always tend to use hash/btree set for this.
I have to explicitly explain and tell Opus that the way you are marking the register can be represented as a bit vector, because the number of registers are bounded (15) and bit vectors will have an ultimate space save. But a next refactor attempt to rewrite the SSA layer into RTL (register transfer language that targets sljit), the same mistake happened again. It turns out my prompt of following previous design didn't work, or Opus simply just couldn't justify that bit vector is the best despite user requirements.
I have to revert that change and do it myself.