about
AutoUpdate: Automatically Recommend Code Updates for Android Apps (arxiv.org)
1 point by PaulHoule on May 11, 2023 | hide | past | pdf | 1 comment on HN

In plain words: They tested code-writing AI tools that rewrite outdated code into its updated form, using two collections of real before-and-after code. The tools scored well when old and new code were mixed but failed on time-based tests and new projects, often outputting no change at all.

Abstract · Automatically Recommend Code Updates: Are We There Yet?

In recent years, large pre-trained Language Models of Code (CodeLMs) have shown promising results on various software engineering tasks. One such task is automatic code update recommendation, which transforms outdated code snippets into their approved and revised counterparts. Although many CodeLM-based approaches have been proposed, claiming high accuracy, their effectiveness and reliability on real-world code update tasks remain questionable. In this paper, we present the first extensive evaluation of state-of-the-art CodeLMs for automatically recommending code updates. We assess their performance on two diverse datasets of paired updated methods, considering factors such as temporal evolution, project specificity, method size, and update complexity. Our results reveal that while CodeLMs perform well in settings that ignore temporal information, they struggle in more realistic time-wise scenarios and generalize poorly to new projects. Furthermore, CodeLM performance decreases significantly for larger methods and more complex updates. Furthermore, we observe that many CodeLM-generated "updates" are actually null, especially in time-wise settings, and meaningful edits remain challenging. Our findings highlight the significant gap between the perceived and actual effectiveness of CodeLMs for real-world code update recommendation and emphasize the need for more research on improving their practicality, robustness, and generalizability.

Yue Liu, Chakkrit Tantithamthavorn, Yonghui Liu, Patanamon Thongtanunam, Li Li
arXiv:2209.07048 · cs.SE · submitted Sep 15, 2022 · updated May 12, 2024
abstract · pdf · html · Under review at a SE journal

add comment on HN

People are often saying that code LLMs do a good job for small greenfield projects but aren’t good for maintenance, where programmers spend most of their time. This study shows excellent performance for one kind of maintenance task.