about
Why most AI coding benchmarks are misleading (COMPASS paper) (arxiv.org)
13 points by jmeaden on Sep 19, 2025 | hide | past | pdf | 2 comments on HN

In plain words: A new test set scores AI-written code on whether it works, how fast it runs, and how clean it is, using 50 contest problems and human results. Models that solved problems correctly still often wrote slow or messy code, so passing tests alone hides weaknesses.

Abstract · COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models

Current code generation benchmarks focus primarily on functional correctness while overlooking two critical aspects of real-world programming: algorithmic efficiency and code quality. We introduce COMPASS (COdility's Multi-dimensional Programming ASSessment), a comprehensive evaluation framework that assesses code generation across three dimensions: correctness, efficiency, and quality. COMPASS consists of 50 competitive programming problems from real Codility competitions, providing authentic human baselines from 393,150 submissions. Unlike existing benchmarks that treat algorithmically inefficient solutions identically to optimal ones provided they pass test cases, COMPASS systematically evaluates runtime efficiency and code quality using industry-standard analysis tools. Our evaluation of three leading reasoning-enhanced models, Anthropic Claude Opus 4, Google Gemini 2.5 Pro, and OpenAI O4-Mini-High, reveals that models achieving high correctness scores do not necessarily produce efficient algorithms or maintainable code. These findings highlight the importance of evaluating more than just correctness to truly understand the real-world capabilities of code generation models. COMPASS serves as a guiding framework, charting a path for future research toward AI systems that are robust, reliable, and ready for production use.

James Meaden, Michał Jarosz, Piotr Jodłowski, Grigori Melnik
arXiv:2508.13757 · cs.SE, cs.AI · submitted Aug 19, 2025
abstract · pdf · html

add comment on HN

Hi, I’m one of the authors. Happy to answer questions about the dataset (LLM coding performance compared to 390k+ human submissions), the scoring approach, or the methodology behind COMPASS. Feedback and critique are welcome.
Hello, I'm someone who does not have a background in CS, so my apologies for not being able to read the paper in-full. Is there any clear-cut strategy you would recommend to model developers so they can improve in not just correctness, but in quality & efficiency? I'm sure it's in the paper & I wish I could understand it in-depth.

If you don't mind me asking a more personal question, I would love to go back to uni for a master's in computer science & hopefully assist with papers like this one day. Do you have any advice for someone with industry CS experience (SWE) vs. academic to make the leap to the academic side? I genuinely love this kind of stuff and already make a decent living so it's not for money.