about
CRQBench: A Benchmark of Code Reasoning Questions (arxiv.org)
1 point by PaulHoule on Aug 28, 2024 | hide | past | pdf | discuss on HN

In plain words: A test set of 100 C++ code reasoning questions, taken from real code review comments and checked by people with an AI helper, tests code understanding apart from software engineering skill. The tested AI assistant answered 65 of 100 correctly, grounded in the context.

Abstract

Large Language Models have demonstrated exceptional proficiency on coding tasks, but it is challenging to precisely evaluate their code reasoning ability. Existing benchmarks are insufficient as they are unrealistic and conflate semantic reasoning ability with performance on software engineering tasks. We introduce CRQBench, a benchmark of 100 C++ code reasoning questions and answers derived from contextualized code review comments. To curate CRQBench, we use an LLM assistant alongside human inspection, reducing manual effort. We conduct an evaluation of GPT-4 on CRQBench and find that it produces correct responses grounded in the given context for 65 of the 100 questions.

Elizabeth Dinella, Satish Chandra, Petros Maniatis
arXiv:2408.08453 · cs.SE · submitted Aug 15, 2024
abstract · pdf · html

add comment on HN