about
Improving Assembly Code Performance with LLMss via Reinforcement Learning (arxiv.org)
13 points by badmonster on May 20, 2025 | hide | past | pdf | 2 comments on HN

In plain words: An AI model is trained to rewrite assembly code into faster versions that still do exactly the same job. The trained model got 95% of programs correct and ran them 1.46 times faster than the standard compiler's best output.

Abstract · SuperCoder: Assembly Program Superoptimization with Large Language Models

Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its input-output behavior. In this work, we investigate whether large language models (LLMs) can serve as superoptimizers, generating assembly programs that outperform code already optimized by industry-standard compilers in end-to-end runtime. We construct the first large-scale benchmark for this problem, consisting of 8,072 assembly programs averaging 130 lines, in contrast to prior datasets restricted to 2-15 straight-line, loop-free programs. We evaluate 23 LLMs on this benchmark and find that the strongest baseline, Claude-opus-4, achieves a 51.5% test-passing rate and a 1.43x average speedup over gcc -O3. To further enhance performance, we fine-tune models with reinforcement learning, optimizing a reward function that integrates correctness and performance speedup. Starting from Qwen2.5-Coder-7B-Instruct (61.4% correctness, 1.10x speedup), the fine-tuned model SuperCoder attains 95.0% correctness and 1.46x average speedup, with additional improvement enabled by Best-of-N sampling and iterative refinement. Our results demonstrate, for the first time, that LLMs can be applied as superoptimizers for assembly programs, establishing a foundation for future research in program performance optimization beyond compiler heuristics.

Anjiang Wei, Tarun Suresh, Huanmi Tan, Yinglun Xu, Gagandeep Singh, Ke Wang, Alex Aiken
arXiv:2505.11480 · cs.CL, cs.AI, cs.PF, cs.PL, cs.SE · submitted May 16, 2025 · updated Aug 8, 2026
abstract · pdf · html

add comment on HN
Also discussed: Jan 2026 (3 points, 0 comments)

reinforcement learning can push LLMs beyond generation and into true performance optimization at the assembly level. achieving a 1.47x speedup over gcc -O3 is no small feat—especially considering -O3 is already highly optimized.
O3 is highly optimized using generic techniques that worked in a variety of scenarios and could get papers published. Given that carefully laid out assembly can outperform by like 10x across size and speed, I think there’s a lot of headroom to play with.