about
Evolution Without an Oracle: Driving Effective Evolution with LLM Judges (arxiv.org)
2 points by PaulHoule 271 days ago | hide | past | pdf | discuss on HN

In plain words: Instead of a computer-scored test, this system lets AI judges score solutions as they are repeatedly improved, breaking vague instructions into specific checkable requirements to keep the judges' noisy feedback steady. It raised software requirement satisfaction from 39.9% to 61.9% over strong competing approaches.

Abstract · Evolution without an Oracle: Driving Effective Evolution with LLM Judges

The integration of Large Language Models (LLMs) with Evolutionary Computation (EC) has unlocked new frontiers in scientific discovery but remains shackled by a fundamental constraint: the reliance on an Oracle--an objective, machine-computable fitness function. This paper breaks this barrier by asking: Can evolution thrive in a purely subjective landscape governed solely by LLM judges? We introduce MADE (Multi-Agent Decomposed Evolution), a framework that tames the inherent noise of subjective evaluation through "Problem Specification." By decomposing vague instructions into specific, verifiable sub-requirements, MADE transforms high-variance LLM feedback into stable, precise selection pressure. The results are transformative: across complex benchmarks like DevAI and InfoBench, MADE outperforms strong baselines by over 50% in software requirement satisfaction (39.9% to 61.9%) and achieves a 95% perfect pass rate on complex instruction following. This work validates a fundamental paradigm shift: moving from optimizing "computable metrics" to "describable qualities," thereby unlocking evolutionary optimization for the vast open-ended domains where no ground truth exists.

Zhe Zhao, Yuheng Yang, Haibin Wen, Xiaojie Qiu, Zaixi Zhang, Qingfu Zhang
arXiv:2511.19489 · cs.SE, cs.AI · submitted Nov 23, 2025
abstract · pdf · html · 14 pages, 5 figures

add comment on HN