about
How Can Reinforcement Learning Achieve Expert-Level [Chip] Placement? (arxiv.org)
4 points by Jimmc414 92 days ago | hide | past | pdf | discuss on HN

In plain words: Instead of scoring chip layouts by wire length alone, this learns a scoring model from finished expert layouts, working backward to the steps that made them. It reaches expert-quality placement from as few as one design and handles new ones well.

Abstract · How Can Reinforcement Learning Achieve Expert-level Placement?

Chip placement is a critical step in physical design. While reinforcement learning (RL)-based methods have recently emerged, their training primarily focuses on wirelength optimization, and therefore often fail to achieve expert-quality layouts. We identify the reward design as the primary cause for the performance gap with experts, and instead of formalizing intricate processes, we circumvent this by directly learning from expert layouts to derive a reward model. Our approach starts from the final expert layouts to infer step-by-step expert trajectories. Using these trajectories as demonstrations or preferences, we train a model that captures the latent implicit rewards in expert results. Experiments show that our framework can efficiently learn from even a single design and generalize well to unseen cases.

Ruo-Tong Chen, Ke Xue, Chengrui Gao, Yunqi Shi, Tian Xu, Peng Xie, Siyuan Xu, Mingxuan Yuan, Chao Qian, Zhi-Hua Zhou
arXiv:2604.25191 · cs.AR, cs.AI, cs.LG · submitted Apr 28, 2026 · updated Jun 1, 2026
abstract · pdf · html · DAC 2026

add comment on HN