In plain words: Balsa learns how to run queries by trial and error, practicing in a simulator then tuning safely on real runs, instead of copying an expert-built optimizer. After two hours it matched hand-written optimizers, and a few more hours made it up to 2.8× faster.
Abstract
Query optimizers are a performance-critical component in every database system. Due to their complexity, optimizers take experts months to write and years to refine. In this work, we demonstrate for the first time that learning to optimize queries without learning from an expert optimizer is both possible and efficient. We present Balsa, a query optimizer built by deep reinforcement learning. Balsa first learns basic knowledge from a simple, environment-agnostic simulator, followed by safe learning in real execution. On the Join Order Benchmark, Balsa matches the performance of two expert query optimizers, both open-source and commercial, with two hours of learning, and outperforms them by up to 2.8$\times$ in workload runtime after a few more hours. Balsa thus opens the possibility of automatically learning to optimize in future compute environments where expert-designed optimizers do not exist.
Zongheng Yang, Wei-Lin Chiang, Sifei Luan, Gautam Mittal, Michael Luo, Ion Stoica
arXiv:2201.01441 · cs.DB, cs.LG · submitted Jan 5, 2022 · updated May 3, 2022
abstract · pdf · html · SIGMOD 2022; code released at: https://github.com/balsa-project/balsa/