about
Learned Indexes for a Google-Scale Disk-Based Database (arxiv.org)
1 point by hamilyon2 on Oct 9, 2024 | hide | past | pdf | discuss on HN

In plain words: A small trained model that predicts where each key sits on disk was built into Bigtable, Google's huge distributed database, to replace the usual B-tree lookup. In end-to-end tests it made reads faster and handled more of them than the B-tree setup.

Abstract · Learned Indexes for a Google-scale Disk-based Database

There is great excitement about learned index structures, but understandable skepticism about the practicality of a new method uprooting decades of research on B-Trees. In this paper, we work to remove some of that uncertainty by demonstrating how a learned index can be integrated in a distributed, disk-based database system: Google's Bigtable. We detail several design decisions we made to integrate learned indexes in Bigtable. Our results show that integrating learned index significantly improves the end-to-end read latency and throughput for Bigtable.

Hussam Abu-Libdeh, Deniz Altınbüken, Alex Beutel, Ed H. Chi, Lyric Doshi, Tim Kraska, Xiaozhou, Li, Andy Ly, Christopher Olston
arXiv:2012.12501 · cs.DB, cs.DC, cs.LG · submitted Dec 23, 2020
abstract · pdf · html · 4 pages, Presented at Workshop on ML for Systems at NeurIPS 2020

add comment on HN