about
KML: Using Machine Learning to Improve Storage Systems (arxiv.org)
2 points by belter on Nov 28, 2021 | hide | past | pdf | 1 comment on HN

In plain words: KML replaces hand-tuned storage settings with tiny learning models inside the operating system that watch the workload and adjust read-ahead and network file read sizes on their own. In tests it boosted throughput up to 15x with almost no CPU cost.

Abstract

Operating systems include many heuristic algorithms designed to improve overall storage performance and throughput. Because such heuristics cannot work well for all conditions and workloads, system designers resorted to exposing numerous tunable parameters to users -- thus burdening users with continually optimizing their own storage systems and applications. Storage systems are usually responsible for most latency in I/O-heavy applications, so even a small latency improvement can be significant. Machine learning (ML) techniques promise to learn patterns, generalize from them, and enable optimal solutions that adapt to changing workloads. We propose that ML solutions become a first-class component in OSs and replace manual heuristics to optimize storage systems dynamically. In this paper, we describe our proposed ML architecture, called KML. We developed a prototype KML architecture and applied it to two case studies: optimizing readahead and NFS read-size values. Our experiments show that KML consumes less than 4KB of dynamic kernel memory, has a CPU overhead smaller than 0.2%, and yet can learn patterns and improve I/O throughput by as much as 2.3x and 15x for two case studies -- even for complex, never-seen-before, concurrently running mixed workloads on different storage devices.

Ibrahim Umit Akgun, Ali Selman Aydin, Andrew Burford, Michael McNeill, Michael Arkhangelskiy, Aadil Shaikh, Lukas Velikov, Erez Zadok
arXiv:2111.11554 · cs.OS, cs.LG · submitted Nov 22, 2021 · updated Jan 26, 2022
abstract · pdf · html · 17 pages, 13 figures

add comment on HN

https://arxiv.org/pdf/2111.11554.pdf

"...In this paper, we describe our proposed ML architecture, called KML. We developed a prototype KML architecture and applied it to two problems: optimal readahead and NFS read-size values. Our experiments show that KML consumes little OS resources, adds negligible latency, and yet can learn patterns that can improve I/O throughput by as much as 2.3x or 15x for the two use cases respectively -- even for complex, never-before-seen, concurrently running mixed workloads on different storage devices..."