In plain words: A new system for running prediction jobs across many machines handles real-time feedback loops, where predictions feed back into decisions within milliseconds and the flow of work changes on the fly. A prototype ran a representative task 63 times faster than today's best system.
Abstract
Machine learning applications are increasingly deployed not only to serve predictions using static models, but also as tightly-integrated components of feedback loops involving dynamic, real-time decision making. These applications pose a new set of requirements, none of which are difficult to achieve in isolation, but the combination of which creates a challenge for existing distributed execution frameworks: computation with millisecond latency at high throughput, adaptive construction of arbitrary task graphs, and execution of heterogeneous kernels over diverse sets of resources. We assert that a new distributed execution framework is needed for such ML applications and propose a candidate approach with a proof-of-concept architecture that achieves a 63x performance improvement over a state-of-the-art execution framework for a representative application.
Robert Nishihara, Philipp Moritz, Stephanie Wang, Alexey Tumanov, William Paul, Johann Schleier-Smith, Richard Liaw, Mehrdad Niknami, Michael I. Jordan, Ion Stoica
arXiv:1703.03924 · cs.DC, cs.AI, cs.LG · submitted Mar 11, 2017 · updated May 19, 2017
abstract · pdf · html · 6 pages, 3 figures