In plain words: Ray spreads AI work across many machines, mixing one-off jobs with long-lived workers that keep state so agents can learn while interacting with the world. It handled over 1.8 million tasks per second and beat systems built for a single style on reinforcement learning.
Abstract
The next generation of AI applications will continuously interact with the environment and learn from these interactions. These applications impose new and demanding systems requirements, both in terms of performance and flexibility. In this paper, we consider these requirements and present Ray---a distributed system to address them. Ray implements a unified interface that can express both task-parallel and actor-based computations, supported by a single dynamic execution engine. To meet the performance requirements, Ray employs a distributed scheduler and a distributed and fault-tolerant store to manage the system's control state. In our experiments, we demonstrate scaling beyond 1.8 million tasks per second and better performance than existing specialized systems for several challenging reinforcement learning applications.
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, Ion Stoica
arXiv:1712.05889 · cs.DC, cs.AI, cs.LG, stat.ML · submitted Dec 16, 2017 · updated Sep 30, 2018
abstract · pdf · html · 17 pages, 14 figures, 13th USENIX Symposium on Operating Systems Design and Implementation, 2018
Of all of the authors listed, none of their previous papers read like the Ray paper nor does the proposal paper (Real-Time Machine Learning : The missing pieces). http://arxiv.org/abs/1703.03924 reads like a corporate/industry grade requirements/proposal doc which is a huge departure from all of the author's prior papers... So, out of the blue, corporate level infrastructure project proposal and completion within the span of a year?
Can any of the paper's authors speak more clearly on who led this project over what span of time, under what direction, and with which industry groups? I see no background or papers from the individuals priority listed in the paper reflective of the sort that creates a formalized and industry grade Distributed Computational Framework such as this. > Robert Nishihara > Philipp Moritz Are priority listed yet have no prior papers leading to such a development. To what degree did : https://rise.cs.berkeley.edu/sponsors/ Drive this?