about
TF-Replicator: Distributed Machine Learning for Researchers (arxiv.org)
1 point by lawrenceyan on Feb 7, 2019 | hide | past | pdf | discuss on HN

In plain words: A tool that lets researchers write training code once and run it on any mix of machines and chips, splitting the work across devices automatically. It scaled three very different models well—image classification, image generation, and robot control—without users needing to know how to coordinate many machines.

Abstract

We describe TF-Replicator, a framework for distributed machine learning designed for DeepMind researchers and implemented as an abstraction over TensorFlow. TF-Replicator simplifies writing data-parallel and model-parallel research code. The same models can be effortlessly deployed to different cluster architectures (i.e. one or many machines containing CPUs, GPUs or TPU accelerators) using synchronous or asynchronous training regimes. To demonstrate the generality and scalability of TF-Replicator, we implement and benchmark three very different models: (1) A ResNet-50 for ImageNet classification, (2) a SN-GAN for class-conditional ImageNet image generation, and (3) a D4PG reinforcement learning agent for continuous control. Our results show strong scalability performance without demanding any distributed systems expertise of the user. The TF-Replicator programming model will be open-sourced as part of TensorFlow 2.0 (see https://github.com/tensorflow/community/pull/25).

Peter Buchlovsky, David Budden, Dominik Grewe, Chris Jones, John Aslanides, Frederic Besse, Andy Brock, Aidan Clark, Sergio Gómez Colmenarejo, Aedan Pope, Fabio Viola, Dan Belov
arXiv:1902.00465 · cs.LG, cs.AI, cs.DC, stat.ML · submitted Feb 1, 2019
abstract · pdf · html

add comment on HN
Also discussed: Feb 2019 (2 points, 0 comments) · Feb 2019 (3 points, 0 comments)