about
Scaling TensorFlow to 300M predictions per second (arxiv.org)
3 points by sebg on Sep 21, 2021 | hide | past | pdf | discuss on HN

In plain words: A team moved its online ad models onto TensorFlow and rebuilt how they serve predictions so responses stay fast at huge scale. The finished system handles 300 million predictions per second while keeping response times low.

Abstract · Scaling TensorFlow to 300 million predictions per second

We present the process of transitioning machine learning models to the TensorFlow framework at a large scale in an online advertising ecosystem. In this talk we address the key challenges we faced and describe how we successfully tackled them; notably, implementing the models in TF and serving them efficiently with low latency using various optimization techniques.

Jan Hartman, Davorin Kopič
arXiv:2109.09541 · cs.LG, cs.PF · submitted Sep 20, 2021
abstract · pdf · html

add comment on HN