In plain words: A bidding agent reads plain-text descriptions of each auction and learns which ads to bid on by chasing profit, instead of following rules hand-tuned by experts. In JD.com's June 18 sale it raised ad revenue from that slice by over 50% and improved advertisers' returns.
Abstract
We present LADDER, the first deep reinforcement learning agent that can successfully learn control policies for large-scale real-world problems directly from raw inputs composed of high-level semantic information. The agent is based on an asynchronous stochastic variant of DQN (Deep Q Network) named DASQN. The inputs of the agent are plain-text descriptions of states of a game of incomplete information, i.e. real-time large scale online auctions, and the rewards are auction profits of very large scale. We apply the agent to an essential portion of JD's online RTB (real-time bidding) advertising business and find that it easily beats the former state-of-the-art bidding policy that had been carefully engineered and calibrated by human experts: during JD.com's June 18th anniversary sale, the agent increased the company's ads revenue from the portion by more than 50%, while the advertisers' ROI (return on investment) also improved significantly.
Yu Wang, Jiayi Liu, Yuxiang Liu, Jun Hao, Yang He, Jinghe Hu, Weipeng P. Yan, Mantian Li
arXiv:1708.05565 · cs.LG, cs.AI, cs.CL, cs.GT · submitted Aug 18, 2017 · updated Sep 1, 2017
abstract · pdf · 8 pages, 12 figures