In plain words: A C program split into threads that pass data through queues is compiled into an FPGA chip, turning software parallelism into parallel circuits that do the heavy image math. On a mid-sized Intel chip it runs VGG-16 at 138 billion operations per second.
Abstract
A deep-learning inference accelerator is synthesized from a C-language software program parallelized with Pthreads. The software implementation uses the well-known producer/consumer model with parallel threads interconnected by FIFO queues. The LegUp high-level synthesis (HLS) tool synthesizes threads into parallel FPGA hardware, translating software parallelism into spatial parallelism. A complete system is generated where convolution, pooling and padding are realized in the synthesized accelerator, with remaining tasks executing on an embedded ARM processor. The accelerator incorporates reduced precision, and a novel approach for zero-weight-skipping in convolution. On a mid-sized Intel Arria 10 SoC FPGA, peak performance on VGG-16 is 138 effective GOPS.
Jin Hee Kim, Brett Grady, Ruolong Lian, John Brothers, Jason H. Anderson
arXiv:1807.10695 · cs.LG, cs.AR, cs.PF, cs.PL, stat.ML · submitted Jul 27, 2018
abstract · pdf · html