about
Face Recognition with Hybrid Efficient Convolution Algorithms on FPGAs (arxiv.org)
3 points by godelmachine on Mar 28, 2018 | hide | past | pdf | discuss on HN

In plain words: They build a face-recognition system on a reprogrammable chip, using shortcut math formulas that skip repeated multiplications and picking the best one for each layer while running many parts at once. It finished 3.75 times faster than an NVIDIA graphics card, beating earlier chips.

Abstract

Deep Convolutional Neural Networks have become a Swiss knife in solving critical artificial intelligence tasks. However, deploying deep CNN models for latency-critical tasks remains to be challenging because of the complex nature of CNNs. Recently, FPGA has become a favorable device to accelerate deep CNNs thanks to its high parallel processing capability and energy efficiency. In this work, we explore different fast convolution algorithms including Winograd and Fast Fourier Transform (FFT), and find an optimal strategy to apply them together on different types of convolutions. We also propose an optimization scheme to exploit parallelism on novel CNN architectures such as Inception modules in GoogLeNet. We implement a configurable IP-based face recognition acceleration system based on FaceNet using High-Level Synthesis. Our implementation on a Xilinx Ultrascale device achieves 3.75x latency speedup compared to a high-end NVIDIA GPU and surpasses previous FPGA results significantly.

Chuanhao Zhuge, Xinheng Liu, Xiaofan Zhang, Sudeep Gummadi, Jinjun Xiong, Deming Chen
arXiv:1803.09004 · cs.CV, cs.DC · submitted Mar 23, 2018
abstract · pdf · html · This paper is accepted in GLSVLSI'18

add comment on HN