In plain words: They design a vision network and its chip circuit together, searching for a network that fits a small programmable chip and writing the circuit code automatically. On object detection, it ran 2.48 times faster than earlier chip designs while using 40% less power.
Abstract · FPGA/DNN Co-Design: An Efficient Design Methodology for IoT Intelligence on the Edge
While embedded FPGAs are attractive platforms for DNN acceleration on edge-devices due to their low latency and high energy efficiency, the scarcity of resources of edge-scale FPGA devices also makes it challenging for DNN deployment. In this paper, we propose a simultaneous FPGA/DNN co-design methodology with both bottom-up and top-down approaches: a bottom-up hardware-oriented DNN model search for high accuracy, and a top-down FPGA accelerator design considering DNN-specific characteristics. We also build an automatic co-design flow, including an Auto-DNN engine to perform hardware-oriented DNN model search, as well as an Auto-HLS engine to generate synthesizable C code of the FPGA accelerator for explored DNNs. We demonstrate our co-design approach on an object detection task using PYNQ-Z1 FPGA. Results show that our proposed DNN model and accelerator outperform the state-of-the-art FPGA designs in all aspects including Intersection-over-Union (IoU) (6.2% higher), frames per second (FPS) (2.48X higher), power consumption (40% lower), and energy efficiency (2.5X higher). Compared to GPU-based solutions, our designs deliver similar accuracy but consume far less energy.
Cong Hao, Xiaofan Zhang, Yuhong Li, Sitao Huang, Jinjun Xiong, Kyle Rupnow, Wen-mei Hwu, Deming Chen
arXiv:1904.04421 · cs.CV · submitted Apr 9, 2019
abstract · pdf · html · Accepted by Design Automation Conference (DAC'2019)
We have known for years now that deep networks are extremely compressible when they're trained. You can drop the vast majority of the weights and still maintain almost all of the performance. Just like you can drop the accuracy from floats, to ints, to int8, and even to bool for the weights and you can still perform well.
This papers is a joke and it deserves to be rejected from anywhere it's submitted. They optimize the shape of the FPGA network but not the GPU network. They don't apply the standard methods to prune weights in networks and they don't compare to those methods. They also pick one GPU network at random, lots of object detectors are faster than Yolo. People have explored the speed-performance tradeoff closely.
About FPGAs in general: there are great reasons why FPGAs have been just over the horizon for a long time. The toolchains suck and they're mostly very closed down. Debugging is very hard. They cost 20x or so as much as GPUs. Your code becomes very specific to one particular FPGA so upgrading is a total pain. And more.. There are plenty of better solutions out there like TPUs.