about
Fast Point R-CNN (arxiv.org)
2 points by sel1 on Aug 12, 2019 | hide | past | pdf | discuss on HN

In plain words: A two-stage system spots 3D objects in point clouds: a fast grid pass proposes rough boxes, then a second pass blends the points inside each box to refine it. It matched the best 3D and top-down results on KITTI at 15 frames per second.

Abstract

We present a unified, efficient and effective framework for point-cloud based 3D object detection. Our two-stage approach utilizes both voxel representation and raw point cloud data to exploit respective advantages. The first stage network, with voxel representation as input, only consists of light convolutional operations, producing a small number of high-quality initial predictions. Coordinate and indexed convolutional feature of each point in initial prediction are effectively fused with the attention mechanism, preserving both accurate localization and context information. The second stage works on interior points with their fused feature for further refining the prediction. Our method is evaluated on KITTI dataset, in terms of both 3D and Bird's Eye View (BEV) detection, and achieves state-of-the-arts with a 15FPS detection rate.

Yilun Chen, Shu Liu, Xiaoyong Shen, Jiaya Jia
arXiv:1908.02990 · cs.CV · submitted Aug 8, 2019 · updated Aug 16, 2019
abstract · pdf · html

add comment on HN