about
You Only Look Once: Unified, Real-Time Object Detection (arxiv.org)
64 points by yogrish on Dec 31, 2015 | hide | past | pdf | 8 comments on HN

In plain words: Instead of running a classifier over many candidate spots, one network looks at the whole image once and outputs boxes and labels directly. It runs in real time at 45 frames per second, with fewer false alarms than other detectors but more misplaced boxes.

Abstract

We present YOLO, a new approach to object detection. Prior work on object detection repurposes classifiers to perform detection. Instead, we frame object detection as a regression problem to spatially separated bounding boxes and associated class probabilities. A single neural network predicts bounding boxes and class probabilities directly from full images in one evaluation. Since the whole detection pipeline is a single network, it can be optimized end-to-end directly on detection performance. Our unified architecture is extremely fast. Our base YOLO model processes images in real-time at 45 frames per second. A smaller version of the network, Fast YOLO, processes an astounding 155 frames per second while still achieving double the mAP of other real-time detectors. Compared to state-of-the-art detection systems, YOLO makes more localization errors but is far less likely to predict false detections where nothing exists. Finally, YOLO learns very general representations of objects. It outperforms all other detection methods, including DPM and R-CNN, by a wide margin when generalizing from natural images to artwork on both the Picasso Dataset and the People-Art Dataset.

Joseph Redmon, Santosh Divvala, Ross Girshick, Ali Farhadi
arXiv:1506.02640 · cs.CV · submitted Jun 8, 2015 · updated May 9, 2016
abstract · pdf · html

add comment on HN
Also discussed: Jul 2020 (1 point, 0 comments)

This is very cool. It looks from the videos like the next step for them is to provide some sort of temporal stability so that detected objects don't get temporarily forgotten across frames and so the bounds expand and contract smoothly. It's obvious that the detection is being run frame-at-a-time.

I also wonder to what extent merging the detection with underlying P-frame information from the video codecs would help. Knowing that a segment of video just moved to the left would mean the detected object could be moved to the left by the same amount, even if it was passing behind another object. Calculating the movement vectors independently seems silly if you can get that data from the underlying video codec itself.

They named their method "YOLO"…

Edit: to add something more "helpful" to this comment, their paper links to a YouTube channel [1] that shows demos of their method, which I think is great.

[1] https://goo.gl/bEs6Cj

This is really cool, even inspiring. Not just because it's one of the first examples I've seen of accurate, real-time detection powered by neural nets, but because they're getting these results via black magic, basically.

The objective function is defined heuristically, and involves about five different sub-objectives (top of page four). Some of the parameters chosen seem to be rough guesses, as does the decision to scale up the images to twice the resolution when moving from classification (the pre-training task) to detection.

It seems miraculous that a process of estimating and refinement, guided by experience, can work on tasks where you have no mathematical guarantee that a good solution can be found. Maybe in time we'll build the theory that explains just why deep learning works so well, but for now I'm just kinda awed and impressed every time one of these stories comes out.

I share your optimism, however it isn't obvious from the paper how many hyper-parameters and different variations of the loss function were tried to get this result. It is still cool that something so heuristic can work so well.
In the paper they use the abbreviation mAP without explaining what it is or providing a reference, such as "Fast YOLO, processes an astounding 155 frames per second while still achieving double the mAP of other real-time detectors"; do folks know what mAP is?
I'm assuming "mAP" in this context is "mean Average Precision": http://fastml.com/what-you-wanted-to-know-about-mean-average...

The other possibility would be "maximum a posteriori" which doesn't fit their usage here.

One of the authors has additional information posted here: http://pjreddie.com/darknet/yolo/