about
End to End Learning for Self-Driving Cars [pdf] (arxiv.org)
1 point by AndreasM on Apr 27, 2016 | hide | past | pdf | 1 comment on HN

In plain words: A neural network turns the front camera image into a steering angle, learning from human steering alone instead of the usual lane-detection, planning and control steps. It drove roads with or without lane markings, parking lots and unpaved roads at 30 frames per second.

Abstract · End to End Learning for Self-Driving Cars

We trained a convolutional neural network (CNN) to map raw pixels from a single front-facing camera directly to steering commands. This end-to-end approach proved surprisingly powerful. With minimum training data from humans the system learns to drive in traffic on local roads with or without lane markings and on highways. It also operates in areas with unclear visual guidance such as in parking lots and on unpaved roads. The system automatically learns internal representations of the necessary processing steps such as detecting useful road features with only the human steering angle as the training signal. We never explicitly trained it to detect, for example, the outline of roads. Compared to explicit decomposition of the problem, such as lane marking detection, path planning, and control, our end-to-end system optimizes all processing steps simultaneously. We argue that this will eventually lead to better performance and smaller systems. Better performance will result because the internal components self-optimize to maximize overall system performance, instead of optimizing human-selected intermediate criteria, e.g., lane detection. Such criteria understandably are selected for ease of human interpretation which doesn't automatically guarantee maximum system performance. Smaller networks are possible because the system learns to solve the problem with the minimal number of processing steps. We used an NVIDIA DevBox and Torch 7 for training and an NVIDIA DRIVE(TM) PX self-driving car computer also running Torch 7 for determining where to drive. The system operates at 30 frames per second (FPS).

Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, et al.
arXiv:1604.07316 · cs.CV, cs.LG, cs.NE · submitted Apr 25, 2016
abstract · pdf · html

add comment on HN
Also discussed: Apr 2016 (13 points, 1 comment) · Apr 2016 (1 point, 0 comments)

In short: NVIDIA has made a neural network which steers a car based on a single camera and trained only with footage from cars being driven manually. No control theory, no rules of the roads, not even the concept of a road is encoded manually.

Video here: https://drive.google.com/open?id=0B9raQzOpizn1TkRIa241ZnBEcj...

This is pretty cool. Granted, it does not handle extremely complex traffic situations (yet), but is it a much simpler approach to what most other groups are doing.