about
Comparative Study of Caffe, Neon, Theano, and Torch for Deep Learning (arxiv.org)
73 points by mindcrime on Dec 12, 2015 | hide | past | pdf | 8 comments on HN

In plain words: Five deep learning tools were compared on how easily they can be extended and how fast they train and run networks on a processor and graphics card. Torch was fastest overall, especially on the processor and for large networks on the graphics card.

Abstract · Comparative Study of Deep Learning Software Frameworks

Deep learning methods have resulted in significant performance improvements in several application domains and as such several software frameworks have been developed to facilitate their implementation. This paper presents a comparative study of five deep learning frameworks, namely Caffe, Neon, TensorFlow, Theano, and Torch, on three aspects: extensibility, hardware utilization, and speed. The study is performed on several types of deep learning architectures and we evaluate the performance of the above frameworks when employed on a single machine for both (multi-threaded) CPU and GPU (Nvidia Titan X) settings. The speed performance metrics used here include the gradient computation time, which is important during the training phase of deep networks, and the forward time, which is important from the deployment perspective of trained networks. For convolutional networks, we also report how each of these frameworks support various convolutional algorithms and their corresponding performance. From our experiments, we observe that Theano and Torch are the most easily extensible frameworks. We observe that Torch is best suited for any deep architecture on CPU, followed by Theano. It also achieves the best performance on the GPU for large convolutional and fully connected networks, followed closely by Neon. Theano achieves the best performance on GPU for training and deployment of LSTM networks. Caffe is the easiest for evaluating the performance of standard deep architectures. Finally, TensorFlow is a very flexible framework, similar to Theano, but its performance is currently not competitive compared to the other studied frameworks.

Soheil Bahrampour, Naveen Ramakrishnan, Lukas Schott, Mohak Shah
arXiv:1511.06435 · cs.LG · submitted Nov 19, 2015 · updated Mar 30, 2016
abstract · pdf · html · Submitted to KDD 2016 with TensorFlow results added. At the time of submission to KDD, TensorFlow was available only with cuDNN v.2 and thus its performance is reported with that version

add comment on HN
Also discussed: May 2016 (3 points, 0 comments)

For those interested in learning about the field: focus on the concepts, not the frameworks. This is true especially of machine learning, which is highly conceptual. I don't think these frameworks are suitable for people not familiar with the concepts. For example, I think it's a good exercise for beginners to implement a basic (no convolution) two layer neural net in plain Python using matrix operations in numpy.

Regardless, these frameworks are more similar than different (one exception is Caffe, which is very focused on computer vision and makes it really easy to do stuff in that area, but makes it really hard to do anything else). In the course of exploring the area, you will probably touch all of them at some point (when you're trying to use a researcher's code).

This very much. Understanding the underlying concepts is crucial. Many people use out-of-the-box tools just plainly wrong, because they didn't bother with fundamentals. I know a guy who reinvented Naive Bayes and Laplace smoothing. Kudos for him, but he wasted about 3-4 months of company's time on it.

That said, I would start with Python based frameworks:

- Caffe is more for production, and Theano and family for experimentation.

- You can play with Theano or Tensorflow using Keras which abstracts them

- Python scientific stack makes it easier to visualize the results later

On that note, one of the best books I've read for teaching the concepts is Learning From Data by Yaser Abu-Mostafa et al. I highly recommend it as a way to refresh or introduce the fundamentals: http://amlbook.com/
It's a bummer they published this just before TensorFlow release...
True. But new "stuff" is coming out so fast in ML these days, that pretty much whenever you publish a comparison, a month later something new is going to drop.
Torch, the scientific computing framework for LuaJIT, is based on the great Lua language. This makes it a good choice - LuaJIT per se is a faster than any Python interpreter/VM.

Caffe is C++, and the rest are Python (Neon, Theano and TensorFlow).

This is a bit misleading. No framework is doing computation in the language you write your model in. The code you write in any of these frameworks just calls wrappers around the underlying C/C++/Cuda implementation.
You are right, my statement was a bit misleading. These Lua and Python based frameworks call the high performance implementation which is coded usually in C/C++/Cuda/Fortran/etc. Nevertheless the Lua <-> C interface is lightweight and it certainly helps if the overhead is low.

I know a guy who wrote a fast SAT solver in C, though called it with various slow bash scripts. I asked him why he uses bash, he mentioned it doesn't matter. In the end Microsoft's SAT solver beat him, exactly because they cared about such details too.