about
PyTorch: An Imperative Style, High-Performance Deep Learning Library [pdf] (arxiv.org)
113 points by stablemap on Dec 6, 2019 | hide | past | pdf | 16 comments on HN

In plain words: A deep learning library where the model is ordinary Python code run step by step, so it is easy to debug and change, while still running fast on GPUs. Benchmarks show it keeps pace with the fastest tools that require you to describe the model as a fixed graph first.

Abstract · PyTorch: An Imperative Style, High-Performance Deep Learning Library

Deep learning frameworks have often focused on either usability or speed, but not both. PyTorch is a machine learning library that shows that these two goals are in fact compatible: it provides an imperative and Pythonic programming style that supports code as a model, makes debugging easy and is consistent with other popular scientific computing libraries, while remaining efficient and supporting hardware accelerators such as GPUs. In this paper, we detail the principles that drove the implementation of PyTorch and how they are reflected in its architecture. We emphasize that every aspect of PyTorch is a regular Python program under the full control of its user. We also explain how the careful and pragmatic implementation of the key components of its runtime enables them to work together to achieve compelling performance. We demonstrate the efficiency of individual subsystems, as well as the overall speed of PyTorch on several common benchmarks.

Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, et al.
arXiv:1912.01703 · cs.LG, cs.MS, stat.ML · submitted Dec 3, 2019
abstract · pdf · 12 pages, 3 figures, NeurIPS 2019

add comment on HN

Finally the solution to all of your PyTorch citation problems! :)
It is surprising to me that they don't put this in the README.

(I do know it is in the root directory[0], it is just common practice to have it in the README. To be stupidly obvious)

[0] https://github.com/pytorch/pytorch/blob/master/CITATION

Man it’s kinda sad as a Lua fan to see so much interest in a project where the main goal is just to not use Lua.

I guess academics like familiarity and Lua insistently refuses to be like other languages (arrays and maps in one type, 1-based arrays, nonstandard builtin patterns, etc).

> the main goal is just to not use Lua.

I don't think that's accurate. People don't really care about Lua, they don't like or dislike it, they just don't know it. The goal of the project is to use python, because people care about Python.

It just happened that porting (lua) torch to Python was chosen, but it could have been another framework.

It's just about package support and the community. If researchers and practitioners were choosing a language based on merit alone it would probably be Julia for native speed and support for scientific computing. It's nice to have a toy language you appreciate but recall the goal is to write math into algorithms; the language is just tool.
As someone who's been a software person for 15 years I am so glad deep learning is centralizing on python. It helps so much to share tools and to use a relatively boring language. Lua offered literally nothing but ecosystem problems...
Academics simply don't have the time to learn languages that do not have substantial ecosystems and relative ease-of-use.
It’s true - there have been attempts to fix it, but nobody has created something other people want to use. The Torch project largely replaced all of the then-popular Lua packages — wxLua was dropped for Torch’s internal qtLua, the Lua concurrency libraries (Lanes and luaproc) were ignored in favor of zeroMQ, LPeg and Lua patterns were generally less popular than PCRE and Re2 bindings, et cetera. Maybe Torch is to blame (NIH syndrome), maybe the Lua packages weren’t up to the task, maybe communication within the community is too hard (Lua lacks centralized discussion channels where experienced users are regularly active), but in the end, Lua didn’t come away looking good here.

Learning a new language wasn’t too hard when that language was Python, after all.

We'd all be happier writing math; writing code is just a nuisance.
The Julia programming language's development started explicitly to address this sentiment.
ah, yes, why squint at a PDF when you can squint at LaTeX compile errors instead
Is there a typo in Listing 1?

The forward function of the conv net should use:

t3 = self.fc(t2)

instead of:

t3 = self.fc(t1)

AFAIK the nn.functional.relu function is NOT inplace by default [1]

https://pytorch.org/docs/stable/nn.functional.html

yes that's a typo
It's a bit funny to call it imperative, when really at the end of the day, the objective is to get something where you have very little insight into what the neural net. is doing to the state.