about
Speed Is All You Need: On-Device Acceleration of Large Diffusion Models (arxiv.org)
56 points by Pelayu on Apr 30, 2023 | hide | past | pdf | 8 comments on HN

In plain words: A set of GPU-focused coding tricks lets big image-generating diffusion models run directly on phones instead of in the cloud. On a Samsung phone it drew a standard-size image in under 12 seconds, the fastest reported, without shrinking the model's number precision.

Abstract · Speed Is All You Need: On-Device Acceleration of Large Diffusion Models via GPU-Aware Optimizations

The rapid development and application of foundation models have revolutionized the field of artificial intelligence. Large diffusion models have gained significant attention for their ability to generate photorealistic images and support various tasks. On-device deployment of these models provides benefits such as lower server costs, offline functionality, and improved user privacy. However, common large diffusion models have over 1 billion parameters and pose challenges due to restricted computational and memory resources on devices. We present a series of implementation optimizations for large diffusion models that achieve the fastest reported inference latency to-date (under 12 seconds for Stable Diffusion 1.4 without int8 quantization on Samsung S23 Ultra for a 512x512 image with 20 iterations) on GPU-equipped mobile devices. These enhancements broaden the applicability of generative AI and improve the overall user experience across a wide range of devices.

Yu-Hui Chen, Raman Sarokin, Juhyun Lee, Jiuqiang Tang, Chuo-Ling Chang, Andrei Kulik, Matthias Grundmann
arXiv:2304.11267 · cs.CV, cs.LG, eess.IV · submitted Apr 21, 2023 · updated Jun 16, 2023
abstract · pdf · html · 4 pages (not including references), 2 figures, 2 tables. Accepted to Efficient Deep Learning for Computer Vision workshop 2023

add comment on HN
Also discussed: Apr 2023 (11 points, 2 comments) · Apr 2023 (9 points, 3 comments)

Interestingly these are OpenCL kernels so in theory some of the optimizations might run out-of-the-box on CPUs.

It would be instructive to compare their speedups on the iPhone to the Apple CoreML implementation: https://github.com/apple/ml-stable-diffusion

This incredible, can't wait to run it. Is there a code sample somewhere to reproduce their Samsung s23 results?
This is definitely a welcome development, but I'm getting so tired of all these papers trying to pay homage to the original Transformer paper in their title. It is neither funny anymore, nor does it give due credit or indicate quality and on top of that the original paper title was a pretty poor choice in hindsight, highlighting how the original authors didn't foresee the gigantic impact of their paper.
Why do you think the original paper title was a poor choice? It very much highlights the main idea, the main aspect which is studied in this paper.

The paper title is "Attention is all you need", for those who don't know.

And attention at that point in time was already very well known and part of the standard translation model. But all those attention-based encoder-decoder models where using LSTMs, or maybe CNNs. Self-attention was also already known at that point, although still rarely used. So the novelty was the study on whether a model where you remove almost everything else, except of attention, whether this still works.

Such study was on the one side just interesting in itself. But then, such model also had some advantages like faster training. In the next few years, the faster training was actually the main advantage over LSTM-based models. For a long time, it was never really clear whether a Transformer is really better than a LSTM-based model when trained the same number of epochs. In most comparisons, Transformer were simply trained much more epochs.

I'm well aware of the research that led to it. I was already working in the field back then and I remember that the community was far from realizing how monumental this paper would end up being. Otherwise the authors probably would have considered a more informative or at least less ambiguous title. It also didn't help that the architecture they described (encoder-decoder) was actually even more complicated than what we have now in GPT and the likes. And the really important thing was not that it could train more epochs than recurrent architectures (although that certainly helped the huge models that came later), but it could drastically extend context length for sequence tasks. They went from a theoretically infinite (but in practice very limited) context length to a fundamentally limited but practically obtainable one.
I am not sure it’s not funny. Elon Musk gave ChatGPT $100 million dollars. There are 9 billion people in the world… he could have made everyone a millionaire many times over! (In SHIB.) I feel like that amount of ShibaCoin would be life changing for most people. Yet he wasted it all on a company that became for-profit and sold shares to Microsoft instead.

(And no, before you say it, my math checks out!)

You must have a PhD in math because it all checks out. No errors :P