about
On-Device Acceleration of Large Diffusion Models (arxiv.org)
9 points by mztwo on Apr 25, 2023 | hide | past | pdf | 3 comments on HN

In plain words: A set of GPU-focused coding tricks lets big image-generating diffusion models run directly on phones instead of in the cloud. On a Samsung phone it drew a standard-size image in under 12 seconds, the fastest reported, without shrinking the model's number precision.

Abstract · Speed Is All You Need: On-Device Acceleration of Large Diffusion Models via GPU-Aware Optimizations

The rapid development and application of foundation models have revolutionized the field of artificial intelligence. Large diffusion models have gained significant attention for their ability to generate photorealistic images and support various tasks. On-device deployment of these models provides benefits such as lower server costs, offline functionality, and improved user privacy. However, common large diffusion models have over 1 billion parameters and pose challenges due to restricted computational and memory resources on devices. We present a series of implementation optimizations for large diffusion models that achieve the fastest reported inference latency to-date (under 12 seconds for Stable Diffusion 1.4 without int8 quantization on Samsung S23 Ultra for a 512x512 image with 20 iterations) on GPU-equipped mobile devices. These enhancements broaden the applicability of generative AI and improve the overall user experience across a wide range of devices.

Yu-Hui Chen, Raman Sarokin, Juhyun Lee, Jiuqiang Tang, Chuo-Ling Chang, Andrei Kulik, Matthias Grundmann
arXiv:2304.11267 · cs.CV, cs.LG, eess.IV · submitted Apr 21, 2023 · updated Jun 16, 2023
abstract · pdf · html · 4 pages (not including references), 2 figures, 2 tables. Accepted to Efficient Deep Learning for Computer Vision workshop 2023

add comment on HN
Also discussed: Apr 2023 (56 points, 8 comments) · Apr 2023 (11 points, 2 comments)

> Google LLC

Now we are talking. Stable Diffusion getting some assistance in on-device usage.

Won't be surprised to see Apple join in.

Apple ported SD to CoreML almost immediately, and pushed it to Github!

That being said, they havent really turned it into a product.

> That being said, they havent really turned it into a product.

I don't think they need to, since they are also AI shovel makers due to their Apple Silicon chips.

So I guess that they will push for on-device LLMs soon. Won't be surprising to see that on iPhones and Macs.