about
Aligning Text-to-Image Models Using Human Feedback (arxiv.org)
2 points by mkaic on Feb 24, 2023 | hide | past | pdf | 1 comment on HN

In plain words: People rate how well generated images match their text prompts, and those ratings train a scorer that the image generator then learns to please. The tuned model draws requested colors, counts, and backgrounds more accurately than the original.

Abstract · Aligning Text-to-Image Models using Human Feedback

Deep generative models have shown impressive results in text-to-image synthesis. However, current text-to-image models often generate images that are inadequately aligned with text prompts. We propose a fine-tuning method for aligning such models using human feedback, comprising three stages. First, we collect human feedback assessing model output alignment from a set of diverse text prompts. We then use the human-labeled image-text dataset to train a reward function that predicts human feedback. Lastly, the text-to-image model is fine-tuned by maximizing reward-weighted likelihood to improve image-text alignment. Our method generates objects with specified colors, counts and backgrounds more accurately than the pre-trained model. We also analyze several design choices and find that careful investigations on such design choices are important in balancing the alignment-fidelity tradeoffs. Our results demonstrate the potential for learning from human feedback to significantly improve text-to-image models.

Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Shixiang Shane Gu
arXiv:2302.12192 · cs.LG, cs.AI, cs.CV · submitted Feb 23, 2023
abstract · pdf · html

add comment on HN

Although this paper is a fairly small-scale proof of concept, I am extremely excited to see where this technique goes. Between ChatGPT and this, I get the feeling that we are just barely touching the surface of what RLHF is capable of in generative models. This paper[0] about mixing in RLHF with supervised learning from the start of LLM training also has me very excited. I feel like we haven't seen a single method make this big of waves in the ML world for a while now.

[0]https://arxiv.org/abs/2302.08582

(Currently empty) discussion thread for [0]: https://news.ycombinator.com/item?id=34869295