In plain words: Style rewriting is usually scored on three things: whether the style changed, whether the meaning stayed, and whether it reads smoothly. This review argues those automatic scores are misleading because the test sentences don't match real uses of style rewriting.
Abstract
The difficulty of textual style transfer lies in the lack of parallel corpora. Numerous advances have been proposed for the unsupervised generation. However, significant problems remain with the auto-evaluation of style transfer tasks. Based on the summary of Pang and Gimpel (2018) and Mir et al. (2019), style transfer evaluations rely on three criteria: style accuracy of transferred sentences, content similarity between original and transferred sentences, and fluency of transferred sentences. We elucidate the problematic current state of style transfer research. Given that current tasks do not represent real use cases of style transfer, current auto-evaluation approach is flawed. This discussion aims to bring researchers to think about the future of style transfer and style transfer evaluation research.
Richard Yuanzhe Pang
arXiv:1910.03747 · cs.CL · submitted Oct 9, 2019 · updated Oct 10, 2019
abstract · pdf · html · Extended abstract in EMNLP Workshop on Neural Generation and Translation (WNGT 2019)