In plain words: A system reads a translation with BERT, a language model that looks at words in both directions, and predicts a quality score. On the 2017 shared translation-scoring test, it matched human judgments better than any other automatic metric for every language translated into English.
Abstract · Machine Translation Evaluation with BERT Regressor
We introduce the metric using BERT (Bidirectional Encoder Representations from Transformers) (Devlin et al., 2019) for automatic machine translation evaluation. The experimental results of the WMT-2017 Metrics Shared Task dataset show that our metric achieves state-of-the-art performance in segment-level metrics task for all to-English language pairs.
Hiroki Shimanaka, Tomoyuki Kajiwara, Mamoru Komachi
arXiv:1907.12679 · cs.CL · submitted Jul 29, 2019
abstract · pdf · html · 6 pages