about
Simplify-Then-Translate:Automatic Preprocessing for BlackBox Machine Translation (arxiv.org)
1 point by xbmcuser on Jun 26, 2020 | hide | past | pdf | 1 comment on HN

In plain words: A learned rewriter simplifies each sentence before it reaches a closed-off translation service, training on pairs made by pushing text through that service and back. For several low-resource languages this beat translating the originals, and human judges preferred the simplified versions side by side.

Abstract · Simplify-then-Translate: Automatic Preprocessing for Black-Box Machine Translation

Black-box machine translation systems have proven incredibly useful for a variety of applications yet by design are hard to adapt, tune to a specific domain, or build on top of. In this work, we introduce a method to improve such systems via automatic pre-processing (APP) using sentence simplification. We first propose a method to automatically generate a large in-domain paraphrase corpus through back-translation with a black-box MT system, which is used to train a paraphrase model that "simplifies" the original sentence to be more conducive for translation. The model is used to preprocess source sentences of multiple low-resource language pairs. We show that this preprocessing leads to better translation performance as compared to non-preprocessed source sentences. We further perform side-by-side human evaluation to verify that translations of the simplified sentences are better than the original ones. Finally, we provide some guidance on recommended language pairs for generating the simplification model corpora by investigating the relationship between ease of translation of a language pair (as measured by BLEU) and quality of the resulting simplification model from back-translations of this language pair (as measured by SARI), and tie this into the downstream task of low-resource translation.

Sneha Mehta, Bahareh Azarnoush, Boris Chen, Avneesh Saluja, Vinith Misra, Ballav Bihani, Ritwik Kumar
arXiv:2005.11197 · cs.CL · submitted May 22, 2020 · updated May 27, 2020
abstract · pdf · html

add comment on HN

Interesting approach concept. I have tried using google translate and other machine translation on subtitles. Currently voice to text in many languages is very accurate so we are able to get subtitles of the spoken language quite accurately but machine translation of those to other languages are bad. The translation breaks when it comes to idioms or colloquial speech. Converting to simpler speech before translating should solve the problem of translation. That is what most human subtitle translators do they don't translate the exact words but more of what the speaker is trying to convey. This concept when used with normal forum posts on the web would probably give better results than the current machine translations when we use google translate plugin.