about
Translator2Vec: Understanding and Representing Human Post-Editors (arxiv.org)
2 points by sel1 on Jul 26, 2019 | hide | past | pdf | discuss on HN

In plain words: A dataset of 66,268 editing sessions records every keystroke, click, and pause, turning each translator's working style into a numeric profile. Watching the action sequence identifies the editor more accurately than comparing only the start and end text, and the profiles better predict editing time.

Abstract

The combination of machines and humans for translation is effective, with many studies showing productivity gains when humans post-edit machine-translated output instead of translating from scratch. To take full advantage of this combination, we need a fine-grained understanding of how human translators work, and which post-editing styles are more effective than others. In this paper, we release and analyze a new dataset with document-level post-editing action sequences, including edit operations from keystrokes, mouse actions, and waiting times. Our dataset comprises 66,268 full document sessions post-edited by 332 humans, the largest of the kind released to date. We show that action sequences are informative enough to identify post-editors accurately, compared to baselines that only look at the initial and final text. We build on this to learn and visualize continuous representations of post-editors, and we show that these representations improve the downstream task of predicting post-editing time.

António Góis, André F. T. Martins
arXiv:1907.10362 · cs.CL · submitted Jul 24, 2019
abstract · pdf · html · Accepted on MT Summit 2019; dataset available here: https://www.github.com/Unbabel/translator2vec; please cite as: @article{gois2019translator2vec, title={Translator2Vec: Understanding and Representing Human Post-Editors}, author={Góis, António and F. T. Martins, André}, year={2019}, publisher={European Association for Machine Translation} }

add comment on HN