Human Evaluation of Machine Translation Output
This is another post from Juan Rowda, this time co-written with his colleague Olga Pospelova, that provides useful information on this subject. Human evaluation of MT output is key to understanding several fundamental questions about MT technology including: Is the MT engine improving? Do the automated metrics (which MT systems developers use to guide development strategy) correlate with independent human judgments of the same MT output? How difficult/easy is the PEMT output task? What is the fair and reasonable PEMT rate? The difficulty with human evaluation of MT output is closely related to the difficulty associated with any discussion of translation quality in general in the industry. It is challenging to come up with an approach that is consistent over time and across different people. Being scientifically objective is particularly a challenge. So, like much else in MT, estimates have to be made, and some rigor needs to be applied to ensure consistency over...