Posts

Post-editese is real

Image
Ever since machine translation was introduced into the professional translation industry, there have been questions about what the impact would be on a final delivered translation service product. For much of the history of MT many translators claimed that while translation production work using a post-edited MT (PEMT) process was faster, the final product was not as good. The research suggests that this has been true from a strictly linguistic perspective, but many of us also know that PEMT worked quite successfully with technical content especially with terminology and consistency even in the days of SMT and RBMT.  As NMT systems proliferate, we are at a turning point, and I suspect that we will see many more NMT systems that are in fact seen as providing useful output that clearly enhances translator productivity, especially on output from systems built by experts. NMT will also quite likely have an influence on the output quality and the difference is also likely to become less...

In a Funk about BLEU

Image
This is a more fleshed-out version of a blog post by Pete Smith and Henry Anderson of the University of Texas at Arlington already published on SDL.com . They describe initial results from a research project they are conducting on MT system quality measurement and related issues.  MT quality measurement, like human translation quality measurement, has been a difficult and challenging subject for both the translation industry and for many MT researchers and systems developers as the most commonly used metric BLEU, is now quite widely understood to be of especially limited value with NMT systems.  Most of the other text-matching NLP scoring measures are just as suspect, and practitioners are reluctant to adopt them as they are either difficult to implement, or the interpretation pitfalls and nuances of these other measures are not well understood. They all can generate a numeric score based on various calculations of Precision and Recall that need to be interpreted with great c...

Adapting Neural MT to Support Digital Transformation

Image
We live in an era where the issue of digital transformation is increasingly recognized as a primary concern, and a key focus of executive management teams in global enterprises. The stakes are high for businesses that fail to embrace change. Since 2000, almost half (52%) of Fortune 500 companies have either gone bankrupt, been acquired, or ceased to exist as a result of digital disruption. It’s also estimated that 75% of today’s S&P 500 will be replaced by 2027, according to Innosight Research.  Responding effectively to the realities of the digital world have now become a matter of survival as well a means to build long term competitive advantage. When we consider what is needed to drive digital transformation in addition to structural integration, we see that large volumes of current, relevant, and accurate content that support the buyer and customer journey are critical to enhancing the digital experience both in B2C and B2B scenarios.  Large volumes of relevant conte...

The Challenge of Open Source MT

Image
This is the raw, first draft, and slightly longer rambling version of a post already published on SDL.COM . MT is considered one of the most difficult problems in the general AI and machine learning field. In the field of artificial intelligence , the most difficult problems are informally known as AI-complete problems , implying that the difficulty of these computational problems is equivalent to that of solving the central artificial intelligence problem— that is, making computers as intelligent as people. It is no surprise that humankind has been working on this problem for almost 70 years now, and is still quite some distance from having solved this problem. “To translate accurately, a machine must be able to understand the text. It must be able to follow the author's argument, so it must have some ability to reason. It must have extensive world knowledge so that it knows what is being discussed — it must at least be familiar with all the same commonsense facts that the average...

Understanding MT Quality - What Really Matters?

Image
This is the second post in our posts series on machine translation quality. Again this is a slightly less polished and raw variant of a version published on the SDL site . The first one focused on BLEU scores , which are often improperly used to make decisions on inferred MT quality, where it clearly is not the best metric to draw this inference. The reality of many of these comparisons today is that scores based on publicly available (i.e. not blind) news domain tests are being used by many companies and LSPs to select MT systems which translate IT, customer support, pharma, financial services domain related content. Clearly, this can only result in sub-optimal choices. The use of machine translation (MT) in the translation industry has historically been heavily focused on localization use cases, with the primary intention to improve efficiency, that is, speed up turnaround and reduce unit word cost. Indeed, machine translation post-editing (MTPE) has been instrumental in helping loca...