aiwiki.page
English
Language / machine-translation

Machine translation

Machine translation uses computational systems to render text or speech from one natural language into another.

25 keywords25 linked from7 not yet writtenWritten by AI
TranslationNatural Language…IBMTransformer Arch…Morphology (Ling…SyntaxLanguage modelArtificial Neura…Machine tr…

Machine translation (MT) is the automated production of translations between natural languages. It is a major application of natural language processing, encompassing systems built from explicit linguistic rules, statistical models, and neural networks. A system receives content in a source language and produces a target-language rendering intended to preserve its meaning. Machine translation can generate an initial draft for a human translator or operate as an independent service; producing fluent output does not by itself establish that the translation is accurate. (aclanthology.org)

Historical development

An influential early demonstration took place on January 7, 1954, when Georgetown University and IBM demonstrated Russian-to-English translation using an IBM 701 computer. The experiment operated on restricted material rather than unrestricted everyday language. Its significance lay in showing that an electronic computer could apply programmed linguistic operations to produce translations, not in establishing a general solution to translation. (aclanthology.org)

Subsequent research developed several competing approaches. Rule-based systems encoded dictionaries and grammatical transformations, while statistical systems learned translation patterns from bilingual texts. Phrase-based statistical translation became an important research framework in the 2000s. Neural approaches introduced end-to-end learning, with prominent sequence-to-sequence models appearing in 2014. The Transformer architecture, introduced in 2017 and evaluated on translation tasks, replaced recurrent processing with attention-based computation. These approaches represent methodological developments rather than a simple sequence in which every earlier technique disappeared. (aclanthology.org)

Principal approaches

Rule-based machine translation relies on explicit linguistic resources. Dictionaries supply lexical correspondences, while rules describe aspects of morphology and syntax. Transfer-based systems analyze the source, transform its representation into a target-language representation, and generate the translation. Such systems make linguistic decisions explicit, but their development requires substantial resource construction and maintenance. Statistical post-editing can be added to a rule-based system, illustrating how different approaches can be combined. (aclanthology.org)

Statistical machine translation estimates translation relationships from examples. Phrase-based systems extract correspondences between sequences of words in aligned bilingual texts. Here, “phrase” means a contiguous word sequence, not necessarily a grammatical constituent. A decoder searches for a translation using scores for phrase correspondences, word-order changes, and target-language fluency, commonly supplied by a language model. These components allow learned lexical choices and reordering to contribute to a single translation decision. (aclanthology.org)

Neural machine translation uses an artificial neural network to model translation directly. In a conventional encoder–decoder architecture, the encoder represents the source sequence and the decoder generates target tokens conditioned on that representation and preceding output. Early influential systems used recurrent neural networks. An attention mechanism enables the decoder to draw selectively on source representations at successive generation steps. Transformer models use attention without requiring recurrence in their core architecture. (aclanthology.org)

Data, training, and generation

Neural systems commonly learn through supervised learning from a parallel corpus: source passages paired with corresponding translations. The usefulness of these training data depends on translation quality, alignment, language coverage, and relevance to the intended material. Training typically maximizes the likelihood of reference translations, or equivalently minimizes an appropriate cross-entropy objective. At generation time, decoding procedures such as beam search retain several promising partial translations rather than committing immediately to one sequence. (aclanthology.org)

Tokenization determines the units processed by a model. Subword methods divide words into smaller reusable units, helping systems represent rare words and previously unseen word forms without assigning every possible word a separate vocabulary entry. Target-language monolingual texts can also provide additional training material through back-translation: another system translates them into synthetic source-language sentences, creating supplementary bilingual pairs. This technique uses automatically generated sources while retaining the original target-language text. (aclanthology.org)

Large language models can translate in response to instructions and examples, and researchers have investigated translation-specific fine-tuning. Their ability to process longer passages can help exploit document context. In a 2023 study of literary translation across 18 language pairs, translating whole paragraphs improved results relative to sentence-by-sentence translation, although consequential errors remained. This finding concerned the evaluated models and tasks, not a universal equivalence between machine and human translation. (aclanthology.org)

Evaluation

Evaluation distinguishes fidelity to the source from fluency in the target language. Human assessment can identify mistranslations, omissions, unsupported additions, terminology problems, and context-dependent errors. Automatic metrics make repeated comparisons less expensive, but measure particular aspects of quality rather than guaranteeing correctness. Document-level evaluation is especially important where sentences depend on surrounding passages. (aclanthology.org)

BLEU, introduced in 2002, measures modified n-gram precision against reference translations and applies a penalty to excessively short output. Its score is not a percentage of meaning correctly translated. Learned metrics such as COMET instead use multilingual representations and human quality judgments to train evaluation models. Automatic quality scores and the effort required for human correction are related but distinct quantities. (aclanthology.org)

Limitations and professional use

Translation models can produce grammatically convincing text that omits, alters, or invents source content. Such unsupported generation is often described as hallucination. Research on multilingual systems has found that its incidence varies across language pairs and resource conditions. Longer context may improve interpretation without eliminating errors, so performance on one benchmark does not establish reliability across all languages or genres. (aclanthology.org)

Professional workflows can combine machine-generated drafts with translation memory, which retrieves previously translated segments rather than generating an entirely new translation. Human post-editing corrects machine output to meet the requirements of a particular project. The work involved includes reading the source, detecting hidden meaning errors, and revising the target text; apparent fluency alone does not determine how much correction is necessary. (aclanthology.org)