Knowledge Graphs Effectiveness in Neural Machine Translation Improvement

Benyamin Ahmadnia; Bonnie J. Dorr; Parisa Kordjamshidi

doi:10.7494/csci.2020.21.3.3701

Knowledge Graphs Effectiveness in Neural Machine Translation Improvement

Computer Science ◽

10.7494/csci.2020.21.3.3701 ◽

2020 ◽

Vol 21 (3) ◽

Author(s):

Benyamin Ahmadnia ◽

Bonnie J. Dorr ◽

Parisa Kordjamshidi

Keyword(s):

Machine Translation ◽

Semantic Representation ◽

Language Translation ◽

Semantic Relations ◽

Training Data ◽

Target Language ◽

Neural Machine Translation ◽

Source Language ◽

Knowledge Graphs ◽

Unknown Words

Neural Machine Translation (NMT) systems require a massive amount of Maintaining semantic relations between words during the translation process yields more accurate target-language output from Neural Machine Translation (NMT). Although difficult to achieve from training data alone, it is possible to leverage Knowledge Graphs (KGs) to retain source-language semantic relations in the corresponding target-language translation. The core idea is to use KG entity relations as embedding constraints to improve the mapping from source to target. This paper describes two embedding constraints, both of which employ Entity Linking (EL)---assigning a unique identity to entities---to associate words in training sentences with those in the KG: (1) a monolingual embedding constraint that supports an enhanced semantic representation of the source words through access to relations between entities in a KG; and (2) a bilingual embedding constraint that forces entity relations in the source-language to be carried over to the corresponding entities in the target-language translation. The method is evaluated for English-Spanish translation exploiting Freebase as a source of knowledge. Our experimental results show that exploiting KG information not only decreases the number of unknown words in the translation but also improves translation quality.

Download Full-text

A Study of Neural Machine Translation from Chinese to Urdu

Journal of Autonomous Intelligence ◽

10.32629/jai.v2i4.82 ◽

2020 ◽

Vol 2 (4) ◽

pp. 28

Author(s):

. Zeeshan

Keyword(s):

Machine Translation ◽

Chinese Language ◽

Language Translation ◽

Target Language ◽

Foreign Languages ◽

Neural Machine Translation ◽

Source Language ◽

Great Progress ◽

Score Method ◽

Translation Methods

Machine Translation (MT) is used for giving a translation from a source language to a target language. Machine translation simply translates text or speech from one language to another language, but this process is not sufficient to give the perfect translation of a text due to the requirement of identification of whole expressions and their direct counterparts. Neural Machine Translation (NMT) is one of the most standard machine translation methods, which has made great progress in the recent years especially in non-universal languages. However, local language translation software for other foreign languages is limited and needs improving. In this paper, the Chinese language is translated to the Urdu language with the help of Open Neural Machine Translation (OpenNMT) in Deep Learning. Firstly, a Chineseto Urdu language sentences datasets were established and supported with Seven million sentences. After that, these datasets were trained by using the Open Neural Machine Translation (OpenNMT) method. At the final stage, the translation was compared to the desired translation with the help of the Bleu Score Method.

Download Full-text

Controlling Neural Machine Translation Formality with Synthetic Supervision

Proceedings of the AAAI Conference on Artificial Intelligence ◽

10.1609/aaai.v34i05.6379 ◽

2020 ◽

Vol 34 (05) ◽

pp. 8568-8575

Author(s):

Xing Niu ◽

Marine Carpuat

Keyword(s):

Machine Translation ◽

Target Language ◽

Sentence Pair ◽

English Sentence ◽

Neural Machine Translation ◽

Source Language ◽

Training Scheme ◽

Training Examples ◽

Language Content ◽

Missing Element

This work aims to produce translations that convey source language content at a formality level that is appropriate for a particular audience. Framing this problem as a neural sequence-to-sequence task ideally requires training triplets consisting of a bilingual sentence pair labeled with target language formality. However, in practice, available training examples are limited to English sentence pairs of different styles, and bilingual parallel sentences of unknown formality. We introduce a novel training scheme for multi-task models that automatically generates synthetic training triplets by inferring the missing element on the fly, thus enabling end-to-end training. Comprehensive automatic and human assessments show that our best model outperforms existing models by producing translations that better match desired formality levels while preserving the source meaning.1

Download Full-text

Query Expansion for Slovak to Bulgarian Language Machine Translation using Parallel Search

WSEAS TRANSACTIONS ON SYSTEMS AND CONTROL ◽

10.37394/23203.2021.16.30 ◽

2021 ◽

pp. 351-357

Author(s):

VELISLAVA STOYKOVA ◽

DANIELA MAJCHRAKOVA

Keyword(s):

Machine Translation ◽

Query Expansion ◽

Statistical Approach ◽

Semantic Relations ◽

Target Language ◽

Parallel Search ◽

Keyword Query ◽

Source Language ◽

Standard Presentation ◽

Standard Semantic

The paper presents results of the application of a statistical approach for Slovak to Bulgarian language machine translation. It uses Information Retrieval inspired search techniques and employs sever alalgorithmic steps of parallel statistical search with query expansion in Slovak-Bulgarian EUROPARL 7 Corpus using the Sketch Engine software and its scoring. The search includes the generation of concordances,collocations, word sketch differences, word sketches, and thesauri of the studied keyword (query) by using a statistical scoring, which is regarded as intermediate (inter-lingual) semantic standard presentation by means of which the studied keyword (from the source language) is mapped together with its possible translation equivalents (onto the target language. The results present the study of adjectival collocabillity in both Slovak and Bulgarian language from the corpus of political speech texts outlining the standard semantic relations based on the evaluation of statistical scoring. Finally, the advantages and shortcomings of the approach are discussed.

Download Full-text

Translation of Sentence Lampung-Indonesian Languages with Neural Machine Translation Attention Based Approach

Inovasi Pembangunan : Jurnal Kelitbangan ◽

10.35450/jip.v6i02.97 ◽

2018 ◽

Vol 6 (02) ◽

pp. 191-206

Author(s):

Zaenal Abidin

Keyword(s):

Machine Translation ◽

Language Translation ◽

Training Process ◽

Neural Machine Translation ◽

New Approach ◽

Source Language ◽

Average Value ◽

Approach Method ◽

Stable Vectors ◽

Network Component

In this research, automatically Lampung language translation into the Indonesian language was using neural machine translation (NMT) attention based approach. NMT, a new approach method in machine translation technology, that has worked by combining the encoder and decoder. The encoder in NMT is a recurrent neural network component that encrypts the source language to several length-stable vectors and the decoder is a recurrent neural networks component that generates translation result comprehensive. NMT Research has begun with creating a pair of 3000 parallel sentences of Lampung language (api dialect) and Indonesian language. Then it continues to decide the NMT parameter model for the data training process. The next step is building NMT model and evaluate it. The testing of this approach has used 25 single sentences without out-of-vocabulary (OOV), 25 single sentences with OOV, 25 plural sentences without OOV, and 25 plural sentences with OOV. The testing translation result using NMT attention shows the bilingual evaluation understudy (BLEU) an average value is 51, 96 %.

Download Full-text

An Experimental Platform for Cross-Language Document Retrieval

Applied Mechanics and Materials ◽

10.4028/www.scientific.net/amm.284-287.3325 ◽

2013 ◽

Vol 284-287 ◽

pp. 3325-3329

Author(s):

Long Yue Wang ◽

Derek F. Wong ◽

Lidia S. Chao

Keyword(s):

Machine Translation ◽

Statistical Machine Translation ◽

Document Retrieval ◽

Training Data ◽

Target Language ◽

Source Language ◽

Experimental Platform ◽

Precision Evaluation ◽

Query Generation ◽

Cross Language

This paper presents a proposed Cross-Language Document Retrieval experimental platform integrated with preprocessing of training data, document translation, query generation, document retrieval and precision evaluation modules. Given a certain document in source language, it will be translated into target language by statistical machine translation module which is trained by selected training data. The query generation module then selects the most relevant words in the translated version of the document as searching query. After all the documents in the target language are ranked by the document retrieval module, the system will choose the N-best documents as its target language versions. Finally, the results can be evaluated by precision evaluator, which can reflect the merits of the strategies. Experimental results showed that this platform was effective and achieved very good performance.

Download Full-text

TRANSLATION OF ENGLISH TASKS INTO INDONESIAN THROUGH ONLINE MACHINE TRANSLATION PROGRAM

IJER - INDONESIAN JOURNAL OF EDUCATIONAL REVIEW ◽

10.21009/ijer.04.01.10 ◽

2017 ◽

Vol 4 (1) ◽

pp. 103

Author(s):

Emzir Emzir ◽

Ninuk Lustyantie ◽

Akbar Akbar

Keyword(s):

Machine Translation ◽

Language Translation ◽

Doctoral Program ◽

Deep Understanding ◽

Language Education ◽

Target Language ◽

State University ◽

Optimum Result ◽

Source Language ◽

Different Cultures

The objective of this research is to obtain a deep understanding about the online machine translation of graduate students in the Language Education Doctoral Program of State University of Jakarta, Indonesia, from source language to target language in order to achieve equivalence in the subject of Language Translation and Education. The approach used is qualitative approach with ethnography method. The translation process is conducted by writing down words or copying-pasting sentences to be translated and then those words/sentences will be automatically translated by machine translation. A repetitive edit, revision and correction process shall be first performed in order to get an optimum result i.e. translated sentences are equal in textual and meanings. The deviations occur due to inaccurate equivalents caused by different cultures between the source language and target language as well as the scope of translated language scientific field. The used strategy is a literal translation. Based on the research results, the translation of English tasks to Indonesian through the online translation program is very useful to facilitate the students’ lecturing process in completing their tasks.

Download Full-text

Language To Language Translation System

International Journal of Scientific Research in Computer Science Engineering and Information Technology ◽

10.32628/cseit206363 ◽

2020 ◽

pp. 289-293

Author(s):

Ms Pratheeksha ◽

Pratheeksha Rai ◽

Ms Vijetha

Keyword(s):

Speech Recognition ◽

Machine Translation ◽

Automatic Speech Recognition ◽

Speech Synthesis ◽

Language Translation ◽

Target Language ◽

Translation System ◽

Text To Speech ◽

Source Language ◽

Text To Speech Synthesis

The system used in Language to Language Translation is the phrases spoken in one language are immediately spoken in other language by the device. Language to Language Translation is a three steps software process which includes Automatic Speech Recognition, Machine Translation and Voice Synthesis. Language to Language system includes the major speech translation projects using different approaches for Speech Recognition, Translation and Text to Speech synthesis highlighting the major pros and cons for the approach being used. Language translation is a process that takes the conversational phrase in one language as an input and translated speech phrases in another language as the output. The three components of language-to-language translation are connected in a sequential order. Automatic Speech Recognition (ASR) is responsible for converting the spoken phrases of source language to the text in the same language followed by machine translation which translates the source language to next target language text and finally the speech synthesizer is responsible for text to speech conversion of target language.

Download Full-text

Unraveling the Contribution of Image Captioning and Neural Machine Translation for Multimodal Machine Translation

Prague Bulletin of Mathematical Linguistics ◽

10.1515/pralin-2017-0020 ◽

2017 ◽

Vol 108 (1) ◽

pp. 197-208 ◽

Cited By ~ 2

Author(s):

Chiraag Lala ◽

Pranava Madhyastha ◽

Josiah Wang ◽

Lucia Specia

Keyword(s):

Recent Work ◽

Machine Translation ◽

Visual Information ◽

Target Language ◽

Image Captioning ◽

Neural Machine Translation ◽

Source Language ◽

Translation Quality ◽

Depth Study ◽

Image Descriptions

AbstractRecent work on multimodal machine translation has attempted to address the problem of producing target language image descriptions based on both the source language description and the corresponding image. However, existing work has not been conclusive on the contribution of visual information. This paper presents an in-depth study of the problem by examining the differences and complementarities of two related but distinct approaches to this task: textonly neural machine translation and image captioning. We analyse the scope for improvement and the effect of different data and settings to build models for these tasks. We also propose ways of combining these two approaches for improved translation quality.

Download Full-text

Handling Unknown Words in Neural Machine Translation System

2020 International Conference on Decision Aid Sciences and Application (DASA) ◽

10.1109/dasa51403.2020.9317169 ◽

2020 ◽

Author(s):

Kamal Deep Garg ◽

Jatin Gupta ◽

Vandana Saini

Keyword(s):

Machine Translation ◽

Translation System ◽

Neural Machine Translation ◽

Machine Translation System ◽

Unknown Words

Download Full-text

Improved neural machine translation for low-resource English–Assamese pair

Journal of Intelligent & Fuzzy Systems ◽

10.3233/jifs-219260 ◽

2021 ◽

pp. 1-12

Author(s):

Sahinur Rahman Laskar ◽

Abdullah Faiz Ur Rahman Khilji ◽

Partha Pakray ◽

Sivaji Bandyopadhyay

Keyword(s):

Machine Translation ◽

Data Augmentation ◽

Language Translation ◽

Linguistically Diverse ◽

Neural Machine Translation ◽

Low Resource ◽

Parallel Data ◽

The World ◽

Translation Accuracy ◽

Vocabulary Problems

Language translation is essential to bring the world closer and plays a significant part in building a community among people of different linguistic backgrounds. Machine translation dramatically helps in removing the language barrier and allows easier communication among linguistically diverse communities. Due to the unavailability of resources, major languages of the world are accounted as low-resource languages. This leads to a challenging task of automating translation among various such languages to benefit indigenous speakers. This article investigates neural machine translation for the English–Assamese resource-poor language pair by tackling insufficient data and out-of-vocabulary problems. We have also proposed an approach of data augmentation-based NMT, which exploits synthetic parallel data and shows significantly improved translation accuracy for English-to-Assamese and Assamese-to-English translation and obtained state-of-the-art results.

Download Full-text