A Character Level Based and Word Level Based Approach for Chinese-Vietnamese Machine Translation

Chinese and Vietnamese have the same isolated language; that is, the words are not delimited by spaces. In machine translation, word segmentation is often done first when translating from Chinese or Vietnamese into different languages (typically English) and vice versa. However, it is a matter for consideration that words may or may not be segmented when translating between two languages in which spaces are not used between words, such as Chinese and Vietnamese. Since Chinese-Vietnamese is a low-resource language pair, the sparse data problem is evident in the translation system of this language pair. Therefore, while translating, whether it should be segmented or not becomes more important. In this paper, we propose a new method for translating Chinese to Vietnamese based on a combination of the advantages of character level and word level translation. In addition, a hybrid approach that combines statistics and rules is used to translate on the word level. And at the character level, a statistical translation is used. The experimental results showed that our method improved the performance of machine translation over that of character or word level translation.

Download Full-text

A Switching Hybrid Approach to Improve Sparse Data Problem of Collaborative Filtering Recommender System

International Journal for Research in Applied Science and Engineering Technology ◽

10.22214/ijraset.2020.31095 ◽

2020 ◽

Vol 8 (8) ◽

pp. 1060-1064

Author(s):

Tuyet-Van Tran Thi

Keyword(s):

Collaborative Filtering ◽

Recommender System ◽

Hybrid Approach ◽

Sparse Data ◽

Data Problem ◽

Sparse Data Problem

Download Full-text

English-Dogri Translation System using MOSES

Circulation in Computer Science ◽

10.22632/ccs-2016-251-25 ◽

2016 ◽

Vol 1 (1) ◽

pp. 45-49

Author(s):

Avinash Singh ◽

Asmeet Kour ◽

Shubhnandan S. Jamwal

Keyword(s):

Natural Language Processing ◽

Machine Translation ◽

Language Processing ◽

Statistical Machine Translation ◽

Translation System ◽

Parallel Corpus ◽

English System ◽

Machine Translation System ◽

Translation Machine ◽

Language Pair

The objective behind this paper is to analyze the English-Dogri parallel corpus translation. Machine translation is the translation from one language into another language. Machine translation is the biggest application of the Natural Language Processing (NLP). Moses is statistical machine translation system allow to train translation models for any language pair. We have developed translation system using Statistical based approach which helps in translating English to Dogri and vice versa. The parallel corpus consists of 98,973 sentences. The system gives accuracy of 80% in translating English to Dogri and the system gives accuracy of 87% in translating Dogri to English system.

Download Full-text

Analyzing Subword Techniques to Improve English to Sinhala Neural Machine Translation

International Journal of Asian Language Processing ◽

10.1142/s2717554520500174 ◽

2021 ◽

pp. 2050017

Author(s):

Rashmini Naranpanawa ◽

Ravinga Perera ◽

Thilakshi Fonseka ◽

Uthayasanker Thayasivam

Keyword(s):

Machine Translation ◽

State Of The Art ◽

Statistical Machine Translation ◽

Translation System ◽

Rare Word ◽

Neural Machine Translation ◽

Parallel Corpus ◽

Low Resource ◽

Word Level ◽

Morphologically Rich Languages

Neural machine translation (NMT) is a remarkable approach which performs much better than the Statistical machine translation (SMT) models when there is an abundance of parallel corpus. However, vanilla NMT is primarily based upon word-level with a fixed vocabulary. Therefore, low resource morphologically rich languages such as Sinhala are mostly affected by the out of vocabulary (OOV) and Rare word problems. Recent advancements in subword techniques have opened up opportunities for low resource communities by enabling open vocabulary translation. In this paper, we extend our recently published state-of-the-art EN-SI translation system using the transformer and explore standard subword techniques on top of it to identify which subword approach has a greater effect on English Sinhala language pair. Our models demonstrate that subword segmentation strategies along with the state-of-the-art NMT can perform remarkably when translating English sentences into a rich morphology language regardless of a large parallel corpus.

Download Full-text

Attention-Based Syllable Level Neural Machine Translation System for Myanmar to English Language Pair

International Journal on Natural Language Computing ◽

10.5121/ijnlc.2019.8201 ◽

2019 ◽

Vol 8 (2) ◽

pp. 01-11

Author(s):

Yi Mon Shwe Sin ◽

Khin Mar Soe

Keyword(s):

Machine Translation ◽

English Language ◽

Translation System ◽

Neural Machine Translation ◽

Machine Translation System ◽

Language Pair

Download Full-text

The Critical Technology Development Status of Machine Translation

Advanced Materials Research ◽

10.4028/www.scientific.net/amr.791-793.1622 ◽

2013 ◽

Vol 791-793 ◽

pp. 1622-1625

Author(s):

Dan Han ◽

Zhi Han Yu

Keyword(s):

Machine Translation ◽

Technology Development ◽

Statistical Machine Translation ◽

Word Segmentation ◽

Translation System ◽

Chinese Word ◽

Segmentation Method ◽

Chinese Word Segmentation ◽

Critical Technology ◽

Translation Machine

In this article, we mainly introduce some basic concepts about machine translation. Machine translation means translating a natural language text to another by software. It can be divided into two categories: rule-based and corpus-based. IBM's statistical machine translation, Microsoft's multi-language machine translation project, AT & T's voice translation system and CMUs PANGLOSS system are three typical machine translation systems. Due to sentences are constructed by words continuously in Chinese. Chinese word segmentation is very essential. Three methods of Chinese word segmentation: segmentation methods based on string matching, segmentation method based on the understanding and segmentation method based on the statistics.

Download Full-text

N-gram based Machine Translation for English-Assamese: Two Languages with High Syntactical Dissimilarity

International Journal of Engineering and Advanced Technology - Regular Issue ◽

10.35940/ijeat.b2320.129219 ◽

2019 ◽

Vol 9 (2) ◽

pp. 2940-2949

Keyword(s):

Machine Translation ◽

Northeast India ◽

Translation System ◽

System Efficiency ◽

Translation Process ◽

The People ◽

Machine Translation System ◽

N Gram ◽

Assamese Language ◽

Language Pair

To bridge the language constraint of the people residing in northeastern region of India, machine translation system is a necessity. Large number of people in this region cannot access many services due to the language incomprehensibility. Among several languages spoken, Assamese is one of the major languages used in northeast India. Machine translation for Assamese language is limited compared to other languages. As a result, large number of people using Assamese language cannot avail lots of benefits associated with it. This paper has focused on the development of the English to Assamese translation system using n-gram model. The n-gram model works very well with the language pair having high dissimilarity in syntax compared to other models. The value of n has a very big role in the quality and efficiency of the system. Bilingual Evaluation Understudy (BLEU) score differs significantly with the change of the n-gram. This model uses tuples to reduce the consumption of excess memory and to accelerate the translation process. Parallel corpus has been used for training the n-gram based decoder called MARIE. The number of translation units extracted using n-gram model is much less than the translation units extracted using phrase based model. This has a high impact on system efficiency.

Download Full-text

Improving Statistical Parser by Recognition of Chinese Number and Quantifier Prefix in Machine Translation

Applied Mechanics and Materials ◽

10.4028/www.scientific.net/amm.427-429.1841 ◽

2013 ◽

Vol 427-429 ◽

pp. 1841-1844

Author(s):

Wen Xiong ◽

Yao Hong Jin ◽

Zhi Ying Liu

Keyword(s):

Experimental Data ◽

Machine Translation ◽

Word Segmentation ◽

Experimental Results ◽

Maximum Matching ◽

Recognition Method ◽

Matching Method ◽

Rule Based ◽

Processing Module ◽

Special Language

By studying the Chinese number and quantifier prefix (CNQP) as a special language phenomenon in machine translation, this paper presents a CNQP recognition method, which is rule based and independent of word segmentation. The method expressed CNQPs compositions using Backus-Naur Form (BNF), and took the numeral as the active information and the quantifiers as the boundaries of the CNQPs. To avoid the word segmentation noise, a forward maximum matching method was used for obtaining the compositions of the CNQPs, which can be fed into the statistical parser for the analysis of the Chinese sentences. The experimental results indicate the proposed method as a pre-processing module can effectively improve the parsing results of the statistical parser without retraining on experimental data constructed manually, which can further enhance the translation qualities.

Download Full-text

Rule-Based Machine Translation for the Italian–Sardinian Language Pair

Prague Bulletin of Mathematical Linguistics ◽

10.1515/pralin-2017-0022 ◽

2017 ◽

Vol 108 (1) ◽

pp. 221-232

Author(s):

Francis M. Tyers ◽

Hèctor Alòs i Font ◽

Gianfranco Fronteddu ◽

Adrià Martín-Mor

Keyword(s):

Machine Translation ◽

Translation System ◽

Rule Based ◽

Romance Language ◽

Machine Translation System ◽

The Mediterranean ◽

Language Pair

AbstractThis paper describes the process of creation of the first machine translation system from Italian to Sardinian, a Romance language spoken on the island of Sardinia in the Mediterranean. The project was carried out by a team of translators and computational linguists. The article focuses on the technology used (Rule-Based Machine Translation) and on some of the rules created, as well as on the orthographic model used for Sardinian.

Download Full-text

Neural text normalization with adapted decoding and POS features

Natural Language Engineering ◽

10.1017/s1351324919000391 ◽

2019 ◽

Vol 25 (5) ◽

pp. 585-605

Author(s):

T. Ruzsics ◽

M. Lusetti ◽

A. Göhring ◽

T. Samardžić ◽

E. Stark

Keyword(s):

Machine Translation ◽

Computer Mediated Communication ◽

Statistical Machine Translation ◽

Language Model ◽

Translation System ◽

Swiss German ◽

Part Of Speech ◽

Word Level ◽

Speech Transcription ◽

Text Normalization

AbstractText normalization is the task of mapping noncanonical language, typical of speech transcription and computer-mediated communication, to a standardized writing. This task is especially important for languages such as Swiss German, with strong regional variation and no written standard. In this paper, we propose a novel solution for normalizing Swiss German WhatsApp messages using the encoder–decoder neural machine translation (NMT) framework. We enhance the performance of a plain character-level NMT model with the integration of a word-level language model and linguistic features in the form of part-of-speech (POS) tags. The two components are intended to improve the performance by addressing two specific issues: the former is intended to improve the fluency of the predicted sequences, whereas the latter aims at resolving cases of word-level ambiguity. Our systematic comparison shows that our proposed solution results in an improvement over a plain NMT system and also over a comparable character-level statistical machine translation system, considered the state of the art in this task till recently. We perform a thorough analysis of the compared systems’ output, showing that our two components produce indeed the intended, complementary improvements.

Download Full-text

Cadlaws – An English–French Parallel Corpus of Legally Equivalent Documents

Mutatis Mutandis Revista Latinoamericana de Traducción ◽

10.17533/udea.mut.v14n2a10 ◽

2021 ◽

Vol 14 (2) ◽

pp. 494-508

Author(s):

Francina Sole-Mauri ◽

Pilar Sánchez-Gijón ◽

Antoni Oliver

Keyword(s):

Machine Translation ◽

Translation System ◽

Neural Machine Translation ◽

Parallel Corpus ◽

Legal Documents ◽

Legal Traditions ◽

Corpus Construction ◽

Machine Translation System ◽

French Corpus ◽

Language Pair

This article presents Cadlaws, a new English–French corpus built from Canadian legal documents, and describes the corpus construction process and preliminary statistics obtained from it. The corpus contains over 16 million words in each language and includes unique features since it is composed of documents that are legally equivalent in both languages but not the result of a translation. The corpus is built upon enactments co-drafted by two jurists to ensure legal equality of each version and to reflect the concepts, terms and institutions of two legal traditions. In this article the corpus definition as a parallel corpus instead of a comparable one is also discussed. Cadlaws has been pre-processed for machine translation and baseline Bilingual Evaluation Understudy (bleu), a score for comparing a candidate translation of text to a gold-standard translation of a neural machine translation system. To the best of our knowledge, this is the largest parallel corpus of texts which convey the same meaning in this language pair and is freely available for non-commercial use.

Download Full-text