Performance of Speaker Independent Language Identification System Under Various Noise Environments

Author(s):  
Phani Kumar Polasi ◽  
K. Sri Rama Krishna
2012 ◽  
Author(s):  
Pavel Matějka ◽  
Oldřich Plchot ◽  
Mehdi Soufifar ◽  
Ondřej Glembek ◽  
Luis Fernando D'Haro ◽  
...  

2020 ◽  
Vol 08 (01) ◽  
pp. 113-131
Author(s):  
Ridouane Tachicart ◽  
Karim Bouzoubaa

With the increase of Web use in Morocco today, Internet has become an important source of information. Specifically, across social media, the Moroccan people use several languages in their communication leaving behind unstructured user-generated text (UGT) that presents several opportunities for Natural Language Processing. Among the languages found in this data, Moroccan Arabic (MA) stands with an important content and several features. In this paper, we investigate online written text generated by Moroccan users in social media with an emphasis on Moroccan Arabic. For this purpose, we follow several steps, using some tools such as a language identification system, in order to conduct a deep study of this data. The most interesting findings that have emerged are the use of code-switching, multi-script and low amount of words in the Moroccan UGT. Moreover, we used the investigated data in order to build a new Moroccan language resource. The latter consists in building a Moroccan words orthographic variants lexicon following an unsupervised approach and using character neural embedding. This lexicon can be useful for several NLP tasks such as spelling normalization.


Sign in / Sign up

Export Citation Format

Share Document