Multilingual Projection for Parsing Truly Low-Resource Languages

We propose a novel approach to cross-lingual part-of-speech tagging and dependency parsing for truly low-resource languages. Our annotation projection-based approach yields tagging and parsing models for over 100 languages. All that is needed are freely available parallel texts, and taggers and parsers for resource-rich languages. The empirical evaluation across 30 test languages shows that our method consistently provides top-level accuracies, close to established upper bounds, and outperforms several competitive baselines.

Download Full-text

Unsupervised Cross-Lingual Part-of-Speech Tagging for Truly Low-Resource Scenarios

10.18653/v1/2020.emnlp-main.391 ◽

2020 ◽

Author(s):

Ramy Eskander ◽

Smaranda Muresan ◽

Michael Collins

Keyword(s):

Low Resource ◽

Part Of Speech Tagging ◽

Part Of Speech ◽

Cross Lingual ◽

Speech Tagging

Download Full-text

An Analysis of Capsule Networks for Part of Speech Tagging in High- and Low-resource Scenarios

10.18653/v1/2020.insights-1.10 ◽

2020 ◽

Author(s):

Andrew Zupon ◽

Faiz Rafique ◽

Mihai Surdeanu

Keyword(s):

Low Resource ◽

Part Of Speech Tagging ◽

Part Of Speech ◽

Speech Tagging

Download Full-text

Sequence Mixup for Zero-Shot Cross-Lingual Part-Of-Speech Tagging

10.18653/v1/2021.mrl-1.22 ◽

2021 ◽

Author(s):

Megh Thakkar ◽

Vishwa Shah ◽

Ramit Sawhney ◽

Debdoot Mukherjee

Keyword(s):

Part Of Speech Tagging ◽

Part Of Speech ◽

Cross Lingual ◽

Speech Tagging

Download Full-text

A Universal Feature Schema for Rich Morphological Annotation and Fine-Grained Cross-Lingual Part-of-Speech Tagging

Systems and Frameworks for Computational Morphology - Communications in Computer and Information Science ◽

10.1007/978-3-319-23980-4_5 ◽

2015 ◽

pp. 72-93 ◽

Cited By ~ 1

Author(s):

John Sylak-Glassman ◽

Christo Kirov ◽

Matt Post ◽

Roger Que ◽

David Yarowsky

Keyword(s):

Fine Grained ◽

Part Of Speech Tagging ◽

Part Of Speech ◽

Universal Feature ◽

Cross Lingual ◽

Speech Tagging

Download Full-text

Distant Supervision from Disparate Sources for Low-Resource Part-of-Speech Tagging

10.18653/v1/d18-1061 ◽

2018 ◽

Cited By ~ 1

Author(s):

Barbara Plank ◽

Željko Agić

Keyword(s):

Low Resource ◽

Part Of Speech Tagging ◽

Distant Supervision ◽

Part Of Speech ◽

Speech Tagging

Download Full-text

Part of speech tagging for Arabic

Natural Language Engineering ◽

10.1017/s1351324911000325 ◽

2011 ◽

Vol 18 (4) ◽

pp. 521-548 ◽

Cited By ~ 8

Author(s):

SANDRA KÜBLER ◽

EMAD MOHAMED

Keyword(s):

Computational Linguistics ◽

Automatic Segmentation ◽

Data Sparseness ◽

Part Of Speech Tagging ◽

Pos Tagging ◽

Part Of Speech ◽

Novel Approach ◽

Pos Tagger ◽

Whole Word ◽

Speech Tagging

AbstractThis paper presents an investigation of part of speech (POS) tagging for Arabic as it occurs naturally, i.e. unvocalized text (without diacritics). We also do not assume any prior tokenization, although this was used previously as a basis for POS tagging. Arabic is a morphologically complex language, i.e. there is a high number of inflections per word; and the tagset is larger than the typical tagset for English. Both factors, the second one being partly dependent on the first, increase the number of word/tag combinations, for which the POS tagger needs to find estimates, and thus they contribute to data sparseness. We present a novel approach to Arabic POS tagging that does not require any pre-processing, such as segmentation or tokenization: whole word tagging. In this approach, the complete word is assigned a complex POS tag, which includes morphological information. A competing approach investigates the effect of segmentation and vocalization on POS tagging to alleviate data sparseness and ambiguity. In the segmentation-based approach, we first automatically segment words and then POS tags the segments. The complex tagset encompasses 993 POS tags, whereas the segment-based tagset encompasses only 139 tags. However, segments are also more ambiguous, thus there are more possible combinations of segment tags. In realistic situations, in which we have no information about segmentation or vocalization, whole word tagging reaches the highest accuracy of 94.74%. If gold standard segmentation or vocalization is available, including this information improves POS tagging accuracy. However, while our automatic segmentation and vocalization modules reach state-of-the-art performance, their performance is not reliable enough for POS tagging and actually impairs POS tagging performance. Finally, we investigate whether a reduction of the complex tagset to the Extra-Reduced Tagset as suggested by Habash and Rambow (Habash, N., and Rambow, O. 2005. Arabic tokenization, part-of-speech tagging and morphological disambiguation in one fell swoop. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL), Ann Arbor, MI, USA, pp. 573–80) will alleviate the data sparseness problem. While the POS tagging accuracy increases due to the smaller tagset, a closer look shows that using a complex tagset for POS tagging and then converting the resulting annotation to the smaller tagset results in a higher accuracy than tagging using the smaller tagset directly.

Download Full-text

Cross-Lingual Part-of-Speech Tagging through Ambiguous Learning

10.3115/v1/d14-1187 ◽

2014 ◽

Cited By ~ 7

Author(s):

Guillaume Wisniewski ◽

Nicolas Pécheux ◽

Souhir Gahbiche-Braham ◽

François Yvon

Keyword(s):

Part Of Speech Tagging ◽

Part Of Speech ◽

Cross Lingual ◽

Speech Tagging

Download Full-text

To Augment or Not to Augment? A Comparative Study on Text Augmentation Techniques for Low-Resource NLP

Computational Linguistics ◽

10.1162/coli_a_00425 ◽

2021 ◽

pp. 1-38

Author(s):

Gözde Gül Şahin

Keyword(s):

Language Models ◽

Semantic Role ◽

Semantic Role Labeling ◽

Dependency Parsing ◽

Low Resource ◽

Part Of Speech Tagging ◽

High Resource ◽

Part Of Speech ◽

Augmentation Techniques ◽

Speech Tagging

Abstract Data-hungry deep neural networks have established themselves as the defacto standard for many NLP tasks including the traditional sequence tagging ones. Despite their state-of-the-art performance on high-resource languages, they still fall behind of their statistical counter-parts in low-resource scenarios. One methodology to counter attack this problem is text augmentation, i.e., generating new synthetic training data points from existing data. Although NLP has recently witnessed a load of textual augmentation techniques, the field still lacks a systematic performance analysis on a diverse set of languages and sequence tagging tasks. To fill this gap, we investigate three categories of text augmentation methodologies which perform changes on the syntax (e.g., cropping sub-sentences), token (e.g., random word insertion) and character (e.g., character swapping) levels.We systematically compare the methods on part-of-speech tagging, dependency parsing and semantic role labeling for a diverse set of language families using various models including the architectures that rely on pretrained multilingual contextualized language models such as mBERT. Augmentation most significantly improves dependency parsing, followed by part-of-speech tagging and semantic role labeling. We find the experimented techniques to be effective on morphologically rich languages in general rather than analytic languages such as Vietnamese. Our results suggest that the augmentation techniques can further improve over strong baselines based on mBERT, especially for dependency parsing. We identify the character-level methods as the most consistent performers, while synonym replacement and syntactic augmenters provide inconsistent improvements. Finally, we discuss that the results most heavily depend on the task, language pair (e.g., syntactic-level techniques mostly benefit higher-level tasks and morphologically richer languages), and the model type (e.g., token-level augmentation provide significant improvements for BPE, while character-level ones give generally higher scores for char and mBERT based models).

Download Full-text