Evaluating effect of context window size, stemming and stop word removal on Hindi word sense disambiguation

Author(s):  
Satyendr Singh ◽  
Tanveer J. Siddiqui
2015 ◽  
Vol 8 (3) ◽  
pp. 21-42 ◽  
Author(s):  
Satyendr Singh ◽  
Tanveer J. Siddiqui

Karakas are an important constituent of Hindi language. Karaka relations express syntactico-semantic or semantico-syntactic relationship between verbs and nouns or pronouns in a sentence. They capture certain level of semantics closer to thematic relations, but different from it. A vibhakti is assigned to each karaka, in Paninian grammar. This paper investigates the role of karaka relations in Hindi Word Sense Disambiguation (WSD) by utilizing vibhaktis. Two supervised WSD algorithms were used for disambiguation. The first algorithm is based on conditional probability of co-occurring words and the second algorithm is Naïve Bayes (NB) classifier. The first algorithm utilizes various heuristics for analyzing the role of karakas in Hindi WSD. The authors obtained an improvement of 14.86% in precision by utilizing content words, vibhaktis and phrases containing them in context vector over the context vector of content words after dropping vibhaktis. A gain of 6.91% in precision was observed by using content words and vibhaktis in context vector over the context vector of content words after dropping vibhaktis of similar context window size. The authors obtained maximum precision of 50.73% by extracting vibhaktis in a ±3 window using WSD algorithm based on conditional probability of co-occurring words. They obtained maximum precision of 56.56% by extracting vibhaktis in a ±4 window using NB classifier.


Author(s):  
Manuel Ladron de Guevara ◽  
Christopher George ◽  
Akshat Gupta ◽  
Daragh Byrne ◽  
Ramesh Krishnamurti

2017 ◽  
Vol 132 ◽  
pp. 47-61 ◽  
Author(s):  
Yoan Gutiérrez ◽  
Sonia Vázquez ◽  
Andrés Montoyo

2005 ◽  
Vol 12 (5) ◽  
pp. 554-565 ◽  
Author(s):  
Martijn J. Schuemie ◽  
Jan A. Kors ◽  
Barend Mons

Sign in / Sign up

Export Citation Format

Share Document