Feature Selection Methods in Sentiment Analysis

Sentiment Analysis of Movie Reviews: A Study of Machine Learning Algorithms with Various Feature Selection Methods

International Journal of Computer Sciences and Engineering ◽

10.26438/ijcse/v5i9.113121 ◽

2017 ◽

Vol 5 (9) ◽

Cited By ~ 1

Author(s):

Rajwinder Kaur

Keyword(s):

Machine Learning ◽

Feature Selection ◽

Sentiment Analysis ◽

Learning Algorithms ◽

Machine Learning Algorithms ◽

Selection Methods

Get full-text (via PubEx)

Ordinal-based and frequency-based integration of feature selection methods for sentiment analysis

Expert Systems with Applications ◽

10.1016/j.eswa.2017.01.009 ◽

2017 ◽

Vol 75 ◽

pp. 80-93 ◽

Cited By ~ 23

Author(s):

Alireza Yousefpour ◽

Roliana Ibrahim ◽

Haza Nuzly Abdel Hamed

Keyword(s):

Feature Selection ◽

Sentiment Analysis ◽

Selection Methods

Get full-text (via PubEx)

Integrated Feature Selection Methods Using Metaheuristic Algorithms for Sentiment Analysis

Intelligent Information and Database Systems - Lecture Notes in Computer Science ◽

10.1007/978-3-662-49381-6_13 ◽

2016 ◽

pp. 129-140 ◽

Cited By ~ 1

Author(s):

Alireza Yousefpour ◽

Roliana Ibrahim ◽

Haza Nuzly Abdul Hamed ◽

Takeru Yokoi

Keyword(s):

Feature Selection ◽

Sentiment Analysis ◽

Metaheuristic Algorithms ◽

Selection Methods

Get full-text (via PubEx)

To use or not to use: Feature selection for sentiment analysis of highly imbalanced data

Natural Language Engineering ◽

10.1017/s1351324917000298 ◽

2017 ◽

Vol 24 (1) ◽

pp. 3-37 ◽

Cited By ~ 5

Author(s):

SANDRA KÜBLER ◽

CAN LIU ◽

ZEESHAN ALI SAYYED

Keyword(s):

Machine Learning ◽

Feature Selection ◽

Sentiment Analysis ◽

Information Gain ◽

Binary Classification ◽

Small Subset ◽

Large Set ◽

Learning Approaches ◽

Selection Methods ◽

Data Set

AbstractWe investigate feature selection methods for machine learning approaches in sentiment analysis. More specifically, we use data from the cooking platform Epicurious and attempt to predict ratings for recipes based on user reviews. In machine learning approaches to such tasks, it is a common approach to use word or part-of-speech n-grams. This results in a large set of features, out of which only a small subset may be good indicators for the sentiment. One of the questions we investigate concerns the extension of feature selection methods from a binary classification setting to a multi-class problem. We show that an inherently multi-class approach, multi-class information gain, outperforms ensembles of binary methods. We also investigate how to mitigate the effects of extreme skewing in our data set by making our features more robust and by using review and recipe sampling. We show that over-sampling is the best method for boosting performance on the minority classes, but it also results in a severe drop in overall accuracy of at least 6 per cent points.

Get full-text (via PubEx)

Comparison of feature selection methods for sentiment analysis on Turkish Twitter data

2017 25th Signal Processing and Communications Applications Conference (SIU) ◽

10.1109/siu.2017.7960388 ◽

2017 ◽

Cited By ~ 1

Author(s):

Tuba Parlar ◽

Esra Sarac ◽

Selma Ayse Ozel

Keyword(s):

Feature Selection ◽

Sentiment Analysis ◽

Selection Methods ◽

Twitter Data

Get full-text (via PubEx)

Feature Selection Methods in Sentiment Analysis and Sentiment Classification of Amazon Product Reviews

International Journal of Computer Trends and Technology ◽

10.14445/22312803/ijctt-v36p139 ◽

2016 ◽

Vol 36 (4) ◽

pp. 225-230 ◽

Cited By ~ 1

Author(s):

Tahura Shaikh ◽

◽

Deepa Deshpande

Keyword(s):

Feature Selection ◽

Sentiment Analysis ◽

Sentiment Classification ◽

Product Reviews ◽

Selection Methods

Get full-text (via PubEx)

Comparison of Feature Selection Methods for Sentiment Analysis

Advances in Artificial Intelligence - Lecture Notes in Computer Science ◽

10.1007/978-3-642-13059-5_30 ◽

2010 ◽

pp. 286-289 ◽

Cited By ~ 25

Author(s):

Chris Nicholls ◽

Fei Song

Keyword(s):

Feature Selection ◽

Sentiment Analysis ◽

Selection Methods

Get full-text (via PubEx)

Performance Assessment of Multiple Classifiers Based on Ensemble Feature Selection Scheme for Sentiment Analysis

Applied Computational Intelligence and Soft Computing ◽

10.1155/2018/8909357 ◽

2018 ◽

Vol 2018 ◽

pp. 1-12 ◽

Cited By ~ 4

Author(s):

Monalisa Ghosh ◽

Goutam Sanyal

Keyword(s):

Machine Learning ◽

Feature Selection ◽

Sentiment Analysis ◽

Gini Index ◽

Feature Vector ◽

Information Gain ◽

Feature Subset ◽

Selection Methods ◽

Prominent Feature ◽

Chi Square

Sentiment classification or sentiment analysis has been acknowledged as an open research domain. In recent years, an enormous research work is being performed in these fields by applying various numbers of methodologies. Feature generation and selection are consequent for text mining as the high-dimensional feature set can affect the performance of sentiment analysis. This paper investigates the inability or incompetency of the widely used feature selection methods (IG, Chi-square, and Gini Index) with unigram and bigram feature set on four machine learning classification algorithms (MNB, SVM, KNN, and ME). The proposed methods are evaluated on the basis of three standard datasets, namely, IMDb movie review and electronics and kitchen product review dataset. Initially, unigram and bigram features are extracted by applying n-gram method. In addition, we generate a composite features vector CompUniBi (unigram + bigram), which is sent to the feature selection methods Information Gain (IG), Gini Index (GI), and Chi-square (CHI) to get an optimal feature subset by assigning a score to each of the features. These methods offer a ranking to the features depending on their score; thus a prominent feature vector (CompIG, CompGI, and CompCHI) can be generated easily for classification. Finally, the machine learning classifiers SVM, MNB, KNN, and ME used prominent feature vector for classifying the review document into either positive or negative. The performance of the algorithm is measured by evaluation methods such as precision, recall, and F-measure. Experimental results show that the composite feature vector achieved a better performance than unigram feature, which is encouraging as well as comparable to the related research. The best results were obtained from the combination of Information Gain with SVM in terms of highest accuracy.

Get full-text (via PubEx)