Short text sentiment classification based on feature extension and ensemble classifier

A large-scale and high-quality training dataset is an important guarantee to learn an ideal classifier for text sentiment classification. However, manually constructing such a training dataset with sentiment labels is a labor-intensive and time-consuming task. Therefore, based on the idea of effectively utilizing unlabeled samples, a synthetical framework that covers the whole process of semi-supervised learning from seed selection, iterative modification of the training text set, to the co-training strategy of the classifier is proposed in this paper for text sentiment classification. To provide an important basis for selecting the seed texts and modifying the training text set, three kinds of measures—the cluster similarity degree of an unlabeled text, the cluster uncertainty degree of a pseudo-label text to a learner, and the reliability degree of a pseudo-label text to a learner—are defined. With these measures, a seed selection method based on Random Swap clustering, a hybrid modification method of the training text set based on active learning and self-learning, and an alternately co-training strategy of the ensemble classifier of the Maximum Entropy and Support Vector Machine are proposed and combined into our framework. The experimental results on three Chinese datasets (COAE2014, COAE2015, and a Hotel review, respectively) and five English datasets (Books, DVD, Electronics, Kitchen, and MR, respectively) in the real world verify the effectiveness of the proposed framework.

Download Full-text

Deep Neural Network for Short-Text Sentiment Classification

Database Systems for Advanced Applications - Lecture Notes in Computer Science ◽

10.1007/978-3-319-32055-7_15 ◽

2016 ◽

pp. 168-175 ◽

Cited By ~ 4

Author(s):

Xiangsheng Li ◽

Jianhui Pang ◽

Biyun Mo ◽

Yanghui Rao ◽

Fu Lee Wang

Keyword(s):

Neural Network ◽

Deep Neural Network ◽

Sentiment Classification ◽

Short Text

Download Full-text

An Ensemble-Classifier Based Approach for Multiclass Emotion Classification of Short Text

2018 7th International Conference on Reliability, Infocom Technologies and Optimization (Trends and Future Directions) (ICRITO) ◽

10.1109/icrito.2018.8748757 ◽

2018 ◽

Cited By ~ 1

Author(s):

Shivangi Chawla ◽

Monica Mehrotra

Keyword(s):

Ensemble Classifier ◽

Emotion Classification ◽

Short Text

Download Full-text

New Method for Sentiment Classification for Short Text

The Open Cybernetics & Systemics Journal ◽

10.2174/1874110x01509010601 ◽

2015 ◽

Vol 9 (1) ◽

pp. 601-607

Author(s):

Hao Fu

Keyword(s):

Sentiment Classification ◽

New Method ◽

Short Text

Download Full-text

Hybrid Ensemble Learning With Feature Selection for Sentiment Classification in Social Media

International Journal of Information Retrieval Research ◽

10.4018/ijirr.2020040103 ◽

2020 ◽

Vol 10 (2) ◽

pp. 40-58 ◽

Cited By ~ 2

Author(s):

Sanur Sharma ◽

Anurag Jain

Keyword(s):

Social Media ◽

Feature Selection ◽

Ensemble Learning ◽

Information Gain ◽

Empirical Evaluation ◽

Ensemble Classifier ◽

Sentiment Classification ◽

Ensemble Classifiers ◽

Chi Squared

This article presents a study on ensemble learning and an empirical evaluation of various ensemble classifiers and ensemble features for sentiment classification of social media data. The data was collected from Twitter in real-time using Twitter API and text pre-processing and ranking-based feature selection is applied to textual data. A framework for a hybrid ensemble learning model is presented where a combination of ensemble features (Information Gain and CHI-Squared) and ensemble classifier that includes Ada Boost with SMO-SVM and Logistic Regression has been implemented. The classification of Twitter data is performed where sentiment analysis is used as a feature. The proposed model has shown improvements as compared to the state-of-the-art methods with an accuracy of 88.2% with a low error rate.

Download Full-text