The Classification of Documents in Malay and Indonesian Using the Naive Bayesian Method Uses Words and Phrases as a Training Set

Malay Language and Indonesian Language are two closely related languages, sharing a lot in common in the meanings of words and grammar. Classifying the two languages automatically using a tool is a challenge because the two languages are very similar. The classification method that is widely used today is the Naive Bayesian method. This method needs to be implemented in a particular way to increase the level of classification accuracy. In this study, a new method was used, by using a training set in the form of words and phrases instead of just using a training set in the form of words only. With this method, the level of classification accuracy of the two languages is increased.

Download Full-text

Comparison of C4.5 algorithm with naive Bayesian method in classification of Diabetes Mellitus (A case study at Hasanuddin University hospital Makassar)

Journal of Physics Conference Series ◽

10.1088/1742-6596/1341/9/092009 ◽

2019 ◽

Vol 1341 ◽

pp. 092009

Author(s):

D R Ente ◽

S Arifin ◽

Andreza ◽

S A Thamrin

Keyword(s):

Diabetes Mellitus ◽

Bayesian Method ◽

University Hospital ◽

Naive Bayesian ◽

Naïve Bayesian ◽

C4.5 Algorithm

Download Full-text

The effect of training set on the classification of honey bee gut microbiota using the Naïve Bayesian Classifier

BMC Microbiology ◽

10.1186/1471-2180-12-221 ◽

2012 ◽

Vol 12 (1) ◽

pp. 221 ◽

Cited By ~ 35

Author(s):

Irene LG Newton ◽

Guus Roeselers

Keyword(s):

Gut Microbiota ◽

Honey Bee ◽

Bayesian Classifier ◽

Training Set ◽

Naïve Bayesian Classifier ◽

Naive Bayesian ◽

Naive Bayesian Classifier ◽

Naïve Bayesian

Download Full-text

Classification of SSVEP-based BCIs using Genetic Algorithm

Journal Of Big Data ◽

10.1186/s40537-021-00478-y ◽

2021 ◽

Vol 8 (1) ◽

Author(s):

Hamideh Soltani ◽

Zahra Einalou ◽

Mehrdad Dadgostar ◽

Keivan Maghooli

Keyword(s):

Genetic Algorithm ◽

Dimension Reduction ◽

Classification Accuracy ◽

Bayesian Method ◽

Computer Interface ◽

Support Vector ◽

Svm Classifier ◽

Effective Dimension ◽

Effective Dimension Reduction

AbstractBrain computer interface (BCI) systems have been regarded as a new way of communication for humans. In this research, common methods such as wavelet transform are applied in order to extract features. However, genetic algorithm (GA), as an evolutionary method, is used to select features. Finally, classification was done using the two approaches support vector machine (SVM) and Bayesian method. Five features were selected and the accuracy of Bayesian classification was measured to be 80% with dimension reduction. Ultimately, the classification accuracy reached 90.4% using SVM classifier. The results of the study indicate a better feature selection and the effective dimension reduction of these features, as well as a higher percentage of classification accuracy in comparison with other studies.

Download Full-text

Patent Text Classification Based on Naive Bayesian Method

2018 11th International Symposium on Computational Intelligence and Design (ISCID) ◽

10.1109/iscid.2018.00020 ◽

2018 ◽

Cited By ~ 1

Author(s):

Lizhong Xiao ◽

Guangzhong Wang ◽

Yuan Liu

Keyword(s):

Text Classification ◽

Bayesian Method ◽

Naive Bayesian ◽

Naïve Bayesian

Download Full-text

1D conditional generative adversarial network for spectrum-to-spectrum translation of simulated chemical reflectance signatures

Journal of Spectral Imaging ◽

10.1255/jsi.2021.a2 ◽

2021 ◽

Author(s):

Cara Murphy ◽

John Kerekes

Keyword(s):

Classification Accuracy ◽

Domain Adaptation ◽

Real Data ◽

Training Set ◽

Generative Adversarial Network ◽

Average Classification Accuracy ◽

Adversarial Network ◽

Chemical Residues ◽

Reflectance Data

The classification of trace chemical residues through active spectroscopic sensing is challenging due to the lack of physics-based models that can accurately predict spectra. To overcome this challenge, we leveraged the field of domain adaptation to translate data from the simulated to the measured domain for training a classifier. We developed the first 1D conditional generative adversarial network (GAN) to perform spectrum-to-spectrum translation of reflectance signatures. We applied the 1D conditional GAN to a library of simulated spectra and quantified the improvement in classification accuracy on real data using the translated spectra for training the classifier. Using the GAN-translated library, the average classification accuracy increased from 0.622 to 0.723 on real chemical reflectance data, including data from chemicals not included in the GAN training set.

Download Full-text

AN INFORMATION-THEORETIC FILTER METHOD FOR FEATURE WEIGHTING IN NAIVE BAYES

International Journal of Pattern Recognition and Artificial Intelligence ◽

10.1142/s0218001414510070 ◽

2014 ◽

Vol 28 (05) ◽

pp. 1451007 ◽

Cited By ~ 2

Author(s):

CHANG-HWAN LEE

Keyword(s):

Data Mining ◽

Bayesian Learning ◽

State Of The Art ◽

Feature Weighting ◽

New Method ◽

Filter Method ◽

Information Theoretic ◽

Naive Bayesian ◽

Naïve Bayesian ◽

Unrealistic Assumption

In spite of its simplicity, naive Bayesian learning has been widely used in many data mining applications. However, the unrealistic assumption that all features are equally important negatively impacts the performance of naive Bayesian learning. In this paper, we propose a new method that uses a Kullback–Leibler measure to calculate the weights of the features analyzed in naive Bayesian learning. Its performance is compared to that of other state-of-the-art methods over a number of datasets.

Download Full-text

A Semi-supervised Naive Bayesian Method for Labeling Heterogeneous Fingerprints

2019 International Conference on Control, Automation and Information Sciences (ICCAIS) ◽

10.1109/iccais46528.2019.9074649 ◽

2019 ◽

Author(s):

Yuan Shao ◽

Lei Wang ◽

Xiansheng Guo

Keyword(s):

Bayesian Method ◽

Naive Bayesian ◽

Naïve Bayesian

Download Full-text

Classification of Imbalanced Malaria Disease Using Naïve Bayesian Algorithm

International Journal of Engineering & Technology ◽

10.14419/ijet.v7i2.7.10978 ◽

2018 ◽

Vol 7 (2.7) ◽

pp. 786 ◽

Cited By ~ 1

Author(s):

T Sajana ◽

M R.Narasingarao

Keyword(s):

Comparative Study ◽

Class Imbalance ◽

Machine Learning Algorithms ◽

Bayesian Algorithm ◽

Naive Bayesian ◽

Class Distribution ◽

Naïve Bayesian ◽

R Programming ◽

The Impact

Malaria disease is one whose presence is rampant in semi urban and non-urban areas especially resource poor developing countries. It is quite evident from the datasets like malaria, dengue, etc., where there is always a possibility of having more negative patients (non-occurrence of the disease) compared to patients suffering from disease (positive cases). Developing a model based decision support system with such unbalanced datasets is a cause of concern and it is indeed necessary to have a model predicting the disease quite accurately. Classification of imbalanced malaria disease data become a crucial task in medical application domain because most of the conventional machine learning algorithms are showing very poor performance to classify whether a patient is affected by malaria disease or not. In imbalanced data, majority (unaffected) class samples are dominates the minority (affected) class samples leading to class imbalance. To overcome the nature of class imbalance problem, balancing the data samples is the best solution which produces the better accuracy in classification of minority samples. The aim of this research is to propose a comparative study on classifying the imbalanced malaria disease data using Naive Bayesian classifier in different environments like weka and using an R-language. We present here, clinical descriptive study on 165 patients of different age group people collected at medical wards of Narasaraopet from 2014-17. Synthetic Minority Oversampling Technique (SMOTE) technique has been used to balance the class distribution and then we performed a comparative study on the dataset using Naïve Bayesian algorithm in various platforms. Out of balanced class distribution data, 70% data was given to train the Naive Bayesian algorithm and the rest of the data was used for testing the model for both weka and R programming environments. Experimental results have indicated that, classification of malaria disease data in weka environment has highest accuracy of 88.5% than the Naive Bayesian algorithm accuracy of 87.5% using R programming language. The impact of vector borne disease is very high in medical applications. Prediction of disease like malaria is an hour of the need and this is possible only with a suitable model for a given dataset. Hence, we have developed a model with Naive Bayesian algorithm is used for current research.

Download Full-text