Feature Selection Optimization for Highlighting Opinions Using Supervised and Unsupervised Learning on Arabic Language

Text mining utilizes machine learning (ML) and natural language processing (NLP) for text implicit knowledge recognition, such knowledge serves many domains as translation, media searching, and business decision making. Opinion mining (OM) is one of the promised text mining fields, which are used for polarity discovering via text and has terminus benefits for business. ML techniques are divided into two approaches: supervised and unsupervised learning, since we herein testified an OM feature selection(FS)using four ML techniques. In this paper, we had implemented number of experiments via four machine learning techniques on the same three Arabic language corpora. This paper aims at increasing the accuracy of opinion highlighting on Arabic language, by using enhanced feature selection approaches. FS proposed model is adopted for enhancing opinion highlighting purpose. The experimental results show the outperformance of the proposed approaches in variant levels of supervisory,i.e. different techniques via distinct data domains. Multiple levels of comparison are carried out and discussed for further understanding of the impact of proposed model on several ML techniques.

Download Full-text

Inflated prediction accuracy of neuropsychiatric biomarkers caused by data leakage in feature selection

Scientific Reports ◽

10.1038/s41598-021-87157-3 ◽

2021 ◽

Vol 11 (1) ◽

Author(s):

Miseon Shim ◽

Seung-Hwan Lee ◽

Han-Jeong Hwang

Keyword(s):

Machine Learning ◽

Feature Selection ◽

Feature Selection Method ◽

Selection Method ◽

Prediction Performance ◽

Machine Learning Techniques ◽

Learning Techniques ◽

Eeg Data ◽

Neuropsychiatric Diseases ◽

The Impact

AbstractIn recent years, machine learning techniques have been frequently applied to uncovering neuropsychiatric biomarkers with the aim of accurately diagnosing neuropsychiatric diseases and predicting treatment prognosis. However, many studies did not perform cross validation (CV) when using machine learning techniques, or others performed CV in an incorrect manner, leading to significantly biased results due to overfitting problem. The aim of this study is to investigate the impact of CV on the prediction performance of neuropsychiatric biomarkers, in particular, for feature selection performed with high-dimensional features. To this end, we evaluated prediction performances using both simulation data and actual electroencephalography (EEG) data. The overall prediction accuracies of the feature selection method performed outside of CV were considerably higher than those of the feature selection method performed within CV for both the simulation and actual EEG data. The differences between the prediction accuracies of the two feature selection approaches can be thought of as the amount of overfitting due to selection bias. Our results indicate the importance of correctly using CV to avoid biased results of prediction performance of neuropsychiatric biomarkers.

Download Full-text

Machine Learning Fake News Classification with Optimal Feature Selection

10.21203/rs.3.rs-835344/v1 ◽

2021 ◽

Author(s):

Muhammad Fayaz ◽

Atif Khan ◽

Muhammad Bilal ◽

Sanaullah Khan

Keyword(s):

Machine Learning ◽

Feature Selection ◽

Information Gain ◽

Machine Learning Techniques ◽

Fake News ◽

Learning Techniques ◽

Proposed Model ◽

Current Events ◽

The Government ◽

Experimental Findings

Abstract Nowadays, information is published in newspapers and social media while transmitted on radio and television about current events and specific fields of interest nationwide and abroad. It becomes difficult to explicit what is real and what is fake due to the explosive growth of online content. As a result, fake news has become epidemic and immensely challenging to analyze fake news to be verified by the producers in the form of data process outlets not to mislead the people. Indeed, it is a big challenge to the government and public to debate the situation depending on case to case. For the purpose several websites were developed for this purpose to classify the news as either real or fake depending on the website logic and algorithm. A mechanism has to be taken on fact-checking rumors and statements, particularly those that get thousands of views and likes before being debunked and refuted by expert sources. Various machine learning techniques have been used to detect and correctly classified of fake news. However, these approaches are restricted in terms of accuracy. This study has applied a Random Forest (RF) classifier to predict fake or real news. For this prpose, twenty-three (23) textual features are extracted from ISOT Fake News Dataset. Four best feature selection techniques like Chi2, Univariate, information gain and Feature importance are used for selecting fourteen best features out of twenty-three. The proposed model and other benchmark techniques are evaluated on the dataset by using best features. Experimental findings show that, the proposed model outperformed state-of-the-art machine learning techniques such as GBM, XGBoost and Ada Boost Regression Model in terms of classification accuracy.

Download Full-text

Rank Prediction for Articles and Conference Papersusing Machine Learning Techniques

International Journal of Recent Technology and Engineering - 2 ◽

10.35940/ijrte.d9257.118419 ◽

2019 ◽

Vol 8 (4) ◽

pp. 9746-9750

Keyword(s):

Artificial Intelligence ◽

Machine Learning ◽

Neural Networks ◽

Unsupervised Learning ◽

Supervised Learning ◽

Research Work ◽

Machine Learning Techniques ◽

Supervised And Unsupervised Learning ◽

Learning Techniques ◽

Day By Day

Searching for an optimal article which was given highest and best priority is quite harder based on requirements. Ranking is one of the best measure or a method to get the best rated and optimal article or a conference or a research paper through this huge Internet World. As Technology been increasing day by day Artificial Intelligence is the first step to get through any problem for a solution Machine learning is also an important aspect of Artificial Intelligence. Machine Learning is best known for classifying, categorizing and predicting. Rank prediction can be done through many different algorithm implementations in machine learning. But choosing the best is important for accurate results. This paper gives the most accurate results of algorithms that can be used for rank predictions for articles. To simplify and resolve this problem, solutions were given in many different ways but to achieve accuracy is necessary, in previous models this is given using supervised learning only. We proposed this research work with perfect results using both supervised and unsupervised learning. Neural Networks is the best algorithm in supervised learning for classifying and predicting within data. In unsupervised learning we used K-means clustering because of grouping the data. This work helps the user(s) for optimal search of an article and also gives a competitive spirit for author to get into the top, totally this is implemented using Machine Learning Techniques of Neural Networks, K-Means Algorithm which is a mixture of supervised and unsupervised learning for predicting ranks.

Download Full-text

On the Impact of Sample Duplication in Machine-Learning-Based Android Malware Detection

ACM Transactions on Software Engineering and Methodology ◽

10.1145/3446905 ◽

2021 ◽

Vol 30 (3) ◽

pp. 1-38

Author(s):

Yanjie Zhao ◽

Li Li ◽

Haoyu Wang ◽

Haipeng Cai ◽

Tegawendé F. Bissyandé ◽

...

Keyword(s):

Machine Learning ◽

Unsupervised Learning ◽

Malware Detection ◽

Experimental Results ◽

Machine Learning Techniques ◽

Detection Rates ◽

Android Malware ◽

Android Malware Detection ◽

Learning Techniques ◽

The Impact

Malware detection at scale in the Android realm is often carried out using machine learning techniques. State-of-the-art approaches such as DREBIN and MaMaDroid are reported to yield high detection rates when assessed against well-known datasets. Unfortunately, such datasets may include a large portion of duplicated samples, which may bias recorded experimental results and insights. In this article, we perform extensive experiments to measure the performance gap that occurs when datasets are de-duplicated. Our experimental results reveal that duplication in published datasets has a limited impact on supervised malware classification models. This observation contrasts with the finding of Allamanis on the general case of machine learning bias for big code. Our experiments, however, show that sample duplication more substantially affects unsupervised learning models (e.g., malware family clustering). Nevertheless, we argue that our fellow researchers and practitioners should always take sample duplication into consideration when performing machine-learning-based (via either supervised or unsupervised learning) Android malware detections, no matter how significant the impact might be.

Download Full-text

Classification of multiwavelength transients with Machine Learning

Monthly Notices of the Royal Astronomical Society ◽

10.1093/mnras/staa3873 ◽

2020 ◽

Author(s):

K Sooknunan ◽

M Lochner ◽

Bruce A Bassett ◽

H V Peiris ◽

R Fender ◽

...

Keyword(s):

Machine Learning ◽

Small Sample ◽

Light Curves ◽

Machine Learning Techniques ◽

Optical Data ◽

Test Time ◽

Test Accuracy ◽

Training Set ◽

The Impact

Abstract With the advent of powerful telescopes such as the Square Kilometer Array and the Vera C. Rubin Observatory, we are entering an era of multiwavelength transient astronomy that will lead to a dramatic increase in data volume. Machine learning techniques are well suited to address this data challenge and rapidly classify newly detected transients. We present a multiwavelength classification algorithm consisting of three steps: (1) interpolation and augmentation of the data using Gaussian processes; (2) feature extraction using wavelets; (3) classification with random forests. Augmentation provides improved performance at test time by balancing the classes and adding diversity into the training set. In the first application of machine learning to the classification of real radio transient data, we apply our technique to the Green Bank Interferometer and other radio light curves. We find we are able to accurately classify most of the eleven classes of radio variables and transients after just eight hours of observations, achieving an overall test accuracy of 78%. We fully investigate the impact of the small sample size of 82 publicly available light curves and use data augmentation techniques to mitigate the effect. We also show that on a significantly larger simulated representative training set that the algorithm achieves an overall accuracy of 97%, illustrating that the method is likely to provide excellent performance on future surveys. Finally, we demonstrate the effectiveness of simultaneous multiwavelength observations by showing how incorporating just one optical data point into the analysis improves the accuracy of the worst performing class by 19%.

Download Full-text

Radiomics side experiments and DAFIT approach in identifying pulmonary hypertension using Cardiac MRI derived radiomics based machine learning models

Scientific Reports ◽

10.1038/s41598-021-92155-6 ◽

2021 ◽

Vol 11 (1) ◽

Author(s):

Sarv Priya ◽

Tanya Aggarwal ◽

Caitlin Ward ◽

Girish Bathla ◽

Mathews Jacob ◽

...

Keyword(s):

Machine Learning ◽

Pulmonary Hypertension ◽

Feature Selection ◽

Subgroup Analysis ◽

Cardiac Mri ◽

Intraclass Correlation ◽

Poor Performance ◽

Superior Performance ◽

Feature Filtering ◽

The Impact

AbstractSide experiments are performed on radiomics models to improve their reproducibility. We measure the impact of myocardial masks, radiomic side experiments and data augmentation for information transfer (DAFIT) approach to differentiate patients with and without pulmonary hypertension (PH) using cardiac MRI (CMRI) derived radiomics. Feature extraction was performed from the left ventricle (LV) and right ventricle (RV) myocardial masks using CMRI in 82 patients (42 PH and 40 controls). Various side study experiments were evaluated: Original data without and with intraclass correlation (ICC) feature-filtering and DAFIT approach (without and with ICC feature-filtering). Multiple machine learning and feature selection strategies were evaluated. Primary analysis included all PH patients with subgroup analysis including PH patients with preserved LVEF (≥ 50%). For both primary and subgroup analysis, DAFIT approach without feature-filtering was the highest performer (AUC 0.957–0.958). ICC approaches showed poor performance compared to DAFIT approach. The performance of combined LV and RV masks was superior to individual masks alone. There was variation in top performing models across all approaches (AUC 0.862–0.958). DAFIT approach with features from combined LV and RV masks provide superior performance with poor performance of feature filtering approaches. Model performance varies based upon the feature selection and model combination.

Download Full-text

FSDroid:- A feature selection technique to detect malware from Android using Machine Learning Techniques

Multimedia Tools and Applications ◽

10.1007/s11042-020-10367-w ◽

2021 ◽

Author(s):

Arvind Mahindru ◽

A.L. Sangal

Keyword(s):

Machine Learning ◽

Feature Selection ◽

Machine Learning Techniques ◽

Feature Selection Technique ◽

Selection Technique ◽

Learning Techniques

Download Full-text

Leveraging Road Characteristics and Contributor Behaviour for Assessing Road Type Quality in OSM

ISPRS International Journal of Geo-Information ◽

10.3390/ijgi10070436 ◽

2021 ◽

Vol 10 (7) ◽

pp. 436

Author(s):

Amerah Alghanim ◽

Musfira Jilani ◽

Michela Bertolotto ◽

Gavin McArdle

Keyword(s):

Machine Learning ◽

Spatial Data ◽

Classification Accuracy ◽

Supervised Machine Learning ◽

Machine Learning Techniques ◽

Data Set ◽

Semantic Inference ◽

Road Type ◽

The Impact

Volunteered Geographic Information (VGI) is often collected by non-expert users. This raises concerns about the quality and veracity of such data. There has been much effort to understand and quantify the quality of VGI. Extrinsic measures which compare VGI to authoritative data sources such as National Mapping Agencies are common but the cost and slow update frequency of such data hinder the task. On the other hand, intrinsic measures which compare the data to heuristics or models built from the VGI data are becoming increasingly popular. Supervised machine learning techniques are particularly suitable for intrinsic measures of quality where they can infer and predict the properties of spatial data. In this article we are interested in assessing the quality of semantic information, such as the road type, associated with data in OpenStreetMap (OSM). We have developed a machine learning approach which utilises new intrinsic input features collected from the VGI dataset. Specifically, using our proposed novel approach we obtained an average classification accuracy of 84.12%. This result outperforms existing techniques on the same semantic inference task. The trustworthiness of the data used for developing and training machine learning models is important. To address this issue we have also developed a new measure for this using direct and indirect characteristics of OSM data such as its edit history along with an assessment of the users who contributed the data. An evaluation of the impact of data determined to be trustworthy within the machine learning model shows that the trusted data collected with the new approach improves the prediction accuracy of our machine learning technique. Specifically, our results demonstrate that the classification accuracy of our developed model is 87.75% when applied to a trusted dataset and 57.98% when applied to an untrusted dataset. Consequently, such results can be used to assess the quality of OSM and suggest improvements to the data set.

Download Full-text

Development of Machine Learning Models to Evaluate the Toughness of OPH Alloys

Materials ◽

10.3390/ma14216713 ◽

2021 ◽

Vol 14 (21) ◽

pp. 6713

Author(s):

Omid Khalaj ◽

Moslem Ghobadi ◽

Ehsan Saebnoori ◽

Alireza Zarezadeh ◽

Mohammadreza Shishesaz ◽

...

Keyword(s):

Machine Learning ◽

Mechanical Properties ◽

Mechanical Alloying ◽

Fuzzy Inference ◽

Oxide Dispersion Strengthened ◽

Machine Learning Techniques ◽

Support Vector ◽

Anfis Model ◽

Inference Systems ◽

The Impact

Oxide Precipitation-Hardened (OPH) alloys are a new generation of Oxide Dispersion-Strengthened (ODS) alloys recently developed by the authors. The mechanical properties of this group of alloys are significantly influenced by the chemical composition and appropriate heat treatment (HT). The main steps in producing OPH alloys consist of mechanical alloying (MA) and consolidation, followed by hot rolling. Toughness was obtained from standard tensile test results for different variants of OPH alloy to understand their mechanical properties. Three machine learning techniques were developed using experimental data to simulate different outcomes. The effectivity of the impact of each parameter on the toughness of OPH alloys is discussed. By using the experimental results performed by the authors, the composition of OPH alloys (Al, Mo, Fe, Cr, Ta, Y, and O), HT conditions, and mechanical alloying (MA) were used to train the models as inputs and toughness was set as the output. The results demonstrated that all three models are suitable for predicting the toughness of OPH alloys, and the models fulfilled all the desired requirements. However, several criteria validated the fact that the adaptive neuro-fuzzy inference systems (ANFIS) model results in better conditions and has a better ability to simulate. The mean square error (MSE) for artificial neural networks (ANN), ANFIS, and support vector regression (SVR) models was 459.22, 0.0418, and 651.68 respectively. After performing the sensitivity analysis (SA) an optimized ANFIS model was achieved with a MSE value of 0.003 and demonstrated that HT temperature is the most significant of these parameters, and this acts as a critical rule in training the data sets.

Download Full-text

A Soft Voting Ensemble-Based Model for the Early Prediction of Idiopathic Pulmonary Fibrosis (IPF) Disease Severity in Lungs Disease Patients

Life ◽

10.3390/life11101092 ◽

2021 ◽

Vol 11 (10) ◽

pp. 1092

Author(s):

Sikandar Ali ◽

Ali Hussain ◽

Satyabrata Aich ◽

Moo Suk Park ◽

Man Pyo Chung ◽

...

Keyword(s):

Machine Learning ◽

Idiopathic Pulmonary Fibrosis ◽

Pulmonary Fibrosis ◽

Lung Disease ◽

Lung Diseases ◽

Early Stage ◽

Confusion Matrix ◽

Machine Learning Techniques ◽

Training Dataset ◽

Proposed Model

Idiopathic pulmonary fibrosis, which is one of the lung diseases, is quite rare but fatal in nature. The disease is progressive, and detection of severity takes a long time as well as being quite tedious. With the advent of intelligent machine learning techniques, and also the effectiveness of these techniques, it was possible to detect many lung diseases. So, in this paper, we have proposed a model that could be able to detect the severity of IPF at the early stage so that fatal situations can be controlled. For the development of this model, we used the IPF dataset of the Korean interstitial lung disease cohort data. First, we preprocessed the data while applying different preprocessing techniques and selected 26 highly relevant features from a total of 502 features for 2424 subjects. Second, we split the data into 80% training and 20% testing sets and applied oversampling on the training dataset. Third, we trained three state-of-the-art machine learning models and combined the results to develop a new soft voting ensemble-based model for the prediction of severity of IPF disease in patients with this chronic lung disease. Hyperparameter tuning was also performed to get the optimal performance of the model. Fourth, the performance of the proposed model was evaluated by calculating the accuracy, AUC, confusion matrix, precision, recall, and F1-score. Lastly, our proposed soft voting ensemble-based model achieved the accuracy of 0.7100, precision 0.6400, recall 0.7100, and F1-scores 0.6600. This proposed model will help the doctors, IPF patients, and physicians to diagnose the severity of the IPF disease in its early stages and assist them to take proactive measures to overcome this disease by enabling the doctors to take necessary decisions pertaining to the treatment of IPF disease.

Download Full-text