Feature Extraction with Machine Learning and Data Mining Algorithms

The development of the phishing sites is by all accounts amazing. Despite the fact that the web clients know about these sorts of phishing assaults, part of clients move toward becoming casualty to these assaults. Quantities of assaults are propelled with the point of making web clients trust that they are speaking with a trusted entity. Phishing is one among them. Phishing is consistently developing since it is anything but difficult to duplicate a whole site utilizing the HTML source code. By rolling out slight improvements in the source code, it is conceivable to guide the victim to the phishing site. Phishers utilize part of strategies to draw the unsuspected web client. Consequently an efficient mechanism is required to recognize the phishing sites from the real sites keeping in mind the end goal to spare credential data. To detect the phishing websites and to identify it as information leaking sites, the system proposes data mining algorithms. In this paper, machine-learning algorithms have been utilized for modeling the prediction task. The process of identity extraction and feature extraction are discussed in this paper and the various experiments carried out to discover the performance of the models are demonstrated.

Download Full-text

Dr. Phish: Phishing Website Detector

E3S Web of Conferences ◽

10.1051/e3sconf/202129701032 ◽

2021 ◽

Vol 297 ◽

pp. 01032

Author(s):

Harish Kumar ◽

Anshal Prasad ◽

Ninad Rane ◽

Nilay Tamane ◽

Anjali Yeole

Keyword(s):

Machine Learning ◽

Data Mining ◽

Machine Learning Algorithms ◽

Machine Learning Techniques ◽

Cyber Crime ◽

Data Mining Algorithms ◽

Learning Techniques ◽

Mining Algorithms ◽

Host Properties ◽

New Strategies

Phishing is a common attack on credulous people by making them disclose their unique information. It is a type of cyber-crime where false sites allure exploited people to give delicate data. This paper deals with methods for detecting phishing websites by analyzing various features of URLs by Machine learning techniques. This experimentation discusses the methods used for detection of phishing websites based on lexical features, host properties and page importance properties. We consider various data mining algorithms for evaluation of the features in order to get a better understanding of the structure of URLs that spread phishing. To protect end users from visiting these sites, we can try to identify the phishing URLs by analyzing their lexical and host-based features.A particular challenge in this domain is that criminals are constantly making new strategies to counter our defense measures. To succeed in this contest, we need Machine Learning algorithms that continually adapt to new examples and features of phishing URLs.

Download Full-text

Benchmarking Data Mining Algorithms

Data Warehousing and Web Engineering ◽

10.4018/978-1-931777-02-5.ch003 ◽

2011 ◽

pp. 77-99

Author(s):

Balaji Rajagopalan ◽

Ravi Krovi

Keyword(s):

Machine Learning ◽

Data Mining ◽

Learning Algorithms ◽

Machine Learning Algorithms ◽

Successful Implementation ◽

Basic Premise ◽

Data Mining Algorithms ◽

External Data ◽

Mining Algorithms ◽

Careful Assessment

Data mining is the process of sifting through the mass of organizational (internal and external) data to identify patterns critical for decision support. Successful implementation of the data mining effort requires a careful assessment of the various tools and algorithms available. The basic premise of this study is that machine-learning algorithms, which are assumption free, should outperform their traditional counterparts when mining business databases. The objective of this study is to test this proposition by investigating the performance of the algorithms for several scenarios. The scenarios are based on simulations designed to reflect the extent to which typical statistical assumptions are violated in the business domain. The results of the computational experiments support the proposition that machine learning algorithms generally outperform their statistical counterparts under certain conditions. These can be used as prescriptive guidelines for the applicability of data mining techniques.

Download Full-text

Bio inspired Ensemble Feature Selection (BEFS) Model with Machine Learning and Data Mining Algorithms for Disease Risk Prediction

2019 5th International Conference On Computing, Communication, Control And Automation (ICCUBEA) ◽

10.1109/iccubea47591.2019.9129304 ◽

2019 ◽

Cited By ~ 1

Author(s):

Syed Javeed Pasha ◽

E. Syed Mohamed

Keyword(s):

Machine Learning ◽

Data Mining ◽

Feature Selection ◽

Risk Prediction ◽

Disease Risk ◽

Data Mining Algorithms ◽

Mining Algorithms

Download Full-text

Ensemble Gain Ratio Feature Selection (EGFS) Model with Machine Learning and Data Mining Algorithms for Disease Risk Prediction

2020 International Conference on Inventive Computation Technologies (ICICT) ◽

10.1109/icict48043.2020.9112406 ◽

2020 ◽

Cited By ~ 1

Author(s):

Syed Javeed Pasha ◽

E.Syed Mohamed

Keyword(s):

Machine Learning ◽

Data Mining ◽

Feature Selection ◽

Risk Prediction ◽

Disease Risk ◽

Gain Ratio ◽

Data Mining Algorithms ◽

Mining Algorithms

Download Full-text

Data Mining Using $\mathcal{MLC}++$ a Machine Learning Library in C++

International Journal of Artificial Intelligence Tools ◽

10.1142/s021821309700027x ◽

1997 ◽

Vol 06 (04) ◽

pp. 537-566 ◽

Cited By ~ 69

Author(s):

Ron Kohavi ◽

Dan Sommerfield ◽

James Dougherty

Keyword(s):

Machine Learning ◽

Data Mining ◽

Pattern Recognition ◽

Statistical Analysis ◽

Classification Algorithms ◽

Pattern Recognition Techniques ◽

Data Mining Algorithms ◽

Multiple Classification ◽

Mining Algorithms ◽

New Algorithms

Data mining algorithms including maching learning, statistical analysis, and pattern recognition techniques can greatly improve our understanding of data warehouses that are now becoming more widespread. In this paper, we focus on classification algorithms and review the need for multiple classification algorithms. We describe a system called [Formula: see text], which was designed to help choose the appropriate classification algorithm for a given dataset by making it easy to compare the utility of different algorithms on a specific dataset of interest. [Formula: see text] not only provides a workbench for such comparisons, but also provides a library of C++ classes to aid in the development of new algorithms, especially hybrid algorithms and multi-strategy algorithms. Such algorithms are generally hard to code from scratch. We discuss design issues, interfaces to other programs, and visualization of the resulting classifiers.

Download Full-text

An Improved Sentiment Analysis Approach to Detect Radical Content on Twitter

International Journal of Information Technology and Web Engineering ◽

10.4018/ijitwe.2021100103 ◽

2021 ◽

Vol 16 (4) ◽

pp. 52-73

Author(s):

kamel Ahsene Djaballah ◽

Kamel Boukhalfa ◽

Omar Boussaid ◽

Yassine Ramdane

Keyword(s):

Machine Learning ◽

Data Mining ◽

Social Networks ◽

Fuzzy Logic ◽

Sentiment Analysis ◽

Experimental Results ◽

Data Mining Algorithms ◽

Terrorist Groups ◽

Radical Content ◽

Mining Algorithms

Social networks are used by terrorist groups and people who support them to propagate their ideas, ideologies, or doctrines and share their views on terrorism. To analyze tweets related to terrorism, several studies have been proposed in the literature. Some works rely on data mining algorithms; others use lexicon-based or machine learning sentiment analysis. Some recent works adopt other methods that combine multi-techniques. This paper proposes an improved approach for sentiment analysis of radical content related to terrorist activity on Twitter. Unlike other solutions, the proposed approach focuses on using a dictionary of weighted terms, the Word2vec method, and trigrams, with a classification based on fuzzy logic. The authors have conducted experiments with 600 manually annotated tweets and 200,000 automatically collected tweets in English and Arabic to evaluate this approach. The experimental results revealed that the new technique provides between 75% to 78% of precision for radicality detection and 61% to 64% to detect radicality degrees.

Download Full-text

Novel Feature Reduction (NFR) Model With Machine Learning and Data Mining Algorithms for Effective Disease Risk Prediction

IEEE Access ◽

10.1109/access.2020.3028714 ◽

2020 ◽

Vol 8 ◽

pp. 184087-184108

Author(s):

Syed Javeed Pasha ◽

E. Syed Mohamed

Keyword(s):

Machine Learning ◽

Data Mining ◽

Risk Prediction ◽

Disease Risk ◽

Feature Reduction ◽

Data Mining Algorithms ◽

Mining Algorithms

Download Full-text

Benchmarking data mining approaches for traveler segmentation

International Journal of Electrical and Computer Engineering (IJECE) ◽

10.11591/ijece.v11i1.pp409-415 ◽

2021 ◽

Vol 11 (1) ◽

pp. 409

Author(s):

Tamer Uçar ◽

Adem Karahoca

Keyword(s):

Machine Learning ◽

Data Mining ◽

Machine Learning Algorithms ◽

Travel Agency ◽

Data Set ◽

Data Mining Algorithms ◽

Travel Agencies ◽

User Data ◽

Hybrid Data ◽

Mining Algorithms

The purpose of this study is proposing a hybrid data mining solution for traveler segmentation in tourism domain which can be used for planning user-oriented trips, arranging travel campaigns or similar services. Data set used in this work have been provided by a travel agency which contains flight and hotel bookings of travelers. Initially, the data set was prepared for running data mining algorithms. Then, various machine learning algorithms were benchmarked for performing accurate traveler segmentation and prediction tasks. Fuzzy C-means and X-means algorithms were applied for clustering user data. J48 and multilayer perceptron (MLP) algorithms were applied for classifying instances based on segmented user data. According to the findings of this study, J48 has the most effective classification results when applied on the data set which is clustered with X-means algorithm. The proposed hybrid data mining solution can be used by travel agencies to plan trip campaigns for similar travelers.

Download Full-text

Student Learning Prediction Using Machine Learning Techniques

International Journal of Engineering and Advanced Technology - Regular Issue ◽

10.35940/ijeat.f8710.088619 ◽

2019 ◽

Vol 8 (6) ◽

pp. 4179-4183

Keyword(s):

Machine Learning ◽

Data Mining ◽

Student Performance ◽

Machine Learning Techniques ◽

Data Mining Algorithms ◽

Learning Techniques ◽

Final Exams ◽

E Learning ◽

Using Data ◽

Mining Algorithms

Now a day’s e-learning is smartly growing technology. This technology is more helpful for students to communicate with their professors through chats or emails. ELearning also removes the obstacle of physical presence of an Elearner. The main aim of this paper is to predict student performance in their final exams using different machine learning techniques. Information like attendance, marks, assignments, class participation, seminar, CA, projects and semester are collected to predict student performance. This prediction helps the instructors to analyze their students based on their performance. For that we have used WEKA tool for the prediction of the student performance. WEKA (Waikato Environment for Knowledge Analysis) is one of the data mining too which is used for the classification and clustering using data mining algorithms. This prediction helps the students and the staffs to know how much effort their students need to be put in their final exams to get good marks.

Download Full-text