COMPARATIVE ANALYSIS OF CLUSTER CONCENTRIC CIRCLE BASED UNDER SAMPLING OVER LOW VERSUS HIGH DIMENSIONAL IMBALANCED DATASETS

Detection of outliers or anomalies is one of the vital issues in pattern-driven data mining. Outlier detection detects the inconsistent behavior of individual objects. It is an important sector in the data mining field with several different applications such as detecting credit card fraud, hacking discovery and discovering criminal activities. It is necessary to develop tools used to uncover the critical information established in the extensive data. This paper investigated a novel method for detecting cluster outliers in a multidimensional dataset, capable of identifying the clusters and outliers for datasets containing noise. The proposed method can detect the groups and outliers left by the clustering process, like instant irregular sets of clusters (C) and outliers (O), to boost the results. The results obtained after applying the algorithm to the dataset improved in terms of several parameters. For the comparative analysis, the accurate average value and the recall value parameters are computed. The accurate average value is 74.05% of the existing COID algorithm, and our proposed algorithm has 77.21%. The average recall value is 81.19% and 89.51% of the existing and proposed algorithm, which shows that the proposed work efficiency is better than the existing COID algorithm.

Download Full-text

A Comparative Analysis of Convergence Rate for Imbalanced Datasets of Active Learning Models

2018 IEEE 23rd International Conference on Digital Signal Processing (DSP) ◽

10.1109/icdsp.2018.8631877 ◽

2018 ◽

Author(s):

Haoke Zhang ◽

Wanqing Wu ◽

Sandeep Pirbhulal Guanglin Li ◽

Hongyi Zhang

Keyword(s):

Comparative Analysis ◽

Active Learning ◽

Convergence Rate ◽

Learning Models ◽

Imbalanced Datasets

Download Full-text

Comparative Analysis of NES and TMD Performance via High-Dimensional Invariant Manifolds

IUTAM Symposium on Exploiting Nonlinear Dynamics for Engineering Systems - IUTAM Bookseries ◽

10.1007/978-3-030-23692-2_13 ◽

2019 ◽

pp. 143-153 ◽

Cited By ~ 1

Author(s):

Giuseppe Habib ◽

Francesco Romeo

Keyword(s):

Comparative Analysis ◽

Invariant Manifolds ◽

High Dimensional

Download Full-text

Navo Minority Over-sampling Technique (NMOTe): A Consistent Performance Booster on Imbalanced Datasets

Journal of Electronics and Informatics - September 2019 ◽

10.36548/jei.2020.2.004 ◽

2020 ◽

Vol 2 (2) ◽

pp. 96-136

Author(s):

Navoneel Chakrabarty ◽

Sanket Biswas

Keyword(s):

State Of The Art ◽

High Dimensional Data ◽

Optimal Solution ◽

Sampling Technique ◽

High Dimensional ◽

Real World Data ◽

Imbalanced Datasets ◽

Comprehensive Overview ◽

Unequal Distribution ◽

Data Imbalance

Imbalanced data refers to a problem in machine learning where there exists unequal distribution of instances for each classes. Performing a classification task on such data can often turn bias in favour of the majority class. The bias gets multiplied in cases of high dimensional data. To settle this problem, there exists many real-world data mining techniques like over-sampling and under-sampling, which can reduce the Data Imbalance. Synthetic Minority Oversampling Technique (SMOTe) provided one such state-of-the-art and popular solution to tackle class imbalancing, even on high-dimensional data platform. In this work, a novel and consistent oversampling algorithm has been proposed that can further enhance the performance of classification, especially on binary imbalanced datasets. It has been named as NMOTe (Navo Minority Oversampling Technique), an upgraded and superior alternative to the existing techniques. A critical analysis and comprehensive overview on the literature has been done to get a deeper insight into the problem statements and nurturing the need to obtain the most optimal solution. The performance of NMOTe on some standard datasets has been established in this work to get a statistical understanding on why it has edged the existing state-of-the-art to become the most robust technique for solving the two-class data imbalance problem.

Download Full-text

Membrane Protein Type Prediction for High-Dimensional Imbalanced Datasets

2018 9th International Conference on Information Technology in Medicine and Education (ITME) ◽

10.1109/itme.2018.00190 ◽

2018 ◽

Author(s):

Lei Guo ◽

Shunfang Wang

Keyword(s):

Membrane Protein ◽

High Dimensional ◽

Imbalanced Datasets ◽

Membrane Protein Type

Download Full-text

Boosted Near-miss Under-sampling on SVM ensembles for concept detection in large-scale imbalanced datasets

Neurocomputing ◽

10.1016/j.neucom.2014.05.096 ◽

2016 ◽

Vol 172 ◽

pp. 198-206 ◽

Cited By ~ 13

Author(s):

Lei Bao ◽

Cao Juan ◽

Jintao Li ◽

Yongdong Zhang

Keyword(s):

Large Scale ◽

Near Miss ◽

Concept Detection ◽

Imbalanced Datasets ◽

Under Sampling

Download Full-text

Glyph sorting: Interactive visualization for multi-dimensional data

Information Visualization ◽

10.1177/1473871613511959 ◽

2013 ◽

Vol 14 (1) ◽

pp. 76-90 ◽

Cited By ~ 26

Author(s):

David HS Chung ◽

Philip A Legg ◽

Matthew L Parry ◽

Rhodri Bown ◽

Iwan W Griffiths ◽

...

Keyword(s):

Comparative Analysis ◽

Visual Search ◽

Conceptual Framework ◽

High Dimensional ◽

Event Analysis ◽

Match Analysis ◽

Technical Aspects ◽

Visualization Process ◽

Intuitive Manner

Glyph-based visualization is an effective tool for depicting multivariate information. Since sorting is one of the most common analytical tasks performed on individual attributes of a multi-dimensional dataset, this motivates the hypothesis that introducing glyph sorting would significantly enhance the usability of glyph-based visualization. In this article, we present a glyph-based conceptual framework as part of a visualization process for interactive sorting of multivariate data. We examine several technical aspects of glyph sorting and provide design principles for developing effective, visually sortable glyphs. Glyphs that are visually sortable provide two key benefits: (1) performing comparative analysis of multiple attributes between glyphs and (2) to support multi-dimensional visual search. We describe a system that incorporates focus and context glyphs to control sorting in a visually intuitive manner and for viewing sorted results in an interactive, multi-dimensional glyph plot that enables users to perform high-dimensional sorting, analyse and examine data trends in detail. To demonstrate the usability of glyph sorting, we present a case study in rugby event analysis for comparing and analysing trends within matches. This work is undertaken in conjunction with a national rugby team. From using glyph sorting, analysts have reported the discovery of new insight beyond traditional match analysis.

Download Full-text