Randomized kernel methods for least-squares support vector machines

The least-squares support vector machine (LS-SVM) is a frequently used kernel method for non-linear regression and classification tasks. Here we discuss several approximation algorithms for the LS-SVM classifier. The proposed methods are based on randomized block kernel matrices, and we show that they provide good accuracy and reliable scaling for multi-class classification problems with relatively large data sets. Also, we present several numerical experiments that illustrate the practical applicability of the proposed methods.

Download Full-text

Support Vector Machines on Large Data Sets: Simple Parallel Approaches

Studies in Classification, Data Analysis, and Knowledge Organization - Data Analysis, Machine Learning and Knowledge Discovery ◽

10.1007/978-3-319-01595-8_10 ◽

2013 ◽

pp. 87-95 ◽

Cited By ~ 5

Author(s):

Oliver Meyer ◽

Bernd Bischl ◽

Claus Weihs

Keyword(s):

Support Vector Machines ◽

Large Data ◽

Large Data Sets ◽

Support Vector ◽

Data Sets ◽

Vector Machines

Download Full-text

Multi-Class Support Vector Machines for Large Data Sets via Minimum Enclosing Ball Clustering

2007 4th International Conference on Electrical and Electronics Engineering ◽

10.1109/iceee.2007.4344994 ◽

2007 ◽

Cited By ~ 2

Author(s):

Jair Cervantes ◽

Xiaoou Li ◽

Wen Yu ◽

Javier Bejarano

Keyword(s):

Support Vector Machines ◽

Large Data ◽

Large Data Sets ◽

Support Vector ◽

Data Sets ◽

Vector Machines ◽

Minimum Enclosing Ball

Download Full-text

Using support vector machines for mining regression classes in large data sets

2002 IEEE Region 10 Conference on Computers, Communications, Control and Power Engineering. TENCOM '02. Proceedings. ◽

10.1109/tencon.2002.1181221 ◽

2004 ◽

Author(s):

Zonghai Sun ◽

Lixin Gao ◽

Youxian Sun

Keyword(s):

Support Vector Machines ◽

Large Data ◽

Large Data Sets ◽

Support Vector ◽

Data Sets ◽

Vector Machines

Download Full-text

Fast classification for large data sets via random selection clustering and Support Vector Machines

Intelligent Data Analysis ◽

10.3233/ida-2012-00558 ◽

2012 ◽

Vol 16 (6) ◽

pp. 897-914 ◽

Cited By ~ 5

Author(s):

Xiaoou Li ◽

Jair Cervantes ◽

Wen Yu

Keyword(s):

Support Vector Machines ◽

Large Data ◽

Random Selection ◽

Large Data Sets ◽

Support Vector ◽

Data Sets ◽

Vector Machines ◽

Fast Classification

Download Full-text

Using the Leader Algorithm with Support Vector Machines for Large Data Sets

Lecture Notes in Computer Science - Artificial Neural Networks and Machine Learning – ICANN 2011 ◽

10.1007/978-3-642-21735-7_28 ◽

2011 ◽

pp. 225-232 ◽

Cited By ~ 1

Author(s):

Enrique Romero

Keyword(s):

Support Vector Machines ◽

Large Data ◽

Large Data Sets ◽

Support Vector ◽

Data Sets ◽

Vector Machines

Download Full-text

Density-Dependent Quantized Least Squares Support Vector Machine for Large Data Sets

IEEE Transactions on Neural Networks and Learning Systems ◽

10.1109/tnnls.2015.2504382 ◽

2017 ◽

Vol 28 (1) ◽

pp. 94-106 ◽

Cited By ~ 24

Author(s):

Shengyu Nan ◽

Lei Sun ◽

Badong Chen ◽

Zhiping Lin ◽

Kar-Ann Toh

Keyword(s):

Support Vector Machine ◽

Least Squares ◽

Large Data ◽

Large Data Sets ◽

Support Vector ◽

Data Sets ◽

Density Dependent

Download Full-text

Distributed training and scalability for the particle clustering method UCluster

EPJ Web of Conferences ◽

10.1051/epjconf/202125102054 ◽

2021 ◽

Vol 251 ◽

pp. 02054

Author(s):

Olga Sunneborn Gudnadottir ◽

Daniel Gedon ◽

Colin Desmarais ◽

Karl Bengtsson Bernander ◽

Raazesh Sainudiin ◽

...

Keyword(s):

Particle Physics ◽

Hadron Collider ◽

Large Data ◽

Large Data Sets ◽

Data Sets ◽

Training Time ◽

Distributed Training ◽

Machine Learning Methods ◽

Multi Class Classification

In recent years, machine-learning methods have become increasingly important for the experiments at the Large Hadron Collider (LHC). They are utilised in everything from trigger systems to reconstruction and data analysis. The recent UCluster method is a general model providing unsupervised clustering of particle physics data, that can be easily modified to provide solutions for a variety of different decision problems. In the current paper, we improve on the UCluster method by adding the option of training the model in a scalable and distributed fashion, and thereby extending its utility to learn from arbitrarily large data sets. UCluster combines a graph-based neural network called ABCnet with a clustering step, using a combined loss function in the training phase. The original code is publicly available in TensorFlow v1.14 and has previously been trained on a single GPU. It shows a clustering accuracy of 81% when applied to the problem of multi-class classification of simulated jet events. Our implementation adds the distributed training functionality by utilising the Horovod distributed training framework, which necessitated a migration of the code to TensorFlow v2. Together with using parquet files for splitting data up between different compute nodes, the distributed training makes the model scalable to any amount of input data, something that will be essential for use with real LHC data sets. We find that the model is well suited for distributed training, with the training time decreasing in direct relation to the number of GPU’s used. However, further improvements by a more exhaustive and possibly distributed hyper-parameter search is required in order to achieve the reported accuracy of the original UCluster method.

Download Full-text

Support vector machine classification for large data sets via minimum enclosing ball clustering

Neurocomputing ◽

10.1016/j.neucom.2007.07.028 ◽

2008 ◽

Vol 71 (4-6) ◽

pp. 611-619 ◽

Cited By ~ 59

Author(s):

Jair Cervantes ◽

Xiaoou Li ◽

Wen Yu ◽

Kang Li

Keyword(s):

Support Vector Machine ◽

Large Data ◽

Large Data Sets ◽

Support Vector ◽

Data Sets ◽

Support Vector Machine Classification ◽

Minimum Enclosing Ball

Download Full-text

Correlation Kernels for Support Vector Machines Classification with Applications in Cancer Data

Computational and Mathematical Methods in Medicine ◽

10.1155/2012/205025 ◽

2012 ◽

Vol 2012 ◽

pp. 1-7 ◽

Cited By ~ 9

Author(s):

Hao Jiang ◽

Wai-Ki Ching

Keyword(s):

Support Vector Machines ◽

Positive Semidefinite ◽

Superior Performance ◽

Support Vector ◽

Svm Classifier ◽

Classification Problems ◽

Correlation Kernel ◽

Cancer Data ◽

Tumor Tissues ◽

Vector Machines

High dimensional bioinformatics data sets provide an excellent and challenging research problem in machine learning area. In particular, DNA microarrays generated gene expression data are of high dimension with significant level of noise. Supervised kernel learning with an SVM classifier was successfully applied in biomedical diagnosis such as discriminating different kinds of tumor tissues. Correlation Kernel has been recently applied to classification problems with Support Vector Machines (SVMs). In this paper, we develop a novel and parsimonious positive semidefinite kernel. The proposed kernel is shown experimentally to have better performance when compared to the usual correlation kernel. In addition, we propose a new kernel based on the correlation matrix incorporating techniques dealing with indefinite kernel. The resulting kernel is shown to be positive semidefinite and it exhibits superior performance to the two kernels mentioned above. We then apply the proposed method to some cancer data in discriminating different tumor tissues, providing information for diagnosis of diseases. Numerical experiments indicate that our method outperforms the existing methods such as the decision tree method and KNN method.

Download Full-text