Critical joint identification for efficient sequencing

Abstract Identifying the optimal sequence of joining is an exhaustive combinatorial optimization problem. On each assembly, there is a specific number of weld points that determine the geometrical deviation of the assembly after joining. The number and sequence of such weld points play a crucial role both for sequencing and assembly planning. While there are studies on identifying the complete sequence of welding, identifying such joints are not addressed. In this paper, based on the principles of machine intelligence, black-box models of the assembly sequences are built using the support vector machines (SVM). To identify the number of the critical weld points, principle component analysis is performed on a proposed data set, evaluated using the SVM models. The approach has been applied to three assemblies of different sizes, and has successfully identified the corresponding critical weld points. It has been shown that a small fraction of the weld points of the assembly can reduce more than 60% of the variability in the assembly deviation after joining.

Download Full-text

Explainable and Interpretable Anomaly Detection Models for Production Data

SPE Journal ◽

10.2118/208586-pa ◽

2021 ◽

pp. 1-15

Author(s):

Basma Alharbi ◽

Zhenwen Liang ◽

Jana M. Aljindan ◽

Ammar K. Agnia ◽

Xiangliang Zhang

Keyword(s):

Anomaly Detection ◽

Global Analysis ◽

Black Box ◽

Prediction Performance ◽

Support Vector ◽

Production Data ◽

K Nearest Neighbor ◽

Data Set ◽

Box Models ◽

Black Box Models

Summary Trusting a machine-learning model is a critical factor that will speed the spread of the fourth industrial revolution. Trust can be achieved by understanding how a model is making decisions. For white-box models, it is easy to “see” the model and examine its prediction. For black-box models, the explanation of the decision process is not straightforward. In this work, we compare the performance of several white- and black-box models on two production data sets in an anomaly detection task. The presence of anomalies in production data can significantly influence business decisions and misrepresent the results of the analysis, if not identified. Therefore, identifying anomalies is a crucial and necessary step to maintain safety and ensure that the wells perform at full capacity. To achieve this, we compare the performance of K-nearest neighbor (KNN), logistic regression (Logit), support vector machines (SVMs), decision tree (DT), random forest (RF), and rule fit classifier (RFC). F1 and complexity are the two main metrics used to compare the prediction performance and interpretability of these models. In one data set, RFC outperformed the remaining models in both F1 and complexity, where F1 = 0.92, and complexity = 0.5. In the second data set, RF outperformed the rest in prediction performance with F1 = 0.84, yet it had the lowest complexity metric (0.04). We further analyzed the best performing models by explaining their predictions using local interpretable model-agnostic explanations, which provide justification for decisions made for each instance. Additionally, we evaluated the global rules learned from white-box models. Local and global analysis enable decision makers to understand how and why models are making certain decisions, which in turn allows trusting the models.

Download Full-text

An approach of classification and parameters estimation, using neural network, for lubricant degradation diagnosis

MATEC Web of Conferences ◽

10.1051/matecconf/201818403009 ◽

2018 ◽

Vol 184 ◽

pp. 03009

Author(s):

Gavril Grebenişan ◽

Nazzal Salem ◽

Sanda Bogdan

Keyword(s):

Machine Tools ◽

Machine Intelligence ◽

Parameters Estimation ◽

Support Vector ◽

Lubricant Degradation ◽

National Program ◽

Data Set ◽

Validation And Verification ◽

Industrial Systems ◽

Degradation Diagnosis

This paper addresses a delicate problem, namely the diagnosis of the state of the oils in the industrial systems, namely the machine tools. Based on measurements (the data set contains over five million records), within a Machine Intelligence for Diagnosis Automation (MIDA) project funded by the National Program PN II, ERA MANUNET: NR 13081221 / 13.08.2013, several applications of MATLAB toolbars are being developed in the field of artificial intelligence, specifically using the Support Vector Machine algorithms and neural networks. The tests were carried out on several distinct situations, followed by validation and verification tests on the devices designed and developed within the project (MIDA, Monitoil).

Download Full-text

Structural Analysis of Product and Computer-Aided Assembly Planning in AssemBL Software Package

Mechanical Engineering and Computer Science ◽

10.24108/0818.0001424 ◽

2018 ◽

pp. 11-33

Author(s):

A. N. Bozhko

Keyword(s):

Computer Aided Design ◽

State Of The Art ◽

Graph Model ◽

Design System ◽

Assembly Planning ◽

Complex Products ◽

Assembly Sequences ◽

Computer Aided ◽

Cad Cam ◽

Aided Design

Computer-aided design of assembly processes (Computer aided assembly planning, CAAP) of complex products is an important and urgent problem of state-of-the-art information technologies. Intensive research on CAAP has been underway since the 1980s. Meanwhile, specialized design systems were created to provide synthesis of assembly plans and product decompositions into assembly units. Such systems as ASPE, RAPID, XAP / 1, FLAPS, Archimedes, PRELEIDES, HAP, etc. can be given, as an example. These experimental developments did not get widespread use in industry, since they are based on the models of products with limited adequacy and require an expert’s active involvement in preparing initial information. The design tools for the state-of-the-art full-featured CAD/CAM systems (Siemens NX, Dassault CATIA and PTC Creo Elements / Pro), which are designed to provide CAAP, mainly take into account the geometric constraints that the design imposes on design solutions. These systems often synthesize technologically incorrect assembly sequences in which known technological heuristics are violated, for example orderliness in accuracy, consistency with the system of dimension chains, etc.An AssemBL software application package has been developed for a structured analysis of products and a synthesis of assembly plans and decompositions. The AssemBL uses a hyper-graph model of a product that correctly describes coherent and sequential assembly operations and processes. In terms of the hyper-graph model, an assembly operation is described as shrinkage of edge, an assembly plan is a sequence of shrinkages that converts a hyper-graph into the point, and a decomposition of product into assembly units is a hyper-graph partition into sub-graphs.The AssemBL solves the problem of minimizing the number of direct checks for geometric solvability when assembling complex products. This task is posed as a plus-sum two-person game of bicoloured brushing of an ordered set. In the paradigm of this model, the brushing operation is to check a certain structured fragment for solvability by collision detection methods. A rational brushing strategy minimizes the number of such checks.The package is integrated into the Siemens NX 10.0 computer-aided design system. This solution allowed us to combine specialized AssemBL tools with a developed toolkit of one of the most powerful and popular integrated CAD/CAM /CAE systems.

Download Full-text

A Computational Method for the Identification of Endolysins and Autolysins

Protein and Peptide Letters ◽

10.2174/0929866526666191002104735 ◽

2020 ◽

Vol 27 (4) ◽

pp. 329-336 ◽

Cited By ~ 1

Author(s):

Lei Xu ◽

Guangmin Liang ◽

Baowen Chen ◽

Xu Tan ◽

Huaikun Xiang ◽

...

Keyword(s):

Support Vector Machine ◽

Cell Wall ◽

Experimental Results ◽

Computational Method ◽

Lytic Enzyme ◽

Support Vector ◽

Lytic Enzymes ◽

Data Set ◽

Optimal Feature ◽

Better Than

Background: Cell lytic enzyme is a kind of highly evolved protein, which can destroy the cell structure and kill the bacteria. Compared with antibiotics, cell lytic enzyme will not cause serious problem of drug resistance of pathogenic bacteria. Thus, the study of cell wall lytic enzymes aims at finding an efficient way for curing bacteria infectious. Compared with using antibiotics, the problem of drug resistance becomes more serious. Therefore, it is a good choice for curing bacterial infections by using cell lytic enzymes. Cell lytic enzyme includes endolysin and autolysin and the difference between them is the purpose of the break of cell wall. The identification of the type of cell lytic enzymes is meaningful for the study of cell wall enzymes. Objective: In this article, our motivation is to predict the type of cell lytic enzyme. Cell lytic enzyme is helpful for killing bacteria, so it is meaningful for study the type of cell lytic enzyme. However, it is time consuming to detect the type of cell lytic enzyme by experimental methods. Thus, an efficient computational method for the type of cell lytic enzyme prediction is proposed in our work. Method: We propose a computational method for the prediction of endolysin and autolysin. First, a data set containing 27 endolysins and 41 autolysins is built. Then the protein is represented by tripeptides composition. The features are selected with larger confidence degree. At last, the classifier is trained by the labeled vectors based on support vector machine. The learned classifier is used to predict the type of cell lytic enzyme. Results: Following the proposed method, the experimental results show that the overall accuracy can attain 97.06%, when 44 features are selected. Compared with Ding's method, our method improves the overall accuracy by nearly 4.5% ((97.06-92.9)/92.9%). The performance of our proposed method is stable, when the selected feature number is from 40 to 70. The overall accuracy of tripeptides optimal feature set is 94.12%, and the overall accuracy of Chou's amphiphilic PseAAC method is 76.2%. The experimental results also demonstrate that the overall accuracy is improved by nearly 18% when using the tripeptides optimal feature set. Conclusion: The paper proposed an efficient method for identifying endolysin and autolysin. In this paper, support vector machine is used to predict the type of cell lytic enzyme. The experimental results show that the overall accuracy of the proposed method is 94.12%, which is better than some existing methods. In conclusion, the selected 44 features can improve the overall accuracy for identification of the type of cell lytic enzyme. Support vector machine performs better than other classifiers when using the selected feature set on the benchmark data set.

Download Full-text

In silico Prediction of Inhibitory Constant of Thrombin Inhibitors Using Machine Learning

Combinatorial Chemistry & High Throughput Screening ◽

10.2174/1386207322666181220130232 ◽

2019 ◽

Vol 21 (9) ◽

pp. 662-669 ◽

Cited By ~ 1

Author(s):

Junnan Zhao ◽

Lu Zhu ◽

Weineng Zhou ◽

Lingfeng Yin ◽

Yuchen Wang ◽

...

Keyword(s):

Machine Learning ◽

Prediction Models ◽

Regression Tree ◽

Large Data ◽

Thrombin Inhibitors ◽

Coagulation Cascade ◽

Gradient Boosting ◽

Support Vector ◽

Data Set ◽

Descriptor Selection

Background: Thrombin is the central protease of the vertebrate blood coagulation cascade, which is closely related to cardiovascular diseases. The inhibitory constant Ki is the most significant property of thrombin inhibitors. Method: This study was carried out to predict Ki values of thrombin inhibitors based on a large data set by using machine learning methods. Taking advantage of finding non-intuitive regularities on high-dimensional datasets, machine learning can be used to build effective predictive models. A total of 6554 descriptors for each compound were collected and an efficient descriptor selection method was chosen to find the appropriate descriptors. Four different methods including multiple linear regression (MLR), K Nearest Neighbors (KNN), Gradient Boosting Regression Tree (GBRT) and Support Vector Machine (SVM) were implemented to build prediction models with these selected descriptors. Results: The SVM model was the best one among these methods with R2=0.84, MSE=0.55 for the training set and R2=0.83, MSE=0.56 for the test set. Several validation methods such as yrandomization test and applicability domain evaluation, were adopted to assess the robustness and generalization ability of the model. The final model shows excellent stability and predictive ability and can be employed for rapid estimation of the inhibitory constant, which is full of help for designing novel thrombin inhibitors.

Download Full-text

Rational Design of Colchicine Derivatives as anti-HIV Agents via QSAR and Molecular Docking

Medicinal Chemistry ◽

10.2174/1573406414666180924163756 ◽

2019 ◽

Vol 15 (4) ◽

pp. 328-340 ◽

Cited By ~ 3

Author(s):

Apilak Worachartcheewan ◽

Napat Songtawee ◽

Suphakit Siriwong ◽

Supaluk Prachayasittikul ◽

Chanin Nantasenamat ◽

...

Keyword(s):

Molecular Docking ◽

Rational Design ◽

External Validation ◽

Rational Drug Design ◽

Support Vector ◽

Data Set ◽

Qsar Models ◽

Anti Hiv Agents ◽

Anti Hiv ◽

Colchicine Derivatives

Background: Human immunodeficiency virus (HIV) is an infective agent that causes an acquired immunodeficiency syndrome (AIDS). Therefore, the rational design of inhibitors for preventing the progression of the disease is required. Objective: This study aims to construct quantitative structure-activity relationship (QSAR) models, molecular docking and newly rational design of colchicine and derivatives with anti-HIV activity. Methods: A data set of 24 colchicine and derivatives with anti-HIV activity were employed to develop the QSAR models using machine learning methods (e.g. multiple linear regression (MLR), artificial neural network (ANN) and support vector machine (SVM)), and to study a molecular docking. Results: The significant descriptors relating to the anti-HIV activity included JGI2, Mor24u, Gm and R8p+ descriptors. The predictive performance of the models gave acceptable statistical qualities as observed by correlation coefficient (Q2) and root mean square error (RMSE) of leave-one out cross-validation (LOO-CV) and external sets. Particularly, the ANN method outperformed MLR and SVM methods that displayed LOO−CV 2 Q and RMSELOO-CV of 0.7548 and 0.5735 for LOOCV set, and Ext 2 Q of 0.8553 and RMSEExt of 0.6999 for external validation. In addition, the molecular docking of virus-entry molecule (gp120 envelope glycoprotein) revealed the key interacting residues of the protein (cellular receptor, CD4) and the site-moiety preferences of colchicine derivatives as HIV entry inhibitors for binding to HIV structure. Furthermore, newly rational design of colchicine derivatives using informative QSAR and molecular docking was proposed. Conclusion: These findings serve as a guideline for the rational drug design as well as potential development of novel anti-HIV agents.

Download Full-text

QSAR Study of PARP Inhibitors by GA-MLR, GA-SVM and GA-ANN Approaches

Current Analytical Chemistry ◽

10.2174/1573411016999200518083359 ◽

2020 ◽

Vol 16 (8) ◽

pp. 1088-1105

Author(s):

Nafiseh Vahedi ◽

Majid Mohammadhosseini ◽

Mehdi Nekoei

Keyword(s):

Present Report ◽

Principal Component ◽

Parp Inhibitors ◽

Support Vector ◽

Ann Model ◽

Statistical Parameters ◽

Qsar Study ◽

Data Set ◽

Test Set ◽

Non Linear

Background: The poly(ADP-ribose) polymerases (PARP) is a nuclear enzyme superfamily present in eukaryotes. Methods: In the present report, some efficient linear and non-linear methods including multiple linear regression (MLR), support vector machine (SVM) and artificial neural networks (ANN) were successfully used to develop and establish quantitative structure-activity relationship (QSAR) models capable of predicting pEC50 values of tetrahydropyridopyridazinone derivatives as effective PARP inhibitors. Principal component analysis (PCA) was used to a rational division of the whole data set and selection of the training and test sets. A genetic algorithm (GA) variable selection method was employed to select the optimal subset of descriptors that have the most significant contributions to the overall inhibitory activity from the large pool of calculated descriptors. Results: The accuracy and predictability of the proposed models were further confirmed using crossvalidation, validation through an external test set and Y-randomization (chance correlations) approaches. Moreover, an exhaustive statistical comparison was performed on the outputs of the proposed models. The results revealed that non-linear modeling approaches, including SVM and ANN could provide much more prediction capabilities. Conclusion: Among the constructed models and in terms of root mean square error of predictions (RMSEP), cross-validation coefficients (Q2 LOO and Q2 LGO), as well as R2 and F-statistical value for the training set, the predictive power of the GA-SVM approach was better. However, compared with MLR and SVM, the statistical parameters for the test set were more proper using the GA-ANN model.

Download Full-text

Comparison of Spectroscopic Techniques Combined with Chemometrics for Cocaine Powder Analysis

Journal of Analytical Toxicology ◽

10.1093/jat/bkaa101 ◽

2020 ◽

Vol 44 (8) ◽

pp. 851-860

Author(s):

Joy Eliaerts ◽

Natalie Meert ◽

Pierre Dardenne ◽

Vincent Baeten ◽

Juan-Antonio Fernandez Pierna ◽

...

Keyword(s):

Gas Chromatography ◽

Near Infrared ◽

Evaluation Criteria ◽

Classification Model ◽

Support Vector ◽

Spectroscopic Techniques ◽

Data Set ◽

Promising Tool ◽

Powder Analysis ◽

Mir Spectra

Abstract Spectroscopic techniques combined with chemometrics are a promising tool for analysis of seized drug powders. In this study, the performance of three spectroscopic techniques [Mid-InfraRed (MIR), Raman and Near-InfraRed (NIR)] was compared. In total, 364 seized powders were analyzed and consisted of 276 cocaine powders (with concentrations ranging from 4 to 99 w%) and 88 powders without cocaine. A classification model (using Support Vector Machines [SVM] discriminant analysis) and a quantification model (using SVM regression) were constructed with each spectral dataset in order to discriminate cocaine powders from other powders and quantify cocaine in powders classified as cocaine positive. The performances of the models were compared with gas chromatography coupled with mass spectrometry (GC–MS) and gas chromatography with flame-ionization detection (GC–FID). Different evaluation criteria were used: number of false negatives (FNs), number of false positives (FPs), accuracy, root mean square error of cross-validation (RMSECV) and determination coefficients (R2). Ten colored powders were excluded from the classification data set due to fluorescence background observed in Raman spectra. For the classification, the best accuracy (99.7%) was obtained with MIR spectra. With Raman and NIR spectra, the accuracy was 99.5% and 98.9%, respectively. For the quantification, the best results were obtained with NIR spectra. The cocaine content was determined with a RMSECV of 3.79% and a R2 of 0.97. The performance of MIR and Raman to predict cocaine concentrations was lower than NIR, with RMSECV of 6.76% and 6.79%, respectively and both with a R2 of 0.90. The three spectroscopic techniques can be applied for both classification and quantification of cocaine, but some differences in performance were detected. The best classification was obtained with MIR spectra. For quantification, however, the RMSECV of MIR and Raman was twice as high in comparison with NIR. Spectroscopic techniques combined with chemometrics can reduce the workload for confirmation analysis (e.g., chromatography based) and therefore save time and resources.

Download Full-text

Approach to hand posture recognition based on hand shape features for human–robot interaction

Complex & Intelligent Systems ◽

10.1007/s40747-021-00333-w ◽

2021 ◽

Author(s):

Jing Qi ◽

Kun Xu ◽

Xilun Ding

Keyword(s):

Gaussian Mixture ◽

Human Robot Interaction ◽

Polar Coordinates ◽

Support Vector ◽

Hand Posture ◽

Data Set ◽

Hand Shape ◽

Hand Posture Recognition ◽

Hand Segmentation ◽

Posture Recognition

AbstractHand segmentation is the initial step for hand posture recognition. To reduce the effect of variable illumination in hand segmentation step, a new CbCr-I component Gaussian mixture model (GMM) is proposed to detect the skin region. The hand region is selected as a region of interest from the image using the skin detection technique based on the presented CbCr-I component GMM and a new adaptive threshold. A new hand shape distribution feature described in polar coordinates is proposed to extract hand contour features to solve the false recognition problem in some shape-based methods and effectively recognize the hand posture in cases when different hand postures have the same number of outstretched fingers. A multiclass support vector machine classifier is utilized to recognize the hand posture. Experiments were carried out on our data set to verify the feasibility of the proposed method. The results showed the effectiveness of the proposed approach compared with other methods.

Download Full-text

Correlation between the structure and skin permeability of compounds

Scientific Reports ◽

10.1038/s41598-021-89587-5 ◽

2021 ◽

Vol 11 (1) ◽

Author(s):

Ruolan Zeng ◽

Jiyong Deng ◽

Limin Dang ◽

Xinliang Yu

Keyword(s):

Large Data ◽

Qsar Model ◽

Coefficient Of Determination ◽

Support Vector ◽

Skin Permeability ◽

Data Set ◽

Test Set ◽

Svm Algorithm ◽

Svm Model ◽

Toxicity Relationship

AbstractA three-descriptor quantitative structure–activity/toxicity relationship (QSAR/QSTR) model was developed for the skin permeability of a sufficiently large data set consisting of 274 compounds, by applying support vector machine (SVM) together with genetic algorithm. The optimal SVM model possesses the coefficient of determination R2 of 0.946 and root mean square (rms) error of 0.253 for the training set of 139 compounds; and a R2 of 0.872 and rms of 0.302 for the test set of 135 compounds. Compared with other models reported in the literature, our SVM model shows better statistical performance in a model that deals with more samples in the test set. Therefore, applying a SVM algorithm to develop a nonlinear QSAR model for skin permeability was achieved.

Download Full-text