Audacity of huge: overcoming challenges of data scarcity and data quality for machine learning in computational materials discovery

A machine learning model is used as a surrogate fitness evaluator in a genetic algorithm (GA) optimization of the atomic distribution of Pt-Au nanoparticles. The machine learning accelerated genetic algorithm (MLaGA) yields a 50-fold reduction of required energy calculations compared to a traditional GA.

Download Full-text

Using Machine Learning for Dependable Outlier Detection in Environmental Monitoring Systems

ACM Transactions on Cyber-Physical Systems ◽

10.1145/3445812 ◽

2021 ◽

Vol 5 (3) ◽

pp. 1-30

Author(s):

Gonçalo Jesus ◽

António Casimiro ◽

Anabela Oliveira

Keyword(s):

Machine Learning ◽

Environmental Monitoring ◽

Data Quality ◽

Outlier Detection ◽

Prediction Models ◽

Sensor Data ◽

Natural Phenomenon ◽

Monitoring Systems ◽

Data Errors ◽

Redundant Data

Sensor platforms used in environmental monitoring applications are often subject to harsh environmental conditions while monitoring complex phenomena. Therefore, designing dependable monitoring systems is challenging given the external disturbances affecting sensor measurements. Even the apparently simple task of outlier detection in sensor data becomes a hard problem, amplified by the difficulty in distinguishing true data errors due to sensor faults from deviations due to natural phenomenon, which look like data errors. Existing solutions for runtime outlier detection typically assume that the physical processes can be accurately modeled, or that outliers consist in large deviations that are easily detected and filtered by appropriate thresholds. Other solutions assume that it is possible to deploy multiple sensors providing redundant data to support voting-based techniques. In this article, we propose a new methodology for dependable runtime detection of outliers in environmental monitoring systems, aiming to increase data quality by treating them. We propose the use of machine learning techniques to model each sensor behavior, exploiting the existence of correlated data provided by other related sensors. Using these models, along with knowledge of processed past measurements, it is possible to obtain accurate estimations of the observed environment parameters and build failure detectors that use these estimations. When a failure is detected, these estimations also allow one to correct the erroneous measurements and hence improve the overall data quality. Our methodology not only allows one to distinguish truly abnormal measurements from deviations due to complex natural phenomena, but also allows the quantification of each measurement quality, which is relevant from a dependability perspective. We apply the methodology to real datasets from a complex aquatic monitoring system, measuring temperature and salinity parameters, through which we illustrate the process for building the machine learning prediction models using a technique based on Artificial Neural Networks, denoted ANNODE ( ANN Outlier Detection ). From this application, we also observe the effectiveness of our ANNODE approach for accurate outlier detection in harsh environments. Then we validate these positive results by comparing ANNODE with state-of-the-art solutions for outlier detection. The results show that ANNODE improves existing solutions regarding accuracy of outlier detection.

Download Full-text

Data Quality Measures and Efficient Evaluation Algorithms for Large-Scale High-Dimensional Data

Applied Sciences ◽

10.3390/app11020472 ◽

2021 ◽

Vol 11 (2) ◽

pp. 472

Author(s):

Hyeongmin Cho ◽

Sangkyun Lee

Keyword(s):

Machine Learning ◽

Data Quality ◽

Large Scale ◽

High Dimensional Data ◽

Quality Measures ◽

Training Data ◽

Measure Data ◽

High Dimensional ◽

Small Scale ◽

Class Separability

Machine learning has been proven to be effective in various application areas, such as object and speech recognition on mobile systems. Since a critical key to machine learning success is the availability of large training data, many datasets are being disclosed and published online. From a data consumer or manager point of view, measuring data quality is an important first step in the learning process. We need to determine which datasets to use, update, and maintain. However, not many practical ways to measure data quality are available today, especially when it comes to large-scale high-dimensional data, such as images and videos. This paper proposes two data quality measures that can compute class separability and in-class variability, the two important aspects of data quality, for a given dataset. Classical data quality measures tend to focus only on class separability; however, we suggest that in-class variability is another important data quality factor. We provide efficient algorithms to compute our quality measures based on random projections and bootstrapping with statistical benefits on large-scale high-dimensional data. In experiments, we show that our measures are compatible with classical measures on small-scale data and can be computed much more efficiently on large-scale high-dimensional datasets.

Download Full-text

A Machine Learning Approach for Data Quality Control of Earth Observation Data Management System

IGARSS 2020 - 2020 IEEE International Geoscience and Remote Sensing Symposium ◽

10.1109/igarss39084.2020.9323615 ◽

2020 ◽

Author(s):

Weiguo Han ◽

Matthew Jochum

Keyword(s):

Machine Learning ◽

Quality Control ◽

Data Quality ◽

Earth Observation ◽

Data Management System ◽

Learning Approach ◽

Observation Data ◽

Data Quality Control ◽

Machine Learning Approach ◽

Earth Observation Data

Download Full-text

Improving Power Grid Monitoring Data Quality: An Efficient Machine Learning Framework for Missing Data Prediction

2015 IEEE 17th International Conference on High Performance Computing and Communications, 2015 IEEE 7th International Symposium on Cyberspace Safety and Security, and 2015 IEEE 12th International Conference on Embedded Software and Systems ◽

10.1109/hpcc-css-icess.2015.16 ◽

2015 ◽

Cited By ~ 10

Author(s):

Weiwei Shi ◽

Yongxin Zhu ◽

Jinkui Zhang ◽

Xiang Tao ◽

Gehao Sheng ◽

...

Keyword(s):

Machine Learning ◽

Missing Data ◽

Data Quality ◽

Power Grid ◽

Monitoring Data ◽

Learning Framework ◽

Data Prediction ◽

Grid Monitoring ◽

Efficient Machine ◽

Missing Data Prediction

Download Full-text

Application of Machine Learning for Lithology-on-Bit Prediction using Drilling Data in Real-Time

10.2118/206622-ms ◽

2021 ◽

Author(s):

Temirlan Zhekenov ◽

Artem Nechaev ◽

Kamilla Chettykbayeva ◽

Alexey Zinovyev ◽

German Sardarov ◽

...

Keyword(s):

Machine Learning ◽

Data Quality ◽

Real Time ◽

Hybrid Modeling ◽

Data Set ◽

Drilling Parameters ◽

Drilling Data ◽

Stable Solutions ◽

Mud Logging

SUMMARY Researchers base their analysis on basic drilling parameters obtained during mud logging and demonstrate impressive results. However, due to limitations imposed by data quality often present during drilling, those solutions often tend to lose their stability and high levels of predictivity. In this work, the concept of hybrid modeling was introduced which allows to integrate the analytical correlations with algorithms of machine learning for obtaining stable solutions consistent from one data set to another.

Download Full-text

Application of Machine Learning for Oilfield Data Quality Improvement

10.2118/191601-18rptc-ms ◽

2018 ◽

Cited By ~ 1

Author(s):

Alla Andrianova ◽

Maxim Simonov ◽

Dmitry Perets ◽

Andrey Margarit ◽

Darya Serebryakova ◽

...

Keyword(s):

Machine Learning ◽

Quality Improvement ◽

Data Quality

Download Full-text

A study of real-world micrograph data quality and machine learning model robustness

npj Computational Materials ◽

10.1038/s41524-021-00616-3 ◽

2021 ◽

Vol 7 (1) ◽

Author(s):

Xiaoting Zhong ◽

Brian Gallagher ◽

Keenan Eves ◽

Emily Robertson ◽

T. Nathan Mundhenk ◽

...

Keyword(s):

Machine Learning ◽

Data Quality ◽

Real World ◽

Image Feature ◽

Molecular Solids ◽

Pixel Intensity ◽

Machine Learning Model ◽

Model Predictions ◽

Intensity Normalization ◽

Model Robustness

AbstractMachine-learning (ML) techniques hold the potential of enabling efficient quantitative micrograph analysis, but the robustness of ML models with respect to real-world micrograph quality variations has not been carefully evaluated. We collected thousands of scanning electron microscopy (SEM) micrographs for molecular solid materials, in which image pixel intensities vary due to both the microstructure content and microscope instrument conditions. We then built ML models to predict the ultimate compressive strength (UCS) of consolidated molecular solids, by encoding micrographs with different image feature descriptors and training a random forest regressor, and by training an end-to-end deep-learning (DL) model. Results show that instrument-induced pixel intensity signals can affect ML model predictions in a consistently negative way. As a remedy, we explored intensity normalization techniques. It is seen that intensity normalization helps to improve micrograph data quality and ML model robustness, but microscope-induced intensity variations can be difficult to eliminate.

Download Full-text

Benchmarking Graph Neural Networks for Materials Chemistry

10.26434/chemrxiv.13615421.v2 ◽

2021 ◽

Author(s):

Victor Fung ◽

Jiaxin Zhang ◽

Eric Juarez ◽

Bobby Sumpter

Keyword(s):

Machine Learning ◽

Neural Networks ◽

Materials Chemistry ◽

Learning Models ◽

High Data ◽

Hyperparameter Selection ◽

Computational Materials ◽

Conventional Models ◽

Graph Neural Networks ◽

Machine Learning Models

Graph neural networks (GNNs) have received intense interest as a rapidly expanding class of machine learning models remarkably well-suited for materials applications. To date, a number of successful GNNs have been proposed and demonstrated for systems ranging from crystal stability to electronic property prediction and to surface chemistry and heterogeneous catalysis. However, a consistent benchmark of these models remains lacking, hindering the development and consistent evaluation of new models in the materials field. Here, we present a workflow and testing platform, MatDeepLearn, for quickly and reproducibly assessing and comparing GNNs and other machine learning models. We use this platform to optimize and evaluate a selection of top performing GNNs on several representative datasets in computational materials chemistry. From our investigations we note the importance of hyperparameter selection and find roughly similar performances for the top models once optimized. We identify several strengths in GNNs over conventional models in cases with compositionally diverse datasets and in its overall flexibility with respect to inputs, due to learned rather than defined representations. Meanwhile several weaknesses of GNNs are also observed including high data requirements, and suggestions for further improvement for applications in materials chemistry are proposed.

Download Full-text

Is data quality enough for a clinical decision?: Apply machine learning and avoid bias

2017 IEEE International Conference on Big Data (Big Data) ◽

10.1109/bigdata.2017.8258221 ◽

2017 ◽

Cited By ~ 1

Author(s):

Kim Hee

Keyword(s):

Machine Learning ◽

Data Quality ◽

Clinical Decision

Download Full-text