Development of a Machine Learning Model Using Electronic Health Record Data to Identify Antibiotic Use Among Hospitalized Patients

Abstract Background Stroke severity is an important predictor of patient outcomes and is commonly measured with the National Institutes of Health Stroke Scale (NIHSS) scores. Because these scores are often recorded as free text in physician reports, structured real-world evidence databases seldom include the severity. The aim of this study was to use machine learning models to impute NIHSS scores for all patients with newly diagnosed stroke from multi-institution electronic health record (EHR) data. Methods NIHSS scores available in the Optum© de-identified Integrated Claims-Clinical dataset were extracted from physician notes by applying natural language processing (NLP) methods. The cohort analyzed in the study consists of the 7149 patients with an inpatient or emergency room diagnosis of ischemic stroke, hemorrhagic stroke, or transient ischemic attack and a corresponding NLP-extracted NIHSS score. A subset of these patients (n = 1033, 14%) were held out for independent validation of model performance and the remaining patients (n = 6116, 86%) were used for training the model. Several machine learning models were evaluated, and parameters optimized using cross-validation on the training set. The model with optimal performance, a random forest model, was ultimately evaluated on the holdout set. Results Leveraging machine learning we identified the main factors in electronic health record data for assessing stroke severity, including death within the same month as stroke occurrence, length of hospital stay following stroke occurrence, aphagia/dysphagia diagnosis, hemiplegia diagnosis, and whether a patient was discharged to home or self-care. Comparing the imputed NIHSS scores to the NLP-extracted NIHSS scores on the holdout data set yielded an R2 (coefficient of determination) of 0.57, an R (Pearson correlation coefficient) of 0.76, and a root-mean-squared error of 4.5. Conclusions Machine learning models built on EHR data can be used to determine proxies for stroke severity. This enables severity to be incorporated in studies of stroke patient outcomes using administrative and EHR databases.

Download Full-text

DEVELOPMENT OF A PREDICTION MODEL FOR INCIDENT MYOCARDIAL INFARCTION USING MACHINE LEARNING APPLIED TO HARMONIZED ELECTRONIC HEALTH RECORD DATA

Journal of the American College of Cardiology ◽

10.1016/s0735-1097(20)30821-4 ◽

2020 ◽

Vol 75 (11) ◽

pp. 194

Author(s):

Divneet Mandair ◽

Premanand Tiwari ◽

Steven Simon ◽

Michael Rosenberg

Keyword(s):

Machine Learning ◽

Myocardial Infarction ◽

Prediction Model ◽

Electronic Health Record ◽

Health Record ◽

Electronic Health Record Data ◽

Record Data ◽

Electronic Health

Download Full-text

Use of electronic health record data and machine learning to identify candidates for HIV pre-exposure prophylaxis: a modelling study

The Lancet HIV ◽

10.1016/s2352-3018(19)30137-7 ◽

2019 ◽

Vol 6 (10) ◽

pp. e688-e695 ◽

Cited By ~ 19

Author(s):

Julia L Marcus ◽

Leo B Hurley ◽

Douglas S Krakower ◽

Stacey Alexeeff ◽

Michael J Silverberg ◽

...

Keyword(s):

Machine Learning ◽

Electronic Health Record ◽

Health Record ◽

Electronic Health Record Data ◽

Record Data ◽

Electronic Health ◽

Exposure Prophylaxis ◽

Modelling Study

Download Full-text

Prediction of Acute Kidney Injury With a Machine Learning Algorithm Using Electronic Health Record Data

Canadian Journal of Kidney Health and Disease ◽

10.1177/2054358118776326 ◽

2018 ◽

Vol 5 ◽

pp. 205435811877632 ◽

Cited By ~ 31

Author(s):

Hamid Mohamadlou ◽

Anna Lynn-Palevsky ◽

Christopher Barton ◽

Uli Chettipally ◽

Lisa Shieh ◽

...

Keyword(s):

Machine Learning ◽

Acute Kidney Injury ◽

Electronic Health Record ◽

Learning Algorithm ◽

Kidney Injury ◽

Health Record ◽

Machine Learning Algorithm ◽

Electronic Health Record Data ◽

Record Data ◽

Electronic Health

Download Full-text

Classifying Pseudogout using Machine Learning Approaches with Electronic Health Record Data

Arthritis Care & Research ◽

10.1002/acr.24132 ◽

2020 ◽

Cited By ~ 1

Author(s):

Sara K. Tedeschi ◽

Tianrun Cai ◽

Zeling He ◽

Yuri Ahuja ◽

Chuan Hong ◽

...

Keyword(s):

Machine Learning ◽

Electronic Health Record ◽

Health Record ◽

Learning Approaches ◽

Electronic Health Record Data ◽

Record Data ◽

Electronic Health

Download Full-text

Subtypes in patients with opioid misuse: A prognostic enrichment strategy using electronic health record data in hospitalized patients

PLoS ONE ◽

10.1371/journal.pone.0219717 ◽

2019 ◽

Vol 14 (7) ◽

pp. e0219717 ◽

Cited By ~ 2

Author(s):

Majid Afshar ◽

Cara Joyce ◽

Dmitriy Dligach ◽

Brihat Sharma ◽

Robert Kania ◽

...

Keyword(s):

Electronic Health Record ◽

Hospitalized Patients ◽

Health Record ◽

Opioid Misuse ◽

Electronic Health Record Data ◽

Record Data ◽

Electronic Health ◽

Enrichment Strategy

Download Full-text

Machine Learning Electronic Health Record Identification of Patients with Rheumatoid Arthritis: Algorithm Pipeline Development and Validation Study (Preprint)

10.2196/preprints.23930 ◽

2020 ◽

Author(s):

Tjardo D Maarseveen ◽

Timo Meinderink ◽

Marcel J T Reinders ◽

Johannes Knitza ◽

Tom W J Huizinga ◽

...

Keyword(s):

Rheumatoid Arthritis ◽

Machine Learning ◽

Electronic Health Record ◽

Support Vector ◽

Free Text ◽

Health Record ◽

Electronic Health Record Data ◽

Data Set ◽

Record Data ◽

Electronic Health

BACKGROUND Financial codes are often used to extract diagnoses from electronic health records. This approach is prone to false positives. Alternatively, queries are constructed, but these are highly center and language specific. A tantalizing alternative is the automatic identification of patients by employing machine learning on format-free text entries. OBJECTIVE The aim of this study was to develop an easily implementable workflow that builds a machine learning algorithm capable of accurately identifying patients with rheumatoid arthritis from format-free text fields in electronic health records. METHODS Two electronic health record data sets were employed: Leiden (n=3000) and Erlangen (n=4771). Using a portion of the Leiden data (n=2000), we compared 6 different machine learning methods and a naïve word-matching algorithm using 10-fold cross-validation. Performances were compared using the area under the receiver operating characteristic curve (AUROC) and the area under the precision recall curve (AUPRC), and F1 score was used as the primary criterion for selecting the best method to build a classifying algorithm. We selected the optimal threshold of positive predictive value for case identification based on the output of the best method in the training data. This validation workflow was subsequently applied to a portion of the Erlangen data (n=4293). For testing, the best performing methods were applied to remaining data (Leiden n=1000; Erlangen n=478) for an unbiased evaluation. RESULTS For the Leiden data set, the word-matching algorithm demonstrated mixed performance (AUROC 0.90; AUPRC 0.33; F1 score 0.55), and 4 methods significantly outperformed word-matching, with support vector machines performing best (AUROC 0.98; AUPRC 0.88; F1 score 0.83). Applying this support vector machine classifier to the test data resulted in a similarly high performance (F1 score 0.81; positive predictive value [PPV] 0.94), and with this method, we could identify 2873 patients with rheumatoid arthritis in less than 7 seconds out of the complete collection of 23,300 patients in the Leiden electronic health record system. For the Erlangen data set, gradient boosting performed best (AUROC 0.94; AUPRC 0.85; F1 score 0.82) in the training set, and applied to the test data, resulted once again in good results (F1 score 0.67; PPV 0.97). CONCLUSIONS We demonstrate that machine learning methods can extract the records of patients with rheumatoid arthritis from electronic health record data with high precision, allowing research on very large populations for limited costs. Our approach is language and center independent and could be applied to any type of diagnosis. We have developed our pipeline into a universally applicable and easy-to-implement workflow to equip centers with their own high-performing algorithm. This allows the creation of observational studies of unprecedented size covering different countries for low cost from already available data in electronic health record systems.

Download Full-text