Comparing Machine Learning to Regression Methods for Mortality Prediction Using Veterans Affairs Electronic Health Record Clinical Data
Datos Bibliográficos
| ID | 9103213 |
|---|---|
| Autores | Bocheng Jing (0000-0002-2021-5055, San Francisco VA Health Care System, autor de correspondencia), W John Boscardin (0000-0003-3121-9526, San Francisco VA Health Care System, autor de correspondencia), W James Deardorff (0000-0002-7947-3008, Division of Geriatrics), Sun Young Jeon (San Francisco VA Health Care System, autor de correspondencia), Alexandra K Lee (0000-0001-9525-3833, San Francisco VA Health Care System, autor de correspondencia), Anne L Donovan (Anesthesia and Perioperative Medicine, University of California, San Francisco, San Francisco, CA), Sei J Lee (0000-0001-7864-5341, San Francisco VA Health Care System, autor de correspondencia) |
| Año | 2022 |
| Volumen | 60 |
| Número | 6 |
| Páginas | 470-479 |
| Fecha de publicación | 2022-06-01 |
| Peer Reviewed | Sí |
| Open Access | No |
| Tipo | ARTICLE |
| Revista | Medical Care (JOURNAL) |
| Identificadores de la revista | ISSN: 0025-7079 • E-ISSN: 1537-1948 |
| Editorial | Ovid Technologies (Wolters Kluwer Health) (PUBLISHER) |
| DOI | 10.1097/mlr.0000000000001720 |
| PMID | 35352701 |
| OpenAlex | W4220736836 |
| Idioma | EN |
| Referencias citadas | 34 |
BACKGROUND: It is unclear whether machine learning methods yield more accurate electronic health record (EHR) prediction models compared with traditional regression methods. OBJECTIVE: The objective of this study was to compare machine learning and traditional regression models for 10-year mortality prediction using EHR data. DESIGN: This was a cohort study. SETTING: Veterans Affairs (VA) EHR data. PARTICIPANTS: Veterans age above 50 with a primary care visit in 2005, divided into separate training and testing cohorts (n= 124,360 each). MEASUREMENTS AND ANALYTIC METHODS: The primary outcome was 10-year all-cause mortality. We considered 924 potential predictors across a wide range of EHR data elements including demographics (3), vital signs (9), medication classes (399), disease diagnoses (293), laboratory results (71), and health care utilization (149). We compared discrimination (c-statistics), calibration metrics, and diagnostic test characteristics (sensitivity, specificity, and positive and negative predictive values) of machine learning and regression models. RESULTS: Our cohort mean age (SD) was 68.2 (10.5), 93.9% were male; 39.4% died within 10 years. Models yielded testing cohort c-statistics between 0.827 and 0.837. Utilizing all 924 predictors, the Gradient Boosting model yielded the highest c-statistic [0.837, 95% confidence interval (CI): 0.835-0.839]. The full (unselected) logistic regression model had the highest c-statistic of regression models (0.833, 95% CI: 0.830-0.835) but showed evidence of overfitting. The discrimination of the stepwise selection logistic model (101 predictors) was similar (0.832, 95% CI: 0.830-0.834) with minimal overfitting. All models were well-calibrated and had similar diagnostic test characteristics. LIMITATION: Our results should be confirmed in non-VA EHRs. CONCLUSION: The differences in c-statistic between the best machine learning model (924-predictor Gradient Boosting) and 101-predictor stepwise logistic models for 10-year mortality prediction were modest, suggesting stepwise regression methods continue to be a reasonable method for VA EHR mortality prediction model development
Cohort · Confidence interval · Logistic regression · Machine learning · Overfitting · Regression · Statistic · Statistics · Stepwise regression · Veterans Affairs · Artificial Intelligence · Artificial Intelligence in Healthcare and Education · Computer Science · Internal Medicine · Machine Learning in Healthcare · Mathematics · Medicine · Sepsis Diagnosis and Treatment
Super Learner
Building Predictive Models in R Using the caret Package
Assessing the Performance of Prediction Models
Bagging predictors
Ranger
A working guide to boosted regression trees
Regularization Paths for Generalized Linear Models via Coordinate Descent
Bagging Predictors
Greedy function approximation
Random Forests
Regression Shrinkage and Selection Via the Lasso
| Velocidad de citación | historical |
|---|---|
| Altamente citado | No |