Use of Sequential Hot-Deck Imputation for Missing Health Care Systems Data for Population Health Research
Bibliographic Data
| ID | 9104036 |
|---|---|
| Authors | Ella A Chrenka (0000-0002-3750-7062, HealthPartners Institute, Bloomington, MN, corresponding author), Steven P Dehmer (0000-0003-3508-9239, HealthPartners Institute, Bloomington, MN, corresponding author), Michael V Maciosek (0000-0002-8164-8858, HealthPartners Institute, Bloomington, MN, corresponding author), Inih J Essien (HealthPartners Institute, Bloomington, MN), Inih Essien (0000-0003-0775-8255, H ea lt hP ar tn er s, corresponding author), Bjorn C Westgard (HealthPartners Institute, Bloomington, MN), Bjorn Westgard (0000-0003-1275-753X, Regions Hospital, corresponding author) |
| Year | 2024 |
| Volume | 62 |
| Issue | 5 |
| Pages | 319-325 |
| Publication date | 2024-05-01 |
| Peer Reviewed | Yes |
| Open Access | No |
| Type | ARTICLE |
| Venue | Medical Care (JOURNAL) |
| Journal identifiers | ISSN: 0025-7079 • E-ISSN: 1537-1948 |
| Publisher | Ovid Technologies (Wolters Kluwer Health) (PUBLISHER) |
| DOI | 10.1097/mlr.0000000000001995 |
| PMID | 38546379 |
| OpenAlex | W4393253378 |
| Language | EN |
| References cited | 26 |
Electronic medical record (EMR) data present many opportunities for population health research. The use of EMR data for population risk models can be impeded by the high proportion of missingness in key patient variables. Common approaches like complete case analysis and multiple imputation may not be appropriate for some population health initiatives that require a single, complete analytic data set. In this study, we demonstrate a sequential hot-deck imputation (HDI) procedure to address missingness in a set of cardiometabolic measures in an EMR data set. We assessed the performance of sequential HDI within the individual variables and a commonly used composite risk score. A data set of cardiometabolic measures based on EMR data from 2 large urban hospitals was used to create a benchmark data set with simulated missingness. Sequential HDI was applied, and the resulting data were used to calculate atherosclerotic cardiovascular disease risk scores. The performance of the imputation approach was assessed using a set of metrics to evaluate the distribution and validity of the imputed data. Of the 567,841 patients, 65% had at least 1 missing cardiometabolic measure. Sequential HDI resulted in the distribution of variables and risk scores that reflected those in the simulated data while retaining correlation. When stratified by age and sex, risk scores were plausible and captured patterns expected in the general population. The use of sequential HDI was shown to be a suitable approach to multivariate missingness in EMR data. Sequential HDI could benefit population health research by providing a straightforward, computationally nonintensive approach to missing EMR data that results in a single analytic data set
Data mining · Data set · Environmental health · Imputation (statistics) · Missing data · Multivariate statistics · Population · Statistics · Chronic Disease Management Strategies · Computer Science · Machine Learning in Healthcare · Mathematics · Medical Coding and Health Information · Medicine
Heart Disease and Stroke Statistics—2019 Update
Heart Disease and Stroke Statistics—2017 Update
Imputation with the R Package VIM
The Impact of eHealth on the Quality and Safety of Health Care
An overview of clinical decision support systems
Dealing with missing data in a multi-question depression scale
A Review of Hot Deck Imputation for Survey Non‐response
| Citation velocity | historical |
|---|---|
| Highly cited | No |