Real-world risk stratification for coronary heart disease
A one-year prediction model using health information exchange data
Dados Bibliográficos
| ID | 15359662 |
|---|---|
| Autores | Yaqi Zhang (0000-0002-7103-9864, Guangdong Polytechnic Normal University, autor correspondente), Yifu Mo (0000-0002-2337-4668, China Southern Power Grid (China)), Noriaki Ozawa (Stanford University), Naoto Ozawa, Takumi Ichikawa (Stanford University), Chao-Jung Huang (0000-0002-4293-9492, National Taiwan University), Zhi Han (0000-0002-5340-0070, Stanford University), Lü Tian (0000-0002-5893-0169, Stanford University), Shaun T Alfreds (0000-0003-1752-4823, Gulf of Maine Research Institute), Karl G Sylvester (0000-0002-8559-0155, Stanford University), Doff B McElhinney (0000-0001-7242-0934, Stanford University), Xuefeng B Ling (0000-0002-5386-3884, Stanford Medicine) |
| Ano | 2025 |
| Volume | 25 |
| Fascículo | 1 |
| Páginas | 3218-3218 |
| Data de publicação | 2025-09-30 |
| Peer Reviewed | Sim |
| Open Access | Sim |
| Tipo | ARTICLE |
| Periódico | BMC Public Health (JOURNAL) |
| Identificadores do periódico | ISSN: 1471-2458 • E-ISSN: 1471-2458 |
| Editora | BioMed Central (PUBLISHER • GB) |
| DOI | 10.1186/s12889-025-24266-y |
| PMID | 41029641 |
| OpenAlex | W4414663968 |
| Idioma | EN |
| Referências citadas | 37 |
Coronary heart disease (CHD), the most common form of heart disease, progresses over years before culminating in serious cardiac events. Early prediction and intervention are critical to reducing CHD-related morbidity, mortality, and healthcare burden. To develop and validate a machine learning model using statewide electronic health records (EHRs) to predict 1-year risk of CHD in the general population of Maine, enabling targeted preventive strategies. Two population-based cohorts were constructed from the Maine Health Information Exchange (HIE): a retrospective cohort for model training and calibration (2015–2017, N = 1,042,124), and a prospective cohort for external validation (2016–2018, N = 1,040,158). EHR features included demographics, diagnoses, procedures, medications, labs, and utilization metrics. A multistage modeling pipeline—comprising statistical filtering, XGBoost-based feature selection, risk prediction, and isotonic regression calibration—was used to construct the final model. Validation included discrimination, calibration, and survival analysis. The final XGBoost model achieved strong discrimination: AUC = 0.952 (95% CI: 0.950–0.954) in the retrospective cohort and 0.888 (95% CI: 0.885–0.890) in the prospective cohort. Based on calibrated risk probabilities, the population was stratified into five risk categories: very low (92.30%, N = 960,021), low (6.79%, N = 70,676), medium (0.85%, N = 8,888), high (0.05%, N = 554), and very high (0.002%, N = 19). Among the very high-risk group, 11 individuals (57.89%) developed CHD within one year. This statewide, HIE-based CHD risk prediction model demonstrates robust performance and real-world applicability. It enables early identification of high-risk individuals and supports population-scale precision prevention through evidence-informed, proactive care
Biostatistics · Cohort · Cohort study · Framingham Risk Score · Health information exchange · Population · Prospective cohort study · Retrospective cohort study · Risk assessment · Artificial Intelligence in Healthcare · Cardiovascular Health and Risk Factors · Machine Learning in Healthcare
| Velocidade de citação | historical |
|---|---|
| Altamente citado | Não |