Using machine learning-based algorithms to construct cardiovascular risk prediction models for Taiwanese adults based on traditional and novel risk factors

Chien-Hsiang Cheng; Bor-Jen Lee; Oswald Ndi Nfor; Chih-Hsuan Hsiao; Yi-Chia Huang; Yung-Po Liaw

doi:10.1186/s12911-024-02603-2

BMC Medical Informatics and Decision Making (Jul 2024)

Using machine learning-based algorithms to construct cardiovascular risk prediction models for Taiwanese adults based on traditional and novel risk factors

Chien-Hsiang Cheng,
Bor-Jen Lee,
Oswald Ndi Nfor,
Chih-Hsuan Hsiao,
Yi-Chia Huang,
Yung-Po Liaw

Affiliations

Chien-Hsiang Cheng: Department of Respiratory Therapy, Taichung Veterans General Hospital
Bor-Jen Lee: Department of Critical Care Medicine, Tungs’ Taichung Metroharbor Hospital
Oswald Ndi Nfor: Department of Public Health, Institute of Public Health, Chung Shan Medical University
Chih-Hsuan Hsiao: Department of Public Health, Institute of Public Health, Chung Shan Medical University
Yi-Chia Huang: Department of Nutrition, Chung Shan Medical University and Chung Shan Medical University Hospital
Yung-Po Liaw: Department of Public Health, Institute of Public Health, Chung Shan Medical University

DOI: https://doi.org/10.1186/s12911-024-02603-2
Journal volume & issue: Vol. 24, no. 1
pp. 1 – 8

Abstract

Read online

Abstract Objective To develop and validate machine learning models for predicting coronary artery disease (CAD) within a Taiwanese cohort, with an emphasis on identifying significant predictors and comparing the performance of various models. Methods This study involved a comprehensive analysis of clinical, demographic, and laboratory data from 8,495 subjects in Taiwan Biobank (TWB) after propensity score matching to address potential confounding factors. Key variables included age, gender, lipid profiles (T-CHO, HDL_C, LDL_C, TG), smoking and alcohol consumption habits, and renal and liver function markers. The performance of multiple machine learning models was evaluated. Results The cohort comprised 1,699 individuals with CAD identified through self-reported questionnaires. Significant differences were observed between CAD and non-CAD individuals regarding demographics and clinical features. Notably, the Gradient Boosting model emerged as the most accurate, achieving an AUC of 0.846 (95% confidence interval [CI] 0.819–0.873), sensitivity of 0.776 (95% CI, 0.732–0.820), and specificity of 0.759 (95% CI, 0.736–0.782), respectively. The accuracy was 0.762 (95% CI, 0.742–0.782). Age was identified as the most influential predictor of CAD risk within the studied dataset. Conclusion The Gradient Boosting machine learning model demonstrated superior performance in predicting CAD within the Taiwanese cohort, with age being a critical predictor. These findings underscore the potential of machine learning models in enhancing the prediction accuracy of CAD, thereby supporting early detection and targeted intervention strategies. Trial registration Not applicable.

Published in BMC Medical Informatics and Decision Making

ISSN: 1472-6947 (Online)
Publisher: BMC
Country of publisher: United Kingdom
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics
Website: http://bmcmedinformdecismak.biomedcentral.com

About the journal

Abstract

Keywords