Enhancing Cardiovascular Risk Prediction: Development of an Advanced Xgboost Model with Hospital-Level Random Effects

Tim Dong; Iyabosola Busola Oronti; Shubhra Sinha; Alberto Freitas; Bing Zhai; Jeremy Chan; Daniel P. Fudulu; Massimo Caputo; Gianni D. Angelini

doi:10.3390/bioengineering11101039

Bioengineering (Oct 2024)

Enhancing Cardiovascular Risk Prediction: Development of an Advanced Xgboost Model with Hospital-Level Random Effects

Tim Dong,
Iyabosola Busola Oronti,
Shubhra Sinha,
Alberto Freitas,
Bing Zhai,
Jeremy Chan,
Daniel P. Fudulu,
Massimo Caputo,
Gianni D. Angelini

Affiliations

Tim Dong: Bristol Heart Institute, Translational Health Sciences, University of Bristol, Bristol BS2 8HW, UK
Iyabosola Busola Oronti: Statistics and Risk Unit (AS&RU), Department of Statistics, School of Engineering, University of Warwick, Coventry CV4 7AL, UK
Shubhra Sinha: Bristol Heart Institute, Translational Health Sciences, University of Bristol, Bristol BS2 8HW, UK
Alberto Freitas: Faculty of Medicine, University of Porto, 4200-319 Porto, Portugal
Bing Zhai: School of Computing Science, Northumbria University, Newcastle upon Tyne NE1 8ST, UK
Jeremy Chan: Bristol Heart Institute, Translational Health Sciences, University of Bristol, Bristol BS2 8HW, UK
Daniel P. Fudulu: Bristol Heart Institute, Translational Health Sciences, University of Bristol, Bristol BS2 8HW, UK
Massimo Caputo: Bristol Heart Institute, Translational Health Sciences, University of Bristol, Bristol BS2 8HW, UK
Gianni D. Angelini: Bristol Heart Institute, Translational Health Sciences, University of Bristol, Bristol BS2 8HW, UK

DOI: https://doi.org/10.3390/bioengineering11101039
Journal volume & issue: Vol. 11, no. 10
p. 1039

Abstract

Read online

Background: Ensemble tree-based models such as Xgboost are highly prognostic in cardiovascular medicine, as measured by the Clinical Effectiveness Metric (CEM). However, their ability to handle correlated data, such as hospital-level effects, is limited. Objectives: The aim of this work is to develop a binary-outcome mixed-effects Xgboost (BME) model that integrates random effects at the hospital level. To ascertain how well the model handles correlated data in cardiovascular outcomes, we aim to assess its performance and compare it to fixed-effects Xgboost and traditional logistic regression models. Methods: A total of 227,087 patients over 17 years of age, undergoing cardiac surgery from 42 UK hospitals between 1 January 2012 and 31 March 2019, were included. The dataset was split into two cohorts: training/validation (n = 157,196; 2012–2016) and holdout (n = 69,891; 2017–2019). The outcome variable was 30-day mortality with hospitals considered as the clustering variable. The logistic regression, mixed-effects logistic regression, Xgboost and binary-outcome mixed-effects Xgboost (BME) were fitted to both standardized and unstandardized datasets across a range of sample sizes and the estimated prediction power metrics were compared to identify the best approach. Results: The exploratory study found high variability in hospital-related mortality across datasets, which supported the adoption of the mixed-effects models. Unstandardized Xgboost BME demonstrated marked improvements in prediction power over the Xgboost model at small sample size ranges, but performance differences decreased as dataset sizes increased. Generalized linear models (glms) and generalized linear mixed-effects models (glmers) followed similar results, with the Xgboost models also excelling at greater sample sizes. Conclusions: These findings suggest that integrating mixed effects into machine learning models can enhance their performance on datasets where the sample size is small.

Published in Bioengineering

ISSN: 2306-5354 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology; Science: Biology (General)
Website: https://www.mdpi.com/journal/bioengineering

About the journal

Abstract

Keywords