Development and validation of prediction models for hypertension risks: A cross-sectional study based on 4,287,407 participants

Weidong Ji; Yushan Zhang; Yinlin Cheng; Yushan Wang; Yi Zhou

doi:10.3389/fcvm.2022.928948

Frontiers in Cardiovascular Medicine (Sep 2022)

Development and validation of prediction models for hypertension risks: A cross-sectional study based on 4,287,407 participants

Weidong Ji,
Yushan Zhang,
Yinlin Cheng,
Yushan Wang,
Yi Zhou

Affiliations

Weidong Ji: Department of Medical Information, Zhongshan School of Medicine, Sun Yat-sen University, Guangzhou, China
Yushan Zhang: Department of Maternal and Child Health, School of Public Health, Sun Yat-sen University, Guangzhou, China
Yinlin Cheng: Department of Medical Information, Zhongshan School of Medicine, Sun Yat-sen University, Guangzhou, China
Yushan Wang: Center of Health Management, The First Affiliated Hospital of Xinjiang Medical University, Urumqi, China
Yi Zhou: Department of Medical Information, Zhongshan School of Medicine, Sun Yat-sen University, Guangzhou, China

DOI: https://doi.org/10.3389/fcvm.2022.928948
Journal volume & issue: Vol. 9

Abstract

Read online

ObjectiveTo develop an optimal screening model to identify the individuals with a high risk of hypertension in China by comparing tree-based machine learning models, such as classification and regression tree, random forest, adaboost with a decision tree, extreme gradient boosting decision tree, and other machine learning models like an artificial neural network, naive Bayes, and traditional logistic regression models.MethodsA total of 4,287,407 adults participating in the national physical examination were included in the study. Features were selected using the least absolute shrinkage and selection operator regression. The Borderline synthetic minority over-sampling technique was used for data balance. Non-laboratory and semi-laboratory analyses were carried out in combination with the selected features. The tree-based machine learning models, other machine learning models, and traditional logistic regression models were constructed to identify individuals with hypertension, respectively. Top features selected using the best algorithm and the corresponding variable importance score were visualized.ResultsA total of 24 variables were finally included for analyses after the least absolute shrinkage and selection operator regression model. The sample size of hypertensive patients in the training set was expanded from 689,025 to 2,312,160 using the borderline synthetic minority over-sampling technique algorithm. The extreme gradient boosting decision tree algorithm showed the best results (area under the receiver operating characteristic curve of non-laboratory: 0.893 and area under the receiver operating characteristic curve of semi-laboratory: 0.894). This study found that age, systolic blood pressure, waist circumference, diastolic blood pressure, albumin, drinking frequency, electrocardiogram, ethnicity (uyghur, hui, and other), body mass index, sex (female), exercise frequency, diabetes mellitus, and total bilirubin are important factors reflecting hypertension. Besides, some algorithms included in the semi-laboratory analyses showed less improvement in the predictive performance compared to the non-laboratory analyses.ConclusionUsing multiple methods, a more significant prediction model can be built, which discovers risk factors and provides new insights into the prediction and prevention of hypertension.

Published in Frontiers in Cardiovascular Medicine

ISSN: 2297-055X (Online)
Publisher: Frontiers Media S.A.
Country of publisher: Switzerland
LCC subjects: Medicine: Internal medicine: Specialties of internal medicine: Diseases of the circulatory (Cardiovascular) system
Website: https://www.frontiersin.org/journals/cardiovascular-medicine

About the journal

Abstract

Keywords