Addressing bias in bagging and boosting regression models

Juliette Ugirumurera; Erik A. Bensen; Joseph Severino; Jibonananda Sanyal

doi:10.1038/s41598-024-68907-5

Scientific Reports (Aug 2024)

Addressing bias in bagging and boosting regression models

Juliette Ugirumurera,
Erik A. Bensen,
Joseph Severino,
Jibonananda Sanyal

Affiliations

Juliette Ugirumurera: Computational Science Center, National Renewable Energy Laboratory
Erik A. Bensen: Department of Statistics and Data Science, Carnegie Mellon University
Joseph Severino: Computational Science Center, National Renewable Energy Laboratory
Jibonananda Sanyal: Computational Science Center, National Renewable Energy Laboratory

DOI: https://doi.org/10.1038/s41598-024-68907-5
Journal volume & issue: Vol. 14, no. 1
pp. 1 – 12

Abstract

Read online

Abstract As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model’s loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

Published in Scientific Reports

ISSN: 2045-2322 (Online)
Publisher: Nature Portfolio
Country of publisher: United Kingdom
LCC subjects: Medicine; Science
Website: https://www.nature.com/srep/

About the journal

Abstract

Keywords