Generalizing Gain Penalization for Feature Selection in Tree-Based Models

Bruna Wundervald; Andrew C. Parnell; Katarina Domijan

doi:10.1109/ACCESS.2020.3032095

IEEE Access (Jan 2020)

Generalizing Gain Penalization for Feature Selection in Tree-Based Models

Bruna Wundervald,
Andrew C. Parnell,
Katarina Domijan

Affiliations

Bruna Wundervald: ORCiD; Hamilton Institute, National University of Ireland Maynooth, Maynooth, Ireland
Andrew C. Parnell: Hamilton Institute, National University of Ireland Maynooth, Maynooth, Ireland
Katarina Domijan: ORCiD; Hamilton Institute, National University of Ireland Maynooth, Maynooth, Ireland

DOI: https://doi.org/10.1109/ACCESS.2020.3032095
Journal volume & issue: Vol. 8
pp. 190231 – 190239

Abstract

Read online

We develop a new approach for feature selection via gain penalization in tree-based models. First, we show that previous methods do not perform sufficient regularization and often exhibit sub-optimal out-of-sample performance, especially when correlated features are present. Instead, we develop a new gain penalization idea that exhibits a general local-global regularization for tree-based models. The new method allows for full flexibility in the choice of feature-specific importance weights, while also applying a global penalization. We validate our method on both simulated and real data, exploring how the hyperparameters interact and we provide the implementation as an extension of the popular R package ranger.

Published in IEEE Access

ISSN: 2169-3536 (Online)
Publisher: IEEE
Country of publisher: United States
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering
Website: https://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=6287639

About the journal

Abstract

Keywords