Contextual Multi-Armed Bandit With Costly Feature Observation in Non-Stationary Environments

Saeed Ghoorchian; Evgenii Kortukov; Setareh Maghsudi

doi:10.1109/OJSP.2024.3389809

IEEE Open Journal of Signal Processing (Jan 2024)

Contextual Multi-Armed Bandit With Costly Feature Observation in Non-Stationary Environments

Saeed Ghoorchian,
Evgenii Kortukov,
Setareh Maghsudi

Affiliations

Saeed Ghoorchian: ORCiD; Faculty of Electrical Engineering and Information Technology, Ruhr-University Bochum, Bochum, Germany
Evgenii Kortukov: ORCiD; Faculty of Mathematics and Natural Sciences, Tübingen University, Tübingen, Germany
Setareh Maghsudi: ORCiD; Faculty of Electrical Engineering and Information Technology, Ruhr-University Bochum, Bochum, Germany

DOI: https://doi.org/10.1109/OJSP.2024.3389809
Journal volume & issue: Vol. 5
pp. 820 – 830

Abstract

Read online

Maximizing long-term rewards is the primary goal in sequential decision-making problems. The majority of existing methods assume that side information is freely available, enabling the learning agent to observe all features' states before making a decision. In real-world problems, however, collecting beneficial information is often costly. That implies that, besides individual arms' reward, learning the observations of the features' states is essential to improve the decision-making strategy. The problem is aggravated in a non-stationary environment where reward and cost distributions undergo abrupt changes over time. To address the aforementioned dual learning problem, we extend the contextual bandit setting and allow the agent to observe subsets of features' states. The objective is to maximize the long-term average gain, which is the difference between the accumulated rewards and the paid costs on average. Therefore, the agent faces a trade-off between minimizing the cost of information acquisition and possibly improving the decision-making process using the obtained information. To this end, we develop an algorithm that guarantees a sublinear regret in time. Numerical results demonstrate the superiority of our proposed policy in a real-world scenario.

Published in IEEE Open Journal of Signal Processing

ISSN: 2644-1322 (Online)
Publisher: IEEE
Country of publisher: United States
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering
Website: https://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=8782710

About the journal

Abstract

Keywords