Scientific Data (Mar 2025)
A Data Engineering Framework for Ethereum Beacon Chain Rewards: From Data Collection to Decentralization Metrics
Abstract
Abstract Ethereum, one of the leading smart contract blockchain platforms, currently operates on a Proof-of-Stake (PoS) consensus mechanism designed to secure the network while incentivizing desired validator behaviors. Despite blockchain technology’s promise of decentralization, limitations and gaps in decentralization persist, posing challenges for analysis and optimization. This study introduces a comprehensive dataset of validator rewards from the Ethereum Beacon chain, categorized into attestation, proposer, and sync committee rewards. By providing granular, transparent, and auditable records of validator activities, the dataset addresses the fragmentation of raw blockchain data and enables robust evaluations of PoS incentive structures. Researchers can leverage this dataset to assess enforceable rules, verify protocol compliance, and analyze long-term validator behavior. In addition, we apply decentralization metrics such as the Shannon entropy, Gini Index, Nakamoto Coefficient, and Herfindahl-Hirschman Index (HHI) to showcase the dataset’s utility in studying decentralization trends. Publicly available on Harvard Dataverse and accompanied by open-source analytical tools on GitHub, this dataset facilitates future research aimed at enhancing blockchain systems’ decentralization, security, and efficiency.