Communications Physics (Oct 2023)

The autoregressive neural network architecture of the Boltzmann distribution of pairwise interacting spins systems

  • Indaco Biazzo

DOI
https://doi.org/10.1038/s42005-023-01416-5
Journal volume & issue
Vol. 6, no. 1
pp. 1 – 10

Abstract

Read online

Abstract Autoregressive Neural Networks (ARNNs) have shown exceptional results in generation tasks across image, language, and scientific domains. Despite their success, ARNN architectures often operate as black boxes without a clear connection to underlying physics or statistical models. This research derives an exact mapping of the Boltzmann distribution of binary pairwise interacting systems in autoregressive form. The parameters of the ARNN are directly related to the Hamiltonian’s couplings and external fields, and commonly used structures like residual connections and recurrent architecture emerge from the derivation. This explicit formulation leverages statistical physics techniques to derive ARNNs for specific systems. Using the Curie–Weiss and Sherrington–Kirkpatrick models as examples, the proposed architectures show superior performance in replicating the associated Boltzmann distributions compared to commonly used designs. The findings foster a deeper connection between physical systems and neural network design, paving the way for tailored architectures and providing a physical lens to interpret existing ones.