On Lyapunov Exponents for RNNs: Understanding Information Propagation Using Dynamical Systems Tools

Ryan Vogt; Maximilian Puelma Touzel; Maximilian Puelma Touzel; Eli Shlizerman; Eli Shlizerman; Guillaume Lajoie; Guillaume Lajoie

doi:10.3389/fams.2022.818799

Frontiers in Applied Mathematics and Statistics (Mar 2022)

On Lyapunov Exponents for RNNs: Understanding Information Propagation Using Dynamical Systems Tools

Ryan Vogt,
Maximilian Puelma Touzel,
Maximilian Puelma Touzel,
Eli Shlizerman,
Eli Shlizerman,
Guillaume Lajoie,
Guillaume Lajoie

Affiliations

Ryan Vogt: Department of Applied Mathematics, University of Washington, Seattle, WA, United States
Maximilian Puelma Touzel: Mathematics and Statistics Department, Université de Montréal, Montréal, QC, Canada
Maximilian Puelma Touzel: Mila - Québec AI Institute,Montréal, QC, Canada
Eli Shlizerman: Department of Applied Mathematics, University of Washington, Seattle, WA, United States
Eli Shlizerman: Department of Electrical and Computer Engineering, University of Washington, Seattle, WA, United States
Guillaume Lajoie: Mathematics and Statistics Department, Université de Montréal, Montréal, QC, Canada
Guillaume Lajoie: Mila - Québec AI Institute,Montréal, QC, Canada

DOI: https://doi.org/10.3389/fams.2022.818799
Journal volume & issue: Vol. 8

Abstract

Read online

Recurrent neural networks (RNNs) have been successfully applied to a variety of problems involving sequential data, but their optimization is sensitive to parameter initialization, architecture, and optimizer hyperparameters. Considering RNNs as dynamical systems, a natural way to capture stability, i.e., the growth and decay over long iterates, are the Lyapunov Exponents (LEs), which form the Lyapunov spectrum. The LEs have a bearing on stability of RNN training dynamics since forward propagation of information is related to the backward propagation of error gradients. LEs measure the asymptotic rates of expansion and contraction of non-linear system trajectories, and generalize stability analysis to the time-varying attractors structuring the non-autonomous dynamics of data-driven RNNs. As a tool to understand and exploit stability of training dynamics, the Lyapunov spectrum fills an existing gap between prescriptive mathematical approaches of limited scope and computationally-expensive empirical approaches. To leverage this tool, we implement an efficient way to compute LEs for RNNs during training, discuss the aspects specific to standard RNN architectures driven by typical sequential datasets, and show that the Lyapunov spectrum can serve as a robust readout of training stability across hyperparameters. With this exposition-oriented contribution, we hope to draw attention to this under-studied, but theoretically grounded tool for understanding training stability in RNNs.

Published in Frontiers in Applied Mathematics and Statistics

ISSN: 2297-4687 (Online)
Publisher: Frontiers Media S.A.
Country of publisher: Switzerland
LCC subjects: Technology: Technology (General): Industrial engineering. Management engineering: Applied mathematics. Quantitative methods; Science: Mathematics: Probabilities. Mathematical statistics
Website: http://journal.frontiersin.org/journal/applied-mathematics-and-statistics#

About the journal

Abstract

Keywords