Identifying Equivalent Training Dynamics

Redman, William T.; Bello-Rivas, Juan M.; Fonoberova, Maria; Mohr, Ryan; Kevrekidis, Ioannis G.; Mezić, Igor

Computer Science > Machine Learning

arXiv:2302.09160 (cs)

[Submitted on 17 Feb 2023 (v1), last revised 31 Oct 2024 (this version, v3)]

Title:Identifying Equivalent Training Dynamics

Authors:William T. Redman, Juan M. Bello-Rivas, Maria Fonoberova, Ryan Mohr, Ioannis G. Kevrekidis, Igor Mezić

View PDF HTML (experimental)

Abstract:Study of the nonlinear evolution deep neural network (DNN) parameters undergo during training has uncovered regimes of distinct dynamical behavior. While a detailed understanding of these phenomena has the potential to advance improvements in training efficiency and robustness, the lack of methods for identifying when DNN models have equivalent dynamics limits the insight that can be gained from prior work. Topological conjugacy, a notion from dynamical systems theory, provides a precise definition of dynamical equivalence, offering a possible route to address this need. However, topological conjugacies have historically been challenging to compute. By leveraging advances in Koopman operator theory, we develop a framework for identifying conjugate and non-conjugate training dynamics. To validate our approach, we demonstrate that comparing Koopman eigenvalues can correctly identify a known equivalence between online mirror descent and online gradient descent. We then utilize our approach to: (a) identify non-conjugate training dynamics between shallow and wide fully connected neural networks; (b) characterize the early phase of training dynamics in convolutional neural networks; (c) uncover non-conjugate training dynamics in Transformers that do and do not undergo grokking. Our results, across a range of DNN architectures, illustrate the flexibility of our framework and highlight its potential for shedding new light on training dynamics.

Comments:	23 pages, 5 figures, 6 supplemental figures
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Dynamical Systems (math.DS)
Cite as:	arXiv:2302.09160 [cs.LG]
	(or arXiv:2302.09160v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2302.09160

Submission history

From: William T Redman [view email]
[v1] Fri, 17 Feb 2023 22:15:20 UTC (3,504 KB)
[v2] Tue, 4 Jun 2024 15:20:15 UTC (6,632 KB)
[v3] Thu, 31 Oct 2024 13:20:42 UTC (8,428 KB)

Computer Science > Machine Learning

Title:Identifying Equivalent Training Dynamics

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Identifying Equivalent Training Dynamics

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators