Stochastic Gradient Descent outperforms Gradient Descent in recovering a high-dimensional signal in a glassy energy landscape

Kamali, Persia Jana; Urbani, Pierfrancesco

Computer Science > Machine Learning

arXiv:2309.04788 (cs)

[Submitted on 9 Sep 2023 (v1), last revised 18 Dec 2023 (this version, v2)]

Title:Stochastic Gradient Descent outperforms Gradient Descent in recovering a high-dimensional signal in a glassy energy landscape

Authors:Persia Jana Kamali, Pierfrancesco Urbani

View PDF HTML (experimental)

Abstract:Stochastic Gradient Descent (SGD) is an out-of-equilibrium algorithm used extensively to train artificial neural networks. However very little is known on to what extent SGD is crucial for to the success of this technology and, in particular, how much it is effective in optimizing high-dimensional non-convex cost functions as compared to other optimization algorithms such as Gradient Descent (GD). In this work we leverage dynamical mean field theory to benchmark its performances in the high-dimensional limit. To do that, we consider the problem of recovering a hidden high-dimensional non-linearly encrypted signal, a prototype high-dimensional non-convex hard optimization problem. We compare the performances of SGD to GD and we show that SGD largely outperforms GD for sufficiently small batch sizes. In particular, a power law fit of the relaxation time of these algorithms shows that the recovery threshold for SGD with small batch size is smaller than the corresponding one of GD.

Comments:	5 pages + appendix. 3 figures
Subjects:	Machine Learning (cs.LG); Disordered Systems and Neural Networks (cond-mat.dis-nn)
Cite as:	arXiv:2309.04788 [cs.LG]
	(or arXiv:2309.04788v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2309.04788

Submission history

From: Pierfrancesco Urbani [view email]
[v1] Sat, 9 Sep 2023 13:29:17 UTC (237 KB)
[v2] Mon, 18 Dec 2023 09:04:57 UTC (238 KB)

Computer Science > Machine Learning

Title:Stochastic Gradient Descent outperforms Gradient Descent in recovering a high-dimensional signal in a glassy energy landscape

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Stochastic Gradient Descent outperforms Gradient Descent in recovering a high-dimensional signal in a glassy energy landscape

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators