Research of Damped Newton Stochastic Gradient Descent Method for Neural Network Training

Zhou, Jingcheng; Wei, Wei; Zheng, Zhiming

Computer Science > Machine Learning

arXiv:2103.16764 (cs)

[Submitted on 31 Mar 2021]

Title:Research of Damped Newton Stochastic Gradient Descent Method for Neural Network Training

Authors:Jingcheng Zhou, Wei Wei, Zhiming Zheng

View PDF

Abstract:First-order methods like stochastic gradient descent(SGD) are recently the popular optimization method to train deep neural networks (DNNs), but second-order methods are scarcely used because of the overpriced computing cost in getting the high-order information. In this paper, we propose the Damped Newton Stochastic Gradient Descent(DN-SGD) method and Stochastic Gradient Descent Damped Newton(SGD-DN) method to train DNNs for regression problems with Mean Square Error(MSE) and classification problems with Cross-Entropy Loss(CEL), which is inspired by a proved fact that the hessian matrix of last layer of DNNs is always semi-definite. Different from other second-order methods to estimate the hessian matrix of all parameters, our methods just accurately compute a small part of the parameters, which greatly reduces the computational cost and makes convergence of the learning process much faster and more accurate than SGD. Several numerical experiments on real datesets are performed to verify the effectiveness of our methods for regression and classification problems.

Comments:	10 pages
Subjects:	Machine Learning (cs.LG); Optimization and Control (math.OC)
Cite as:	arXiv:2103.16764 [cs.LG]
	(or arXiv:2103.16764v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2103.16764

Submission history

From: Jingcheng Zhou [view email]
[v1] Wed, 31 Mar 2021 02:07:18 UTC (329 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2021-03

Change to browse by:

cs
math
math.OC

References & Citations

DBLP - CS Bibliography

listing | bibtex

Wei Wei
Zhiming Zheng

export BibTeX citation

Computer Science > Machine Learning

Title:Research of Damped Newton Stochastic Gradient Descent Method for Neural Network Training

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Research of Damped Newton Stochastic Gradient Descent Method for Neural Network Training

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators