Understanding Dropout as an Optimization Trick

Hahn, Sangchul; Choi, Heeyoul

Computer Science > Machine Learning

arXiv:1806.09783 (cs)

[Submitted on 26 Jun 2018 (v1), last revised 9 Oct 2019 (this version, v3)]

Title:Understanding Dropout as an Optimization Trick

Authors:Sangchul Hahn, Heeyoul Choi

View PDF

Abstract:As one of standard approaches to train deep neural networks, dropout has been applied to regularize large models to avoid overfitting, and the improvement in performance by dropout has been explained as avoiding co-adaptation between nodes. However, when correlations between nodes are compared after training the networks with or without dropout, one question arises if co-adaptation avoidance explains the dropout effect completely. In this paper, we propose an additional explanation of why dropout works and propose a new technique to design better activation functions. First, we show that dropout can be explained as an optimization technique to push the input towards the saturation area of nonlinear activation function by accelerating gradient information flowing even in the saturation area in backpropagation. Based on this explanation, we propose a new technique for activation functions, {\em gradient acceleration in activation function (GAAF)}, that accelerates gradients to flow even in the saturation area. Then, input to the activation function can climb onto the saturation area which makes the network more robust because the model converges on a flat region. Experiment results support our explanation of dropout and confirm that the proposed GAAF technique improves image classification performance with expected properties.

Comments:	16 pages
Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1806.09783 [cs.LG]
	(or arXiv:1806.09783v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1806.09783

Submission history

From: Heeyoul Choi [view email]
[v1] Tue, 26 Jun 2018 03:43:17 UTC (1,162 KB)
[v2] Sun, 7 Apr 2019 09:11:43 UTC (2,018 KB)
[v3] Wed, 9 Oct 2019 06:27:52 UTC (2,017 KB)

Computer Science > Machine Learning

Title:Understanding Dropout as an Optimization Trick

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Understanding Dropout as an Optimization Trick

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators