Layer-Wise Multi-View Learning for Neural Machine Translation

Wang, Qiang; Li, Changliang; Zhang, Yue; Xiao, Tong; Zhu, Jingbo

Computer Science > Computation and Language

arXiv:2011.01482 (cs)

[Submitted on 3 Nov 2020]

Title:Layer-Wise Multi-View Learning for Neural Machine Translation

Authors:Qiang Wang, Changliang Li, Yue Zhang, Tong Xiao, Jingbo Zhu

View PDF

Abstract:Traditional neural machine translation is limited to the topmost encoder layer's context representation and cannot directly perceive the lower encoder layers. Existing solutions usually rely on the adjustment of network architecture, making the calculation more complicated or introducing additional structural restrictions. In this work, we propose layer-wise multi-view learning to solve this problem, circumventing the necessity to change the model structure. We regard each encoder layer's off-the-shelf output, a by-product in layer-by-layer encoding, as the redundant view for the input sentence. In this way, in addition to the topmost encoder layer (referred to as the primary view), we also incorporate an intermediate encoder layer as the auxiliary view. We feed the two views to a partially shared decoder to maintain independent prediction. Consistency regularization based on KL divergence is used to encourage the two views to learn from each other. Extensive experimental results on five translation tasks show that our approach yields stable improvements over multiple strong baselines. As another bonus, our method is agnostic to network architectures and can maintain the same inference speed as the original model.

Comments:	COLING 2020
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2011.01482 [cs.CL]
	(or arXiv:2011.01482v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2011.01482

Submission history

From: Qiang Wang [view email]
[v1] Tue, 3 Nov 2020 05:06:37 UTC (38 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2020-11

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Qiang Wang
Changliang Li
Yue Zhang
Tong Xiao
Jingbo Zhu

export BibTeX citation

Computer Science > Computation and Language

Title:Layer-Wise Multi-View Learning for Neural Machine Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Layer-Wise Multi-View Learning for Neural Machine Translation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators