Multilingual Transfer Learning for QA Using Translation as Data Augmentation

Bornea, Mihaela; Pan, Lin; Rosenthal, Sara; Florian, Radu; Sil, Avirup

Computer Science > Computation and Language

arXiv:2012.05958 (cs)

[Submitted on 10 Dec 2020]

Title:Multilingual Transfer Learning for QA Using Translation as Data Augmentation

Authors:Mihaela Bornea, Lin Pan, Sara Rosenthal, Radu Florian, Avirup Sil

View PDF

Abstract:Prior work on multilingual question answering has mostly focused on using large multilingual pre-trained language models (LM) to perform zero-shot language-wise learning: train a QA model on English and test on other languages. In this work, we explore strategies that improve cross-lingual transfer by bringing the multilingual embeddings closer in the semantic space. Our first strategy augments the original English training data with machine translation-generated data. This results in a corpus of multilingual silver-labeled QA pairs that is 14 times larger than the original training set. In addition, we propose two novel strategies, language adversarial training and language arbitration framework, which significantly improve the (zero-resource) cross-lingual transfer performance and result in LM embeddings that are less language-variant. Empirically, we show that the proposed models outperform the previous zero-shot baseline on the recently introduced multilingual MLQA and TyDiQA datasets.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2012.05958 [cs.CL]
	(or arXiv:2012.05958v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2012.05958
Journal reference:	AAAI 2021

Submission history

From: Mihaela Bornea [view email]
[v1] Thu, 10 Dec 2020 20:29:34 UTC (33 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2020-12

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Lin Pan
Sara Rosenthal
Radu Florian
Avirup Sil

export BibTeX citation

Computer Science > Computation and Language

Title:Multilingual Transfer Learning for QA Using Translation as Data Augmentation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Multilingual Transfer Learning for QA Using Translation as Data Augmentation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators