Mixed Precision of Quantization of Transformer Language Models for Speech Recognition

Xu, Junhao; Hu, Shoukang; Yu, Jianwei; Liu, Xunying; Meng, Helen

Computer Science > Computation and Language

arXiv:2112.11540 (cs)

[Submitted on 29 Nov 2021]

Title:Mixed Precision of Quantization of Transformer Language Models for Speech Recognition

Authors:Junhao Xu, Shoukang Hu, Jianwei Yu, Xunying Liu, Helen Meng

View PDF

Abstract:State-of-the-art neural language models represented by Transformers are becoming increasingly complex and expensive for practical applications. Low-bit deep neural network quantization techniques provides a powerful solution to dramatically reduce their model size. Current low-bit quantization methods are based on uniform precision and fail to account for the varying performance sensitivity at different parts of the system to quantization errors. To this end, novel mixed precision DNN quantization methods are proposed in this paper. The optimal local precision settings are automatically learned using two techniques. The first is based on a quantization sensitivity metric in the form of Hessian trace weighted quantization perturbation. The second is based on mixed precision Transformer architecture search. Alternating direction methods of multipliers (ADMM) are used to efficiently train mixed precision quantized DNN systems. Experiments conducted on Penn Treebank (PTB) and a Switchboard corpus trained LF-MMI TDNN system suggest the proposed mixed precision Transformer quantization techniques achieved model size compression ratios of up to 16 times over the full precision baseline with no recognition performance degradation. When being used to compress a larger full precision Transformer LM with more layers, overall word error rate (WER) reductions up to 1.7% absolute (18% relative) were obtained.

Comments:	arXiv admin note: substantial text overlap with arXiv:2112.11438, arXiv:2111.14479
Subjects:	Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2112.11540 [cs.CL]
	(or arXiv:2112.11540v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2112.11540

Submission history

From: Junhao Xu [view email]
[v1] Mon, 29 Nov 2021 09:57:00 UTC (878 KB)

Computer Science > Computation and Language

Title:Mixed Precision of Quantization of Transformer Language Models for Speech Recognition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Mixed Precision of Quantization of Transformer Language Models for Speech Recognition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators