Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation

Lashkarashvili, Nineli; Wu, Wen; Sun, Guangzhi; Woodland, Philip C.

doi:10.1109/ICASSP48485.2024.10446272

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2402.11747 (eess)

[Submitted on 19 Feb 2024]

Title:Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation

Authors:Nineli Lashkarashvili, Wen Wu, Guangzhi Sun, Philip C. Woodland

View PDF

Abstract:Foundation models have shown superior performance for speech emotion recognition (SER). However, given the limited data in emotion corpora, finetuning all parameters of large pre-trained models for SER can be both resource-intensive and susceptible to overfitting. This paper investigates parameter-efficient finetuning (PEFT) for SER. Various PEFT adaptors are systematically studied for both classification of discrete emotion categories and prediction of dimensional emotional attributes. The results demonstrate that the combination of PEFT methods surpasses full finetuning with a significant reduction in the number of trainable parameters. Furthermore, a two-stage adaptation strategy is proposed to adapt models trained on acted emotion data, which is more readily available, to make the model more adept at capturing natural emotional expressions. Both intra- and cross-corpus experiments validate the efficacy of the proposed approach in enhancing the performance on both the source and target domains.

Subjects:	Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
Cite as:	arXiv:2402.11747 [eess.AS]
	(or arXiv:2402.11747v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2402.11747
Journal reference:	ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, Republic of, 2024, pp. 10986-10990
Related DOI:	https://doi.org/10.1109/ICASSP48485.2024.10446272

Submission history

From: Nineli Lashkarashvili [view email]
[v1] Mon, 19 Feb 2024 00:21:07 UTC (572 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators