An Analysis of Personalized Speech Recognition System Development for the Deaf and Hard-of-Hearing

Violeta, Lester Phillip; Toda, Tomoki

Computer Science > Sound

arXiv:2306.13953 (cs)

[Submitted on 24 Jun 2023]

Title:An Analysis of Personalized Speech Recognition System Development for the Deaf and Hard-of-Hearing

Authors:Lester Phillip Violeta, Tomoki Toda

View PDF

Abstract:Deaf or hard-of-hearing (DHH) speakers typically have atypical speech caused by deafness. With the growing support of speech-based devices and software applications, more work needs to be done to make these devices inclusive to everyone. To do so, we analyze the use of openly-available automatic speech recognition (ASR) tools with a DHH Japanese speaker dataset. As these out-of-the-box ASR models typically do not perform well on DHH speech, we provide a thorough analysis of creating personalized ASR systems. We collected a large DHH speaker dataset of four speakers totaling around 28.05 hours and thoroughly analyzed the performance of different training frameworks by varying the training data sizes. Our findings show that 1000 utterances (or 1-2 hours) from a target speaker can already significantly improve the model performance with minimal amount of work needed, thus we recommend researchers to collect at least 1000 utterances to make an efficient personalized ASR system. In cases where 1000 utterances is difficult to collect, we also discover significant improvements in using previously proposed data augmentation techniques such as intermediate fine-tuning when only 200 utterances are available.

Comments:	Submitted to APSIPA 2023
Subjects:	Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2306.13953 [cs.SD]
	(or arXiv:2306.13953v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2306.13953

Submission history

From: Lester Phillip Violeta [view email]
[v1] Sat, 24 Jun 2023 12:49:49 UTC (198 KB)

Computer Science > Sound

Title:An Analysis of Personalized Speech Recognition System Development for the Deaf and Hard-of-Hearing

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:An Analysis of Personalized Speech Recognition System Development for the Deaf and Hard-of-Hearing

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators