Lipper: Synthesizing Thy Speech using Multi-View Lipreading

Kumar, Yaman; Jain, Rohit; Salik, Khwaja Mohd.; Shah, Rajiv Ratn; yin, Yifang; Zimmermann, Roger

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:1907.01367 (eess)

[Submitted on 28 Jun 2019]

Title:Lipper: Synthesizing Thy Speech using Multi-View Lipreading

Authors:Yaman Kumar, Rohit Jain, Khwaja Mohd. Salik, Rajiv Ratn Shah, Yifang yin, Roger Zimmermann

View PDF

Abstract:Lipreading has a lot of potential applications such as in the domain of surveillance and video conferencing. Despite this, most of the work in building lipreading systems has been limited to classifying silent videos into classes representing text phrases. However, there are multiple problems associated with making lipreading a text-based classification task like its dependence on a particular language and vocabulary mapping. Thus, in this paper we propose a multi-view lipreading to audio system, namely Lipper, which models it as a regression task. The model takes silent videos as input and produces speech as the output. With multi-view silent videos, we observe an improvement over single-view speech reconstruction results. We show this by presenting an exhaustive set of experiments for speaker-dependent, out-of-vocabulary and speaker-independent settings. Further, we compare the delay values of Lipper with other speechreading systems in order to show the real-time nature of audio produced. We also perform a user study for the audios produced in order to understand the level of comprehensibility of audios produced using Lipper.

Comments:	Accepted at AAAI 2019
Subjects:	Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Machine Learning (stat.ML)
Cite as:	arXiv:1907.01367 [eess.AS]
	(or arXiv:1907.01367v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.1907.01367

Submission history

From: Rajiv Ratn Shah [view email]
[v1] Fri, 28 Jun 2019 10:26:23 UTC (990 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Lipper: Synthesizing Thy Speech using Multi-View Lipreading

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Lipper: Synthesizing Thy Speech using Multi-View Lipreading

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators