Investigating Contributions of Speech and Facial Landmarks for Talking Head Generation <BR>(3 minutes introduction)

Investigating Contributions of Speech and Facial Landmarks for Talking Head Generation
(3 minutes introduction)

Ege Kesim (Koç University, Turkey), Engin Erzin (Koç University, Turkey)

Talking head generation is an active research problem. It has been widely studied as a direct speech-to-video or two stage speech-to-landmarks-to-video mapping problem. In this study, our main motivation is to assess individual and joint contributions of the speech and facial landmarks to the talking head generation quality through a state-of-the-art generative adversarial network (GAN) architecture. Incorporating frame and sequence discriminators and a feature matching loss, we investigate performances of speech only, landmark only and joint speech and landmark driven talking head generation on the CREMA-D dataset. Objective evaluations using the peak signal-to-noise ratio (PSNR), structural similarity index (SSIM) and landmark distance (LMD) indicate that while landmarks bring PSNR and SSIM improvements to the speech driven system, speech brings LMD improvement to the landmark driven system. Furthermore, feature matching is observed to improve the speech driven talking head generation models significantly.

Search in Audio

Related Recordings

Cross-lingual Speaker Adaptation using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis
(longer introduction)

Detai Xin , Yuki Saito , Shinnosuke Takamichi , Tomoki Koriyama , Hiroshi Saruwatari

Speech2Video: Cross-Modal Distillation for Speech to Video Generation
(3 minutes introduction)

Shijing Si , Jianzong Wang , Xiaoyang Qu , Ning Cheng , Wenqi Wei , Xinghua Zhu , Jing Xiao

InterSpeech 2021

Investigating Contributions of Speech and Facial Landmarks for Talking Head Generation (3 minutes introduction)

Search in Audio

Related Recordings

Cross-lingual Speaker Adaptation using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis (longer introduction)

Speech2Video: Cross-Modal Distillation for Speech to Video Generation (3 minutes introduction)

Investigating Contributions of Speech and Facial Landmarks for Talking Head Generation
(3 minutes introduction)

Cross-lingual Speaker Adaptation using Domain Adaptation and Speaker Consistency Loss for Text-To-Speech Synthesis
(longer introduction)

Speech2Video: Cross-Modal Distillation for Speech to Video Generation
(3 minutes introduction)