Residual Networks for Resisting Noise: Analysis of an Embeddings-based Spoofing Countermeasure

Bence Halpern, Finnian Kelly, Rob van Son, Anil Alexander

In this paper we propose a spoofing countermeasure based on Constant Q-transform (CQT) features with a ResNet embeddings extractor and a Gaussian Mixture Model (GMM) classifier. We present a detailed analysis of this approach using the Logical Access portion of the ASVspoof2019 evaluation database, and demonstrate that it provides complementary information to the baseline evaluation systems. We additionally evaluate the CQT-ResNet approach in the presence of various types of real noise, and show that it is more robust than the baseline systems. Finally, we explore some explainable audio approaches to offer the human listener insight into the types of information exploited by the network in discriminating spoofed speech from real speech.　

An Explainability Study of the Constant Q Cepstral Coefficient Spoofing Countermeasure for Automatic Speaker Verification

Hemlata Tak, Jose Patino, Andreas Nautsch, Nicholas Evans, Massimiliano Todisco

Odyssey 2020

The Speaker and Language Recognition Workshop

Residual Networks for Resisting Noise: Analysis of an Embeddings-based Spoofing Countermeasure

Search in Audio

Speech Transcript

Related Recordings

Phase Spectrum of Time-flipped Speech Signals for Robust Spoofing Detection

An Explainability Study of the Constant Q Cepstral Coefficient Spoofing Countermeasure for Automatic Speaker Verification