SuperLectures.com

LINGUISTIC INFLUENCES ON BOTTOM-UP AND TOP-DOWN CLUSTERING FOR SPEAKER DIARIZATION

Full Paper at IEEE Xplore

Speaker Diarization

Presented by: Simon Bozonnet, Author(s): Simon Bozonnet, Dong Wang, Nicholas Evans, Raphaël Troncy, EURECOM, France

While bottom-up approaches have emerged as the standard, default approach to clustering for speaker diarization we have always found the top-down approach gives equivalent or superior performance. Our recent work shows that significant gains in performance can be obtained when cluster purification is applied to the output of top-down systems but that it can degrade performance when applied to the output of bottom-up systems. This paper demonstrates that these observations can be accounted for by factors unrelated to the speaker and that they can impact more strongly on the performance of bottom-up clustering strategies than top-down strategies. Experimental results confirm that clusters produced through top-down clustering are better normalized against phone variation than those produced through bottom-up clustering and that this accounts for the observed inconsistencies in purification performance. The work highlights the need for marginalization strategies which should encourage convergence toward different speakers rather than toward nuisance factors such as that those related to the linguistic content.


  Speech Transcript

|

  Slides

Enlarge the slide | Show all slides in a pop-up window

0:00:16

  1. slide

0:00:54

  2. slide

0:01:57

  3. slide

0:02:41

  4. slide

0:03:32

  5. slide

0:04:15

  6. slide

0:05:26

  7. slide

0:05:58

  8. slide

0:06:28

  9. slide

0:07:20

 10. slide

0:08:10

 11. slide

0:08:58

 12. slide

0:09:31

 13. slide

0:09:52

 14. slide

0:10:46

 15. slide

0:12:10

 16. slide

0:12:54

 17. slide

0:15:15

 18. slide

0:16:02

 19. slide

0:17:04

 20. slide

  Comments

Please sign in to post your comment!

  Lecture Information

Recorded: 2011-05-24 14:25 - 14:45, Panorama
Added: 16. 6. 2011 18:59
Number of views: 20
Video resolution: 1024x576 px, 512x288 px
Video length: 0:19:26
Audio track: MP3 [6.57 MB], 0:19:26