Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks <BR>(3 minutes introduction)

Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks
(3 minutes introduction)

Herman Kamper (Stellenbosch University, South Africa), Benjamin van Niekerk (Stellenbosch University, South Africa)

We investigate segmenting and clustering speech into low-bitrate phone-like sequences without supervision. We specifically constrain pretrained self-supervised vector-quantized (VQ) neural networks so that blocks of contiguous feature vectors are assigned to the same code, thereby giving a variable-rate segmentation of the speech into discrete units. Two segmentation methods are considered. In the first, features are greedily merged until a prespecified number of segments are reached. The second uses dynamic programming to optimize a squared error with a penalty term to encourage fewer but longer segments. We show that these VQ segmentation methods can be used without alteration across a wide range of tasks: unsupervised phone segmentation, ABX phone discrimination, same-different word discrimination, and as inputs to a symbolic word segmentation algorithm. The penalized dynamic programming method generally performs best. While performance on individual tasks is only comparable to the state-of-the-art in some cases, in all tasks a reasonable competing approach is outperformed at a substantially lower bitrate.

Search in Audio

Related Recordings

Speech SimCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
(3 minutes introduction)

Dongwei Jiang , Wubo Li , Miao Cao , Wei Zou , Xiangang Li

Speech SimCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
(longer introduction)

Dongwei Jiang , Wubo Li , Miao Cao , Wei Zou , Xiangang Li

InterSpeech 2021

Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks (3 minutes introduction)

Search in Audio

Related Recordings

Speech SimCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning (3 minutes introduction)

Speech SimCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning (longer introduction)

Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks
(3 minutes introduction)

Speech SimCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
(3 minutes introduction)

Speech SimCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
(longer introduction)