Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/120231
PIRA download icon_1.1View/Download Full Text
Title: Asymmetric clean segments-guided self-supervised learning for robust speaker verification
Authors: Gan, CX 
Mak, MW 
Lin, W 
Chien, JT
Issue Date: 2024
Source: In 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing: Proceedings, p. 11081-11085
Abstract: Contrastive self-supervised learning (CSL) for speaker verification (SV) has drawn increasing interest recently due to its ability to exploit unlabeled data. Performing data augmentation on raw waveforms, such as adding noise or reverberation, plays a pivotal role in achieving promising results in SV. Data augmentation, however, demands meticulous calibration to ensure intact speaker-specific information, which is difficult to achieve without speaker labels. To address this issue, we introduce a novel framework by incorporating clean and augmented segments into the contrastive training pipeline. The clean segments are repurposed to pair with noisy segments to form additional positive and negative pairs. Moreover, the contrastive loss is weighted to increase the difference between the clean and augmented embeddings of different speakers. Experimental results on Voxceleb1 suggest that the proposed framework can achieve a remarkable 19% improvement over the conventional methods, and it surpasses many existing state-of-the-art techniques.
Keywords: Contrastive learning
Hard negative pairs
Self-supervised learning
Speaker verification
Weighted contrastive loss
Publisher: Institute of Electrical and Electronics Engineers
ISBN: 979-8-3503-4485-1 (Electronic)
979-8-3503-4486-8 (Print on Demand(PoD))
DOI: 10.1109/ICASSP48485.2024.10446161
Description: ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 14-19 April 2024, COEX, Seoul, Korea
Rights: © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
The following publication C. -X. Gan, M. -W. Mak, W. Lin and J. -T. Chien, "Asymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker Verification," ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, Republic of, 2024, pp. 11081-11085 is available at https://doi.org/10.1109/ICASSP48485.2024.10446161.
Appears in Collections:Conference Paper

Files in This Item:
File Description SizeFormat 
Gan_Asymmetric_Clean_Segments.pdfPre-Published version1.15 MBAdobe PDFView/Open
Open Access Information
Status open access
File Version Final Accepted Manuscript
Access
View full-text via PolyU eLinks SFX Query
Show full item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.