Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/120231
| Title: | Asymmetric clean segments-guided self-supervised learning for robust speaker verification | Authors: | Gan, CX Mak, MW Lin, W Chien, JT |
Issue Date: | 2024 | Source: | In 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing: Proceedings, p. 11081-11085 | Abstract: | Contrastive self-supervised learning (CSL) for speaker verification (SV) has drawn increasing interest recently due to its ability to exploit unlabeled data. Performing data augmentation on raw waveforms, such as adding noise or reverberation, plays a pivotal role in achieving promising results in SV. Data augmentation, however, demands meticulous calibration to ensure intact speaker-specific information, which is difficult to achieve without speaker labels. To address this issue, we introduce a novel framework by incorporating clean and augmented segments into the contrastive training pipeline. The clean segments are repurposed to pair with noisy segments to form additional positive and negative pairs. Moreover, the contrastive loss is weighted to increase the difference between the clean and augmented embeddings of different speakers. Experimental results on Voxceleb1 suggest that the proposed framework can achieve a remarkable 19% improvement over the conventional methods, and it surpasses many existing state-of-the-art techniques. | Keywords: | Contrastive learning Hard negative pairs Self-supervised learning Speaker verification Weighted contrastive loss |
Publisher: | Institute of Electrical and Electronics Engineers | ISBN: | 979-8-3503-4485-1 (Electronic) 979-8-3503-4486-8 (Print on Demand(PoD)) |
DOI: | 10.1109/ICASSP48485.2024.10446161 | Description: | ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 14-19 April 2024, COEX, Seoul, Korea | Rights: | © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. The following publication C. -X. Gan, M. -W. Mak, W. Lin and J. -T. Chien, "Asymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker Verification," ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, Republic of, 2024, pp. 11081-11085 is available at https://doi.org/10.1109/ICASSP48485.2024.10446161. |
| Appears in Collections: | Conference Paper |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| Gan_Asymmetric_Clean_Segments.pdf | Pre-Published version | 1.15 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.



