Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/120234
PIRA download icon_1.1View/Download Full Text
Title: Disentangling speaker and content in pre-trained speech models with latent diffusion for robust speaker verification
Authors: Li, Z 
Mak, MW 
Chien, JT
Pilanci, M
Jin, Z 
Meng, H
Issue Date: 2025
Source: In 26th edition of the Interspeech Conference, to be held August 17-21, 2025, in Rotterdam, The Netherlands, p. 1108-1112
Abstract: Disentangled speech representation learning for speaker verification aims to separate spoken content and speaker timbre into distinct representations. However, existing variational autoencoder (VAE)--based methods for speech disentanglement rely on latent variables that lack semantic meaning, limiting their effectiveness for speaker verification. To address this limitation, we propose a diffusion-based method that disentangles and separates speaker features and speech content in the latent space. Building upon the VAE framework, we employ a speaker encoder to learn latent variables representing speaker features while using frame-specific latent variables to capture content. Unlike previous sequential VAE approaches, our method utilizes a conditional diffusion model in the latent space to derive speaker-aware representations. Experiments on the VoxCeleb datasets demonstrate that our method effectively isolates speaker features from speech content using pre-trained speech
Keywords: Diffusion models
Disentanglement
Pre-trained speech models
Speaker verification
VAE
Publisher: International Speech Communication Association
DOI: 10.21437/Interspeech.2025-1865
Description: 26th edition of the Interspeech Conference, August 17-21, 2025, Rotterdam, The Netherlands
Rights: The following publication Li, Z., Mak, M.-W., Chien, J.-T., Pilanci, M., Jin, Z., Meng, H. (2025) Disentangling Speaker and Content in Pre-trained Speech Models with Latent Diffusion for Robust Speaker Verification. Proc. Interspeech 2025, 1108-1112 is available at https://doi.org/10.21437/Interspeech.2025-1865.
Appears in Collections:Conference Paper

Files in This Item:
File Description SizeFormat 
li25z_interspeech.pdf1.07 MBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Access
View full-text via PolyU eLinks SFX Query
Show full item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.