Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/120234
| Title: | Disentangling speaker and content in pre-trained speech models with latent diffusion for robust speaker verification | Authors: | Li, Z Mak, MW Chien, JT Pilanci, M Jin, Z Meng, H |
Issue Date: | 2025 | Source: | In 26th edition of the Interspeech Conference, to be held August 17-21, 2025, in Rotterdam, The Netherlands, p. 1108-1112 | Abstract: | Disentangled speech representation learning for speaker verification aims to separate spoken content and speaker timbre into distinct representations. However, existing variational autoencoder (VAE)--based methods for speech disentanglement rely on latent variables that lack semantic meaning, limiting their effectiveness for speaker verification. To address this limitation, we propose a diffusion-based method that disentangles and separates speaker features and speech content in the latent space. Building upon the VAE framework, we employ a speaker encoder to learn latent variables representing speaker features while using frame-specific latent variables to capture content. Unlike previous sequential VAE approaches, our method utilizes a conditional diffusion model in the latent space to derive speaker-aware representations. Experiments on the VoxCeleb datasets demonstrate that our method effectively isolates speaker features from speech content using pre-trained speech | Keywords: | Diffusion models Disentanglement Pre-trained speech models Speaker verification VAE |
Publisher: | International Speech Communication Association | DOI: | 10.21437/Interspeech.2025-1865 | Description: | 26th edition of the Interspeech Conference, August 17-21, 2025, Rotterdam, The Netherlands | Rights: | The following publication Li, Z., Mak, M.-W., Chien, J.-T., Pilanci, M., Jin, Z., Meng, H. (2025) Disentangling Speaker and Content in Pre-trained Speech Models with Latent Diffusion for Robust Speaker Verification. Proc. Interspeech 2025, 1108-1112 is available at https://doi.org/10.21437/Interspeech.2025-1865. |
| Appears in Collections: | Conference Paper |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| li25z_interspeech.pdf | 1.07 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.



