Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/120234
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of Electrical and Electronic Engineering | - |
| dc.creator | Li, Z | - |
| dc.creator | Mak, MW | - |
| dc.creator | Chien, JT | - |
| dc.creator | Pilanci, M | - |
| dc.creator | Jin, Z | - |
| dc.creator | Meng, H | - |
| dc.date.accessioned | 2026-07-28T00:26:36Z | - |
| dc.date.available | 2026-07-28T00:26:36Z | - |
| dc.identifier.uri | http://hdl.handle.net/10397/120234 | - |
| dc.description | 26th edition of the Interspeech Conference, August 17-21, 2025, Rotterdam, The Netherlands | en_US |
| dc.language.iso | en | en_US |
| dc.publisher | International Speech Communication Association | en_US |
| dc.rights | The following publication Li, Z., Mak, M.-W., Chien, J.-T., Pilanci, M., Jin, Z., Meng, H. (2025) Disentangling Speaker and Content in Pre-trained Speech Models with Latent Diffusion for Robust Speaker Verification. Proc. Interspeech 2025, 1108-1112 is available at https://doi.org/10.21437/Interspeech.2025-1865. | en_US |
| dc.subject | Diffusion models | en_US |
| dc.subject | Disentanglement | en_US |
| dc.subject | Pre-trained speech models | en_US |
| dc.subject | Speaker verification | en_US |
| dc.subject | VAE | en_US |
| dc.title | Disentangling speaker and content in pre-trained speech models with latent diffusion for robust speaker verification | en_US |
| dc.type | Conference Paper | en_US |
| dc.identifier.spage | 1108 | - |
| dc.identifier.epage | 1112 | - |
| dc.identifier.doi | 10.21437/Interspeech.2025-1865 | - |
| dcterms.abstract | Disentangled speech representation learning for speaker verification aims to separate spoken content and speaker timbre into distinct representations. However, existing variational autoencoder (VAE)--based methods for speech disentanglement rely on latent variables that lack semantic meaning, limiting their effectiveness for speaker verification. To address this limitation, we propose a diffusion-based method that disentangles and separates speaker features and speech content in the latent space. Building upon the VAE framework, we employ a speaker encoder to learn latent variables representing speaker features while using frame-specific latent variables to capture content. Unlike previous sequential VAE approaches, our method utilizes a conditional diffusion model in the latent space to derive speaker-aware representations. Experiments on the VoxCeleb datasets demonstrate that our method effectively isolates speaker features from speech content using pre-trained speech | - |
| dcterms.accessRights | open access | en_US |
| dcterms.bibliographicCitation | In 26th edition of the Interspeech Conference, to be held August 17-21, 2025, in Rotterdam, The Netherlands, p. 1108-1112 | - |
| dcterms.issued | 2025 | - |
| dc.identifier.scopus | 2-s2.0-105020094439 | - |
| dc.relation.ispartofbook | 26th edition of the Interspeech Conference, to be held August 17-21, 2025, in Rotterdam, The Netherlands | - |
| dc.relation.conference | Conference of the International Speech Communication Association [INTERSPEECH] | - |
| dc.description.validate | 202607 bcch | - |
| dc.description.oa | Version of Record | en_US |
| dc.identifier.FolderNumber | a4744 | en_US |
| dc.identifier.SubFormID | 53842 | en_US |
| dc.description.fundingSource | RGC | en_US |
| dc.description.pubStatus | Published | en_US |
| dc.description.oaCategory | VoR allowed | en_US |
| Appears in Collections: | Conference Paper | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| li25z_interspeech.pdf | 1.07 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.



