Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/120234
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Electrical and Electronic Engineering-
dc.creatorLi, Z-
dc.creatorMak, MW-
dc.creatorChien, JT-
dc.creatorPilanci, M-
dc.creatorJin, Z-
dc.creatorMeng, H-
dc.date.accessioned2026-07-28T00:26:36Z-
dc.date.available2026-07-28T00:26:36Z-
dc.identifier.urihttp://hdl.handle.net/10397/120234-
dc.description26th edition of the Interspeech Conference, August 17-21, 2025, Rotterdam, The Netherlandsen_US
dc.language.isoenen_US
dc.publisherInternational Speech Communication Associationen_US
dc.rightsThe following publication Li, Z., Mak, M.-W., Chien, J.-T., Pilanci, M., Jin, Z., Meng, H. (2025) Disentangling Speaker and Content in Pre-trained Speech Models with Latent Diffusion for Robust Speaker Verification. Proc. Interspeech 2025, 1108-1112 is available at https://doi.org/10.21437/Interspeech.2025-1865.en_US
dc.subjectDiffusion modelsen_US
dc.subjectDisentanglementen_US
dc.subjectPre-trained speech modelsen_US
dc.subjectSpeaker verificationen_US
dc.subjectVAEen_US
dc.titleDisentangling speaker and content in pre-trained speech models with latent diffusion for robust speaker verificationen_US
dc.typeConference Paperen_US
dc.identifier.spage1108-
dc.identifier.epage1112-
dc.identifier.doi10.21437/Interspeech.2025-1865-
dcterms.abstractDisentangled speech representation learning for speaker verification aims to separate spoken content and speaker timbre into distinct representations. However, existing variational autoencoder (VAE)--based methods for speech disentanglement rely on latent variables that lack semantic meaning, limiting their effectiveness for speaker verification. To address this limitation, we propose a diffusion-based method that disentangles and separates speaker features and speech content in the latent space. Building upon the VAE framework, we employ a speaker encoder to learn latent variables representing speaker features while using frame-specific latent variables to capture content. Unlike previous sequential VAE approaches, our method utilizes a conditional diffusion model in the latent space to derive speaker-aware representations. Experiments on the VoxCeleb datasets demonstrate that our method effectively isolates speaker features from speech content using pre-trained speech-
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationIn 26th edition of the Interspeech Conference, to be held August 17-21, 2025, in Rotterdam, The Netherlands, p. 1108-1112-
dcterms.issued2025-
dc.identifier.scopus2-s2.0-105020094439-
dc.relation.ispartofbook26th edition of the Interspeech Conference, to be held August 17-21, 2025, in Rotterdam, The Netherlands-
dc.relation.conferenceConference of the International Speech Communication Association [INTERSPEECH]-
dc.description.validate202607 bcch-
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4744en_US
dc.identifier.SubFormID53842en_US
dc.description.fundingSourceRGCen_US
dc.description.pubStatusPublisheden_US
dc.description.oaCategoryVoR alloweden_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
li25z_interspeech.pdf1.07 MBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Access
View full-text via PolyU eLinks SFX Query
Show simple item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.