Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/120232
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Electrical and Electronic Engineering-
dc.creatorGan, CX-
dc.creatorLi, Z-
dc.creatorJin, Z-
dc.creatorHuang, Z-
dc.creatorMak, MW-
dc.creatorLee, KA-
dc.date.accessioned2026-07-28T00:26:33Z-
dc.date.available2026-07-28T00:26:33Z-
dc.identifier.urihttp://hdl.handle.net/10397/120232-
dc.description26th edition of the Interspeech Conference, August 17-21, 2025, Rotterdam, The Netherlandsen_US
dc.language.isoenen_US
dc.publisherInternational Speech Communication Associationen_US
dc.rightsThe following publication Gan, C.-X., Li, Z., Jin, Z., Huang, Z., Mak, M.-W., Lee, K.A. (2025) IDIR: Identifying and Distilling Informative Relations for Speaker Verification. Proc. Interspeech 2025, 5758-5762 is available at https://doi.org/10.21437/Interspeech.2025-736.en_US
dc.subjectFeature distillationen_US
dc.subjectKnowledge distillationen_US
dc.subjectSpeaker recognitionen_US
dc.titleIDIR : identifying and distilling informative relations for speaker verificationen_US
dc.typeConference Paperen_US
dc.identifier.spage5758-
dc.identifier.epage5762-
dc.identifier.doi10.21437/Interspeech.2025-736-
dcterms.abstractTraditional feature-based knowledge distillation aligns the student’s features with the teacher’s features. However, these one-to-one alignments overlook the structural relations between speakers in a mini-batch. Also, the large capacity gap between the two networks causes significant discrepancies between their features. To address these limitations, we propose distilling the inter- and intra-speaker relations. Instead of mimicking all pairwise relations between the student's and teacher's feature vectors, we propose Identifying and Distilling Informative Relations (IDIR), enabling the student network to acquire speakers' relationships from the teacher. Moreover, a margin is added to the similarity scores of the informative pairs, further reducing intra-speaker variances and increasing inter-speaker separations. Evaluations with a simple x-vector student network demonstrate the method's superb performance across three test sets, showcasing its merits and effectiveness.-
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationIn 26th edition of the Interspeech Conference, to be held August 17-21, 2025, in Rotterdam, The Netherlands, p. 5758-5762-
dcterms.issued2025-
dc.identifier.scopus2-s2.0-105020065762-
dc.relation.ispartofbook26th edition of the Interspeech Conference, to be held August 17-21, 2025, in Rotterdam, The Netherlands-
dc.relation.conferenceConference of the International Speech Communication Association [INTERSPEECH]-
dc.description.validate202607 bcch-
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4743en_US
dc.identifier.SubFormID53840en_US
dc.description.fundingSourceRGCen_US
dc.description.pubStatusPublisheden_US
dc.description.oaCategoryVoR alloweden_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
gan25_interspeech.pdf685.78 kBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Access
View full-text via PolyU eLinks SFX Query
Show simple item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.