Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/120229
PIRA download icon_1.1View/Download Full Text
Title: Distilling attention knowledge for speaker verification
Authors: Jin, Z 
Liu, S
Li, Z
Gan, CX 
Huang, Z 
Mak, MW 
Issue Date: 2026
Source: In CASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP): Proceedings, p. 16447-16451
Abstract: In recent years, the adoption of pre-trained speech models as feature extractors for speaker verification (SV) has surged. To reduce model complexity, knowledge from these pre-trained models has been transferred to compact student models, enabling them to achieve performance unattainable by conventional methods. While label-level and feature-level knowledge distillation (KD) have both yielded promising results in SV, attention-level KD remains under explored. Different from the other two methods, attention-level KD aims to guide the student network to focus more on the informative and important parts of the feature. To explore the attention-level KD in SV tasks, in this paper, we develop two distinct attention maps that enable the student model to learn the teacher’s attention along both the frequency and temporal dimensions, named Frequency-attentive KD (FREQ-AKD) and Temporal-attentive KD (TEMPO-AKD), respectively. Experiments conducted on the VoxCeleb and CN-Celeb datasets demonstrate the effectiveness of both FREQ-AKD and TEMPO-AKD.
Keywords: Attention maps
Knowledge distillation
Speaker verification
Publisher: Institute of Electrical and Electronics Engineers
ISBN: 979-8-3315-6701-9 (Electronic)
979-8-3315-6702-6 (Print on Demand(PoD))
DOI: 10.1109/ICASSP55912.2026.11464971
Description: ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4-8 May 2026, Barcelona, Spain
Rights: © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
The following publication Z. Jin et al., "Distilling Attention Knowledge for Speaker Verification," ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2026, pp. 16447-16451 is available at https://doi.org/10.1109/ICASSP55912.2026.11464971.
Appears in Collections:Conference Paper

Files in This Item:
File Description SizeFormat 
Jin_Distilling_Attention_Knowledge.pdfPre-Published version1.22 MBAdobe PDFView/Open
Open Access Information
Status open access
File Version Final Accepted Manuscript
Access
View full-text via PolyU eLinks SFX Query
Show full item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.