Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/121719
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Electrical and Electronic Engineeringen_US
dc.creatorHuang, Zen_US
dc.creatorLee, KAen_US
dc.creatorGan, CXen_US
dc.creatorJin, Zen_US
dc.creatorZuo, Ren_US
dc.creatorMak, MWen_US
dc.date.accessioned2026-10-05T07:05:18Z-
dc.date.available2026-10-05T07:05:18Z-
dc.identifier.urihttp://hdl.handle.net/10397/121719-
dc.descriptionInterspeech 2026: Sydney, Australia, 27 September - 1 October 2026en_US
dc.language.isoenen_US
dc.publisherInternational Speech Communication Associationen_US
dc.rightsThe following publication Huang, Z., Lee, K.A., Gan, C.-x., Jin, Z., Zuo, R., Mak, M.-W. (2026) EII-SCL: Harnessing Emotional Inertia for Multimodal Emotion Recognition in Conversation. Proc. Interspeech 2026, 1904-1908. DOI: 10.21437/Interspeech.2026-3532 is available at https://www.isca-archive.org/interspeech_2026/huang26p_interspeech.html.en_US
dc.subjectContrastive learningen_US
dc.subjectEmotional inertiaen_US
dc.subjectEmotion recognition in conversationen_US
dc.subjectMultimodal networken_US
dc.titleEII-SCL : harnessing emotional inertia for multimodal emotion recognition in conversationen_US
dc.typeConference Paperen_US
dc.identifier.spage1904en_US
dc.identifier.epage1908en_US
dc.identifier.doi10.21437/Interspeech.2026-3532en_US
dcterms.abstractMultimodal emotion recognition in conversation (MERC) achieves accurate predictions by integrating multimodal and contextual information in dialogues. While current MERC approaches focus on modeling complex contextual dependencies in conversation, they often overlook the impact of contextual emotional inertia in emotion shift, leading to suboptimal performance. To address this issue, we propose a novel Emotional Inertia-Informed Supervised Contrastive Learning module (EII-SCL) that informs the contrastive objective by constructing inertia-affected samples within temporal windows, effectively leveraging emotional inertia as a prior while enabling seamless integration with existing MERC models without requiring additional data. Extensive experiments on IEMOCAP and MELD show that our approach consistently outperforms state-of-the-art methods.en_US
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationInterspeech 2026: 27 September - 1 October 2026, Sydney, Australia, p. 1904-1908en_US
dcterms.issued2026-
dc.description.validate202610 bcchen_US
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4748-
dc.identifier.SubFormID53846-
dc.description.fundingSourceOthersen_US
dc.description.fundingTextThe work presented in this article is supported by the Research Platform for Advanced Audio and Speech Signal Processing (P0049192) funded by Innovation Technology Co. Ltd.en_US
dc.description.pubStatusPublisheden_US
dc.description.oaCategoryVoR alloweden_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
huang26p_interspeech.pdf1.23 MBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Access
View full-text via PolyU eLinks SFX Query
Show simple item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.