Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/121717
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Electrical and Electronic Engineeringen_US
dc.creatorHuang, Zen_US
dc.creatorLee, KAen_US
dc.creatorLi, Jen_US
dc.creatorLi, Zen_US
dc.creatorMak, MWen_US
dc.date.accessioned2026-10-05T07:00:21Z-
dc.date.available2026-10-05T07:00:21Z-
dc.identifier.urihttp://hdl.handle.net/10397/121717-
dc.descriptionInterspeech 2026: Sydney, Australia, 27 September - 1 October 2026en_US
dc.language.isoenen_US
dc.publisherInternational Speech Communication Associationen_US
dc.rightsThe following publication Huang, Z., Lee, K.A., Li, J., Li, Z., Mak, M.-W. (2026) EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation. Proc. Interspeech 2026, 587-592. DOI: 10.21437/Interspeech.2026-1996 is available at https://www.isca-archive.org/interspeech_2026/huang26n_interspeech.html .en_US
dc.subjectEmotion recognition in conversationen_US
dc.subjectExplicit uncertainty supervisionen_US
dc.subjectMultimodal fusionen_US
dc.titleEmoEUS : uncertainty supervision for multimodal emotion recognition in conversationen_US
dc.typeConference Paperen_US
dc.identifier.spage587en_US
dc.identifier.epage592en_US
dc.identifier.doi10.21437/Interspeech.2026-1996en_US
dcterms.abstractMultimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC often ignore modality-specific uncertainty across utterances caused by conflicting cues, varying noise, and missing modality-specific signals. We propose EmoEUS, an explicit uncertainty supervision framework for MERC. EmoEUS performs uncertainty-aware multimodal fusion by dynamically weighting modalities using learned variance estimates. We also introduce an explicitly supervised loss that aligns each utterance's predicted variance with the distance between the utterance's distributional representation and its emotion-and modality-specific cluster center. Experiments on IEMOCAP and MELD show that EmoEUS consistently outperforms state-of-the-art methods.en_US
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationInterspeech 2026: 27 September - 1 October 2026, Sydney, Australia, p. 587-592en_US
dcterms.issued2026-
dc.description.validate202610 bcchen_US
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4748-
dc.identifier.SubFormID53847-
dc.description.fundingSourceOthersen_US
dc.description.fundingTextThe work presented in this article is supported by the Research Platform for Advanced Audio and Speech Signal Processing (P0049192) funded by Innovation Technology Co. Ltd.en_US
dc.description.pubStatusPublisheden_US
dc.description.oaCategoryVoR alloweden_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
huang26n_interspeech.pdf944.93 kBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Access
View full-text via PolyU eLinks SFX Query
Show simple item record

Google ScholarTM

Check

Altmetric


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.