Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/121717
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of Electrical and Electronic Engineering | en_US |
| dc.creator | Huang, Z | en_US |
| dc.creator | Lee, KA | en_US |
| dc.creator | Li, J | en_US |
| dc.creator | Li, Z | en_US |
| dc.creator | Mak, MW | en_US |
| dc.date.accessioned | 2026-10-05T07:00:21Z | - |
| dc.date.available | 2026-10-05T07:00:21Z | - |
| dc.identifier.uri | http://hdl.handle.net/10397/121717 | - |
| dc.description | Interspeech 2026: Sydney, Australia, 27 September - 1 October 2026 | en_US |
| dc.language.iso | en | en_US |
| dc.publisher | International Speech Communication Association | en_US |
| dc.rights | The following publication Huang, Z., Lee, K.A., Li, J., Li, Z., Mak, M.-W. (2026) EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation. Proc. Interspeech 2026, 587-592. DOI: 10.21437/Interspeech.2026-1996 is available at https://www.isca-archive.org/interspeech_2026/huang26n_interspeech.html . | en_US |
| dc.subject | Emotion recognition in conversation | en_US |
| dc.subject | Explicit uncertainty supervision | en_US |
| dc.subject | Multimodal fusion | en_US |
| dc.title | EmoEUS : uncertainty supervision for multimodal emotion recognition in conversation | en_US |
| dc.type | Conference Paper | en_US |
| dc.identifier.spage | 587 | en_US |
| dc.identifier.epage | 592 | en_US |
| dc.identifier.doi | 10.21437/Interspeech.2026-1996 | en_US |
| dcterms.abstract | Multimodal emotion recognition in conversation (MERC) can leverage multimodal and contextual cues to boost recognition performance. However, existing fusion approaches in MERC often ignore modality-specific uncertainty across utterances caused by conflicting cues, varying noise, and missing modality-specific signals. We propose EmoEUS, an explicit uncertainty supervision framework for MERC. EmoEUS performs uncertainty-aware multimodal fusion by dynamically weighting modalities using learned variance estimates. We also introduce an explicitly supervised loss that aligns each utterance's predicted variance with the distance between the utterance's distributional representation and its emotion-and modality-specific cluster center. Experiments on IEMOCAP and MELD show that EmoEUS consistently outperforms state-of-the-art methods. | en_US |
| dcterms.accessRights | open access | en_US |
| dcterms.bibliographicCitation | Interspeech 2026: 27 September - 1 October 2026, Sydney, Australia, p. 587-592 | en_US |
| dcterms.issued | 2026 | - |
| dc.description.validate | 202610 bcch | en_US |
| dc.description.oa | Version of Record | en_US |
| dc.identifier.FolderNumber | a4748 | - |
| dc.identifier.SubFormID | 53847 | - |
| dc.description.fundingSource | Others | en_US |
| dc.description.fundingText | The work presented in this article is supported by the Research Platform for Advanced Audio and Speech Signal Processing (P0049192) funded by Innovation Technology Co. Ltd. | en_US |
| dc.description.pubStatus | Published | en_US |
| dc.description.oaCategory | VoR allowed | en_US |
| Appears in Collections: | Conference Paper | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| huang26n_interspeech.pdf | 944.93 kB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.



