Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/118783
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Data Science and Artificial Intelligenceen_US
dc.creatorLiu, Ren_US
dc.creatorChen, Zen_US
dc.creatorPeng, Sen_US
dc.creatorYou, Wen_US
dc.creatorHuang, ZAen_US
dc.creatorWu, Jen_US
dc.creatorTan, KCen_US
dc.date.accessioned2026-05-19T08:15:33Z-
dc.date.available2026-05-19T08:15:33Z-
dc.identifier.urihttp://hdl.handle.net/10397/118783-
dc.descriptionThe Fourteenth International Conference on Learning Representations, ICLR 2026, Rio de Janeiro, Brazil, Apr 23 2026en_US
dc.language.isoenen_US
dc.publisherOpenReview.neten_US
dc.rightsCC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)en_US
dc.rightsThe following publication Liu, R., Chen, Z., You, W., Huang, Z. A., Wu, J., & Tan, K. C. (2026). MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic Alignment. In The Fourteenth International Conference on Learning Representations is available at https://openreview.net/forum?id=1ifQzlETeG.en_US
dc.titleMindMix : a multimodal foundation model for auditory perception decoding via deep neural-acoustic alignmenten_US
dc.typeConference Paperen_US
dcterms.abstractDecoding complex auditory experiences from non-invasive EEG is a rapidly emerging field that holds significant promise for advancing both fundamental neuroscience and human-machine interaction technologies. Recent developments in EEG foundation models have yielded powerful neural representations that are promising for auditory decoding. However, the effectiveness of these models remains fundamentally constrained by their limited integration with acoustic stimulus information. Specifically, the lack of deep coupling between neural signals and auditory inputs hampers the models’ ability to generalize effectively across diverse auditory tasks. To bridge this gap, we introduce MindMix, a multimodal foundation model designed to bridge the gap between unimodal EEG foundations and task-specific auditory decoders. MindMix employs a two-stage training strategy: first, a high-capacity EEG encoder is pre-trained on over 3,000 hours of EEG data to learn generalized EEG features that can transfer across tasks and subjects. Second, the model learns the neural-acoustic mapping using over 100 hours of paired data, facilitated by our novel Cross-Attention Low-Rank Alignment module, which facilitates fine-grained, cross-modal information integration. Experimental results demonstrate that MindMix substantially surpassing existing baselines across a range of auditory decoding tasks, including auditory attention decoding, auditory emotion recognition, and cross-modal retrieval. This work thus establishes a foundation for future research in multimodal brain decoding and auditory brain-computer interfaces. Our code is available at https://github.com/CookieMikeLiu/MindMix.en_US
dcterms.accessRightsopen accessen_US
dcterms.bibliographicCitationThe Fourteenth International Conference on Learning Representations, ICLR 2026, Rio de Janeiro, Brazil, Apr 23 2026, https://openreview.net/forum?id=1ifQzlETeGen_US
dcterms.issued2026-
dc.relation.conferenceInternational Conference on Learning Representations [ICLR]en_US
dc.description.validate202605 bcchen_US
dc.description.oaVersion of Recorden_US
dc.identifier.FolderNumbera4422a-
dc.identifier.SubFormID52755-
dc.description.fundingSourceRGCen_US
dc.description.fundingSourceOthersen_US
dc.description.fundingTextThis work was partially supported by the National Natural Science Foundation of China (Grant No. 62306259 and 62572413), the Research Grants Council of the Hong Kong SAR (Grant No. C5052-23G, PolyU25216423, PolyU15217424, and SRFS2526-5S04), The Hong Kong Polytechnic University (P0058445).en_US
dc.description.pubStatusUnpublishen_US
dc.description.oaCategoryCCen_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
9526_MindMix_A_Multimodal_Foun.pdf6.57 MBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Show simple item record

Google ScholarTM

Check


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.