Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/118783
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of Data Science and Artificial Intelligence | en_US |
| dc.creator | Liu, R | en_US |
| dc.creator | Chen, Z | en_US |
| dc.creator | Peng, S | en_US |
| dc.creator | You, W | en_US |
| dc.creator | Huang, ZA | en_US |
| dc.creator | Wu, J | en_US |
| dc.creator | Tan, KC | en_US |
| dc.date.accessioned | 2026-05-19T08:15:33Z | - |
| dc.date.available | 2026-05-19T08:15:33Z | - |
| dc.identifier.uri | http://hdl.handle.net/10397/118783 | - |
| dc.description | The Fourteenth International Conference on Learning Representations, ICLR 2026, Rio de Janeiro, Brazil, Apr 23 2026 | en_US |
| dc.language.iso | en | en_US |
| dc.publisher | OpenReview.net | en_US |
| dc.rights | CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) | en_US |
| dc.rights | The following publication Liu, R., Chen, Z., You, W., Huang, Z. A., Wu, J., & Tan, K. C. (2026). MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic Alignment. In The Fourteenth International Conference on Learning Representations is available at https://openreview.net/forum?id=1ifQzlETeG. | en_US |
| dc.title | MindMix : a multimodal foundation model for auditory perception decoding via deep neural-acoustic alignment | en_US |
| dc.type | Conference Paper | en_US |
| dcterms.abstract | Decoding complex auditory experiences from non-invasive EEG is a rapidly emerging field that holds significant promise for advancing both fundamental neuroscience and human-machine interaction technologies. Recent developments in EEG foundation models have yielded powerful neural representations that are promising for auditory decoding. However, the effectiveness of these models remains fundamentally constrained by their limited integration with acoustic stimulus information. Specifically, the lack of deep coupling between neural signals and auditory inputs hampers the models’ ability to generalize effectively across diverse auditory tasks. To bridge this gap, we introduce MindMix, a multimodal foundation model designed to bridge the gap between unimodal EEG foundations and task-specific auditory decoders. MindMix employs a two-stage training strategy: first, a high-capacity EEG encoder is pre-trained on over 3,000 hours of EEG data to learn generalized EEG features that can transfer across tasks and subjects. Second, the model learns the neural-acoustic mapping using over 100 hours of paired data, facilitated by our novel Cross-Attention Low-Rank Alignment module, which facilitates fine-grained, cross-modal information integration. Experimental results demonstrate that MindMix substantially surpassing existing baselines across a range of auditory decoding tasks, including auditory attention decoding, auditory emotion recognition, and cross-modal retrieval. This work thus establishes a foundation for future research in multimodal brain decoding and auditory brain-computer interfaces. Our code is available at https://github.com/CookieMikeLiu/MindMix. | en_US |
| dcterms.accessRights | open access | en_US |
| dcterms.bibliographicCitation | The Fourteenth International Conference on Learning Representations, ICLR 2026, Rio de Janeiro, Brazil, Apr 23 2026, https://openreview.net/forum?id=1ifQzlETeG | en_US |
| dcterms.issued | 2026 | - |
| dc.relation.conference | International Conference on Learning Representations [ICLR] | en_US |
| dc.description.validate | 202605 bcch | en_US |
| dc.description.oa | Version of Record | en_US |
| dc.identifier.FolderNumber | a4422a | - |
| dc.identifier.SubFormID | 52755 | - |
| dc.description.fundingSource | RGC | en_US |
| dc.description.fundingSource | Others | en_US |
| dc.description.fundingText | This work was partially supported by the National Natural Science Foundation of China (Grant No. 62306259 and 62572413), the Research Grants Council of the Hong Kong SAR (Grant No. C5052-23G, PolyU25216423, PolyU15217424, and SRFS2526-5S04), The Hong Kong Polytechnic University (P0058445). | en_US |
| dc.description.pubStatus | Unpublish | en_US |
| dc.description.oaCategory | CC | en_US |
| Appears in Collections: | Conference Paper | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 9526_MindMix_A_Multimodal_Foun.pdf | 6.57 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.


