Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/118783
| Title: | MindMix : a multimodal foundation model for auditory perception decoding via deep neural-acoustic alignment | Authors: | Liu, R Chen, Z Peng, S You, W Huang, ZA Wu, J Tan, KC |
Issue Date: | 2026 | Source: | The Fourteenth International Conference on Learning Representations, ICLR 2026, Rio de Janeiro, Brazil, Apr 23 2026, https://openreview.net/forum?id=1ifQzlETeG | Abstract: | Decoding complex auditory experiences from non-invasive EEG is a rapidly emerging field that holds significant promise for advancing both fundamental neuroscience and human-machine interaction technologies. Recent developments in EEG foundation models have yielded powerful neural representations that are promising for auditory decoding. However, the effectiveness of these models remains fundamentally constrained by their limited integration with acoustic stimulus information. Specifically, the lack of deep coupling between neural signals and auditory inputs hampers the models’ ability to generalize effectively across diverse auditory tasks. To bridge this gap, we introduce MindMix, a multimodal foundation model designed to bridge the gap between unimodal EEG foundations and task-specific auditory decoders. MindMix employs a two-stage training strategy: first, a high-capacity EEG encoder is pre-trained on over 3,000 hours of EEG data to learn generalized EEG features that can transfer across tasks and subjects. Second, the model learns the neural-acoustic mapping using over 100 hours of paired data, facilitated by our novel Cross-Attention Low-Rank Alignment module, which facilitates fine-grained, cross-modal information integration. Experimental results demonstrate that MindMix substantially surpassing existing baselines across a range of auditory decoding tasks, including auditory attention decoding, auditory emotion recognition, and cross-modal retrieval. This work thus establishes a foundation for future research in multimodal brain decoding and auditory brain-computer interfaces. Our code is available at https://github.com/CookieMikeLiu/MindMix. | Publisher: | OpenReview.net | Description: | The Fourteenth International Conference on Learning Representations, ICLR 2026, Rio de Janeiro, Brazil, Apr 23 2026 | Rights: | CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) The following publication Liu, R., Chen, Z., You, W., Huang, Z. A., Wu, J., & Tan, K. C. (2026). MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic Alignment. In The Fourteenth International Conference on Learning Representations is available at https://openreview.net/forum?id=1ifQzlETeG. |
| Appears in Collections: | Conference Paper |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 9526_MindMix_A_Multimodal_Foun.pdf | 6.57 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.


