Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/118783
PIRA download icon_1.1View/Download Full Text
Title: MindMix : a multimodal foundation model for auditory perception decoding via deep neural-acoustic alignment
Authors: Liu, R 
Chen, Z 
Peng, S 
You, W 
Huang, ZA
Wu, J 
Tan, KC 
Issue Date: 2026
Source: The Fourteenth International Conference on Learning Representations, ICLR 2026, Rio de Janeiro, Brazil, Apr 23 2026, https://openreview.net/forum?id=1ifQzlETeG
Abstract: Decoding complex auditory experiences from non-invasive EEG is a rapidly emerging field that holds significant promise for advancing both fundamental neuroscience and human-machine interaction technologies. Recent developments in EEG foundation models have yielded powerful neural representations that are promising for auditory decoding. However, the effectiveness of these models remains fundamentally constrained by their limited integration with acoustic stimulus information. Specifically, the lack of deep coupling between neural signals and auditory inputs hampers the models’ ability to generalize effectively across diverse auditory tasks. To bridge this gap, we introduce MindMix, a multimodal foundation model designed to bridge the gap between unimodal EEG foundations and task-specific auditory decoders. MindMix employs a two-stage training strategy: first, a high-capacity EEG encoder is pre-trained on over 3,000 hours of EEG data to learn generalized EEG features that can transfer across tasks and subjects. Second, the model learns the neural-acoustic mapping using over 100 hours of paired data, facilitated by our novel Cross-Attention Low-Rank Alignment module, which facilitates fine-grained, cross-modal information integration. Experimental results demonstrate that MindMix substantially surpassing existing baselines across a range of auditory decoding tasks, including auditory attention decoding, auditory emotion recognition, and cross-modal retrieval. This work thus establishes a foundation for future research in multimodal brain decoding and auditory brain-computer interfaces. Our code is available at https://github.com/CookieMikeLiu/MindMix.
Publisher: OpenReview.net
Description: The Fourteenth International Conference on Learning Representations, ICLR 2026, Rio de Janeiro, Brazil, Apr 23 2026
Rights: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
The following publication Liu, R., Chen, Z., You, W., Huang, Z. A., Wu, J., & Tan, K. C. (2026). MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic Alignment. In The Fourteenth International Conference on Learning Representations is available at https://openreview.net/forum?id=1ifQzlETeG.
Appears in Collections:Conference Paper

Files in This Item:
File Description SizeFormat 
9526_MindMix_A_Multimodal_Foun.pdf6.57 MBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Show full item record

Google ScholarTM

Check


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.