Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/119865
| Title: | R1-SyntheticVL : is synthetic data from generative models ready for multimodal large language model? | Authors: | Zhang, J Lin, T Yao, H Lan, X Liu, S Huang, J |
Issue Date: | 2026 | Source: | Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea, https://openreview.net/forum?id=t1Ki6pM2EG&referrer=%5Bthe%20profile%20of%20Jingyi%20Zhang%5D%28%2Fprofile%3Fid%3D~Jingyi_Zhang7%29 | Abstract: | In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks. To this end, we propose Collective Adversarial Data Synthesis (CADS), a novel and general approach to synthesize high-quality, diverse and challenging multimodal data for MLLMs. The core idea of CADS is to leverage collective intelligence to ensure high-quality and diverse generation, while exploring adversarial learning to synthesize challenging samples for effectively driving model improvement. Specifically, CADS operates with two cyclic phases, i.e., Collective Adversarial Data Generation (CAD-Generate) and Collective Adversarial Data Judgment (CAD-Judge). CAD-Generate leverages collective knowledge to jointly generate new and diverse multimodal data, while CAD-Judge collaboratively assesses the quality of synthesized data. In addition, CADS introduces an Adversarial Context Optimization mechanism to optimize the generation context to encourage challenging and high-value data generation. With CADS, we construct MMSynthetic-20K and train our model R1-SyntheticVL, which demonstrates superior performance on various benchmarks. | Description: | Forty-Third International Conference on Machine Learning, Seoul, South Korea, July 6th - 11th, 2026 | Rights: | Copyright 2026 by the author(s). CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) The following publication Zhang, J., Lin, T., Yao, H., Lan, X., Liu, S., & Huang, J. (2026). R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? is available at https://openreview.net/forum?id=t1Ki6pM2EG&referrer=%5Bthe%20profile%20of%20Jingyi%20Zhang%5D%28%2Fprofile%3Fid%3D~Jingyi_Zhang7%29. |
| Appears in Collections: | Conference Paper |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| R1_SyntheticVL_Is_Synthet.pdf | 2.63 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.


