Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/119865
PIRA download icon_1.1View/Download Full Text
Title: R1-SyntheticVL : is synthetic data from generative models ready for multimodal large language model?
Authors: Zhang, J 
Lin, T 
Yao, H
Lan, X
Liu, S
Huang, J 
Issue Date: 2026
Source: Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea, https://openreview.net/forum?id=t1Ki6pM2EG&referrer=%5Bthe%20profile%20of%20Jingyi%20Zhang%5D%28%2Fprofile%3Fid%3D~Jingyi_Zhang7%29
Abstract: In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks. To this end, we propose Collective Adversarial Data Synthesis (CADS), a novel and general approach to synthesize high-quality, diverse and challenging multimodal data for MLLMs. The core idea of CADS is to leverage collective intelligence to ensure high-quality and diverse generation, while exploring adversarial learning to synthesize challenging samples for effectively driving model improvement. Specifically, CADS operates with two cyclic phases, i.e., Collective Adversarial Data Generation (CAD-Generate) and Collective Adversarial Data Judgment (CAD-Judge). CAD-Generate leverages collective knowledge to jointly generate new and diverse multimodal data, while CAD-Judge collaboratively assesses the quality of synthesized data. In addition, CADS introduces an Adversarial Context Optimization mechanism to optimize the generation context to encourage challenging and high-value data generation. With CADS, we construct MMSynthetic-20K and train our model R1-SyntheticVL, which demonstrates superior performance on various benchmarks.
Description: Forty-Third International Conference on Machine Learning, Seoul, South Korea, July 6th - 11th, 2026
Rights: Copyright 2026 by the author(s).
CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
The following publication Zhang, J., Lin, T., Yao, H., Lan, X., Liu, S., & Huang, J. (2026). R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? is available at https://openreview.net/forum?id=t1Ki6pM2EG&referrer=%5Bthe%20profile%20of%20Jingyi%20Zhang%5D%28%2Fprofile%3Fid%3D~Jingyi_Zhang7%29.
Appears in Collections:Conference Paper

Files in This Item:
File Description SizeFormat 
R1_SyntheticVL_Is_Synthet.pdf2.63 MBAdobe PDFView/Open
Open Access Information
Status open access
File Version Version of Record
Show full item record

Google ScholarTM

Check


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.