Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/119865
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of Data Science and Artificial Intelligence | en_US |
| dc.creator | Zhang, J | en_US |
| dc.creator | Lin, T | en_US |
| dc.creator | Yao, H | en_US |
| dc.creator | Lan, X | en_US |
| dc.creator | Liu, S | en_US |
| dc.creator | Huang, J | en_US |
| dc.date.accessioned | 2026-07-13T07:11:10Z | - |
| dc.date.available | 2026-07-13T07:11:10Z | - |
| dc.identifier.uri | http://hdl.handle.net/10397/119865 | - |
| dc.description | Forty-Third International Conference on Machine Learning, Seoul, South Korea, July 6th - 11th, 2026 | en_US |
| dc.language.iso | en | en_US |
| dc.rights | Copyright 2026 by the author(s). | en_US |
| dc.rights | CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) | en_US |
| dc.rights | The following publication Zhang, J., Lin, T., Yao, H., Lan, X., Liu, S., & Huang, J. (2026). R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model? is available at https://openreview.net/forum?id=t1Ki6pM2EG&referrer=%5Bthe%20profile%20of%20Jingyi%20Zhang%5D%28%2Fprofile%3Fid%3D~Jingyi_Zhang7%29. | en_US |
| dc.title | R1-SyntheticVL : is synthetic data from generative models ready for multimodal large language model? | en_US |
| dc.type | Conference Paper | en_US |
| dcterms.abstract | In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks. To this end, we propose Collective Adversarial Data Synthesis (CADS), a novel and general approach to synthesize high-quality, diverse and challenging multimodal data for MLLMs. The core idea of CADS is to leverage collective intelligence to ensure high-quality and diverse generation, while exploring adversarial learning to synthesize challenging samples for effectively driving model improvement. Specifically, CADS operates with two cyclic phases, i.e., Collective Adversarial Data Generation (CAD-Generate) and Collective Adversarial Data Judgment (CAD-Judge). CAD-Generate leverages collective knowledge to jointly generate new and diverse multimodal data, while CAD-Judge collaboratively assesses the quality of synthesized data. In addition, CADS introduces an Adversarial Context Optimization mechanism to optimize the generation context to encourage challenging and high-value data generation. With CADS, we construct MMSynthetic-20K and train our model R1-SyntheticVL, which demonstrates superior performance on various benchmarks. | en_US |
| dcterms.accessRights | open access | en_US |
| dcterms.bibliographicCitation | Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea, https://openreview.net/forum?id=t1Ki6pM2EG&referrer=%5Bthe%20profile%20of%20Jingyi%20Zhang%5D%28%2Fprofile%3Fid%3D~Jingyi_Zhang7%29 | en_US |
| dcterms.issued | 2026 | - |
| dc.relation.conference | International Conference on Machine Learning [ICML] | en_US |
| dc.description.validate | 202607 bcch | en_US |
| dc.description.oa | Version of Record | en_US |
| dc.identifier.FolderNumber | a4518 | - |
| dc.identifier.SubFormID | 53022 | - |
| dc.description.fundingSource | RGC | en_US |
| dc.description.pubStatus | Unpublish | en_US |
| dc.description.oaCategory | CC | en_US |
| Appears in Collections: | Conference Paper | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| R1_SyntheticVL_Is_Synthet.pdf | 2.63 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.


