Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/121064
PIRA download icon_1.1View/Download Full Text
DC FieldValueLanguage
dc.contributorDepartment of Data Science and Artificial Intelligenceen_US
dc.creatorZhang, Sen_US
dc.creatorChen, Xen_US
dc.creatorShen, Yen_US
dc.creatorYe, Zen_US
dc.creatorWu, Jen_US
dc.date.accessioned2026-09-14T06:34:34Z-
dc.date.available2026-09-14T06:34:34Z-
dc.identifier.urihttp://hdl.handle.net/10397/121064-
dc.descriptionThe IEEE/CVF Conference on Computer Vision and Pattern Recognition 2026, June 3 - Sun June 7, 2026, Colorado Convention Centeren_US
dc.language.isoenen_US
dc.titleReLaX : reasoning with latent exploration for large reasoning modelsen_US
dc.typeConference Paperen_US
dcterms.abstractReinforcement Learning with Verifiable Rewards (RLVR) has recently demonstrated remarkable potential in enhancing the reasoning capability of Large Reasoning Models (LRMs). However, RLVR often drives the policy toward over-determinism, resulting in ineffective exploration and premature policy convergence. While promoting token-level diversity has shown promise in mitigating entropy collapse, we argue that the latent dynamics underlying token generation encode a far richer computational structure for steering policy optimization toward a more effective exploration–exploitation tradeoff. To enable tractable analysis and intervention of the latent dynamics of LRMs, we leverage Koopman operator theory to obtain a linearized representation of their hidden state dynamics. This enables us to introduce Dynamic Spectral Dispersion (DSD), a new metric to quantify the heterogeneity of the model’s latent dynamics, serving as a direct indicator of policy exploration. Building upon these foundations, we propose Reasoning with Latent eXploration (ReLaX), a framework that explicitly incorporates latent dynamics to regulate exploration and exploitation during policy optimization. Comprehensive experiments across a wide range of multimodal and text-only reasoning benchmarks show that ReLaX consistently incentivizes reasoning capability and outperforms existing token-level methods. Our project is available at https://github.com/ZhangShimin1/ReLaX.en_US
dcterms.accessRightsembargoed accessen_US
dcterms.bibliographicCitationThe IEEE/CVF Conference on Computer Vision and Pattern Recognition 2026, June 3 - June 7, 2026, Colorado Convention Center, https://openaccess.thecvf.com/content/CVPR2026/html/Zhang_ReLaX_Reasoning_with_Latent_Exploration_for_Large_Reasoning_Models_CVPR_2026_paper.htmlen_US
dcterms.issued2026-
dc.relation.conferenceComputer Vision and Pattern Recognition [CVPR]en_US
dc.description.validate202607 bcchen_US
dc.description.oaMetadata onlyen_US
dc.identifier.FolderNumbera4422d-
dc.identifier.SubFormID52765-
dc.description.fundingSourceRGCen_US
dc.description.fundingSourceOthersen_US
dc.description.fundingTextThis work was partially supported by the Research Grants Council of the Hong Kong SAR (Grant No. PolyU25216423, PolyU15217424, and C5052-23G), the National Natural Science Foundation of China (Grant No. 62306259), and The Hong Kong Polytechnic University (P0058445).en_US
dc.date.embargo0000-00-00 (to be updated)en_US
dc.description.oaCategoryGreen (AAM)en_US
Appears in Collections:Conference Paper
Files in This Item:
File Description SizeFormat 
121064_link.htm225 BHTMLView/Open
Open Access Information
Status embargoed access
Embargo End Date 0000-00-00 (to be updated)
Show simple item record

Google ScholarTM

Check


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.