Please use this identifier to cite or link to this item:
http://hdl.handle.net/10397/121064
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of Data Science and Artificial Intelligence | en_US |
| dc.creator | Zhang, S | en_US |
| dc.creator | Chen, X | en_US |
| dc.creator | Shen, Y | en_US |
| dc.creator | Ye, Z | en_US |
| dc.creator | Wu, J | en_US |
| dc.date.accessioned | 2026-09-14T06:34:34Z | - |
| dc.date.available | 2026-09-14T06:34:34Z | - |
| dc.identifier.uri | http://hdl.handle.net/10397/121064 | - |
| dc.description | The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2026, June 3 - Sun June 7, 2026, Colorado Convention Center | en_US |
| dc.language.iso | en | en_US |
| dc.title | ReLaX : reasoning with latent exploration for large reasoning models | en_US |
| dc.type | Conference Paper | en_US |
| dcterms.abstract | Reinforcement Learning with Verifiable Rewards (RLVR) has recently demonstrated remarkable potential in enhancing the reasoning capability of Large Reasoning Models (LRMs). However, RLVR often drives the policy toward over-determinism, resulting in ineffective exploration and premature policy convergence. While promoting token-level diversity has shown promise in mitigating entropy collapse, we argue that the latent dynamics underlying token generation encode a far richer computational structure for steering policy optimization toward a more effective exploration–exploitation tradeoff. To enable tractable analysis and intervention of the latent dynamics of LRMs, we leverage Koopman operator theory to obtain a linearized representation of their hidden state dynamics. This enables us to introduce Dynamic Spectral Dispersion (DSD), a new metric to quantify the heterogeneity of the model’s latent dynamics, serving as a direct indicator of policy exploration. Building upon these foundations, we propose Reasoning with Latent eXploration (ReLaX), a framework that explicitly incorporates latent dynamics to regulate exploration and exploitation during policy optimization. Comprehensive experiments across a wide range of multimodal and text-only reasoning benchmarks show that ReLaX consistently incentivizes reasoning capability and outperforms existing token-level methods. Our project is available at https://github.com/ZhangShimin1/ReLaX. | en_US |
| dcterms.accessRights | embargoed access | en_US |
| dcterms.bibliographicCitation | The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2026, June 3 - June 7, 2026, Colorado Convention Center, https://openaccess.thecvf.com/content/CVPR2026/html/Zhang_ReLaX_Reasoning_with_Latent_Exploration_for_Large_Reasoning_Models_CVPR_2026_paper.html | en_US |
| dcterms.issued | 2026 | - |
| dc.relation.conference | Computer Vision and Pattern Recognition [CVPR] | en_US |
| dc.description.validate | 202607 bcch | en_US |
| dc.description.oa | Metadata only | en_US |
| dc.identifier.FolderNumber | a4422d | - |
| dc.identifier.SubFormID | 52765 | - |
| dc.description.fundingSource | RGC | en_US |
| dc.description.fundingSource | Others | en_US |
| dc.description.fundingText | This work was partially supported by the Research Grants Council of the Hong Kong SAR (Grant No. PolyU25216423, PolyU15217424, and C5052-23G), the National Natural Science Foundation of China (Grant No. 62306259), and The Hong Kong Polytechnic University (P0058445). | en_US |
| dc.date.embargo | 0000-00-00 (to be updated) | en_US |
| dc.description.oaCategory | Green (AAM) | en_US |
| Appears in Collections: | Conference Paper | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 121064_link.htm | 225 B | HTML | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.


