站内快照 · 国内可打开。外网原文可能无法访问。
- 论文公开站arXiv
Decoupling Exploration from Optimization in RLVR
Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In pra