DCAI
← 返回全部动态
arXiv 机器学习规则精选09月25日 12:00

RLVR landscapes for iterated multiplications can be benign: Insights from spin-glass theory

arXiv:2609.28625v1 Announce Type: new Abstract: Despite the importance of reinforcement learning with verifiable rewards (RLVR), the extent to which it can learn new reasoning capabilities remains debated. Here we study the optimization landscape of RLVR on algorithmic tasks, such as iterated group and quasigroup multiplication. To this end, we map entropy-regularized RLVR over myopic tabular policies onto an energy-based (spin-glass) model over deterministic policies. This mapping upper-bounds what RLVR can achieve, and lets us rigorously characterize the landscape in this tabular setting. We show, both theoretically and experimentally, that

阅读 arXiv 机器学习 原文 ↗