DCAI
← 返回全部动态
arXiv 人工智能规则精选09月25日 12:00

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

arXiv:2609.29050v1 Announce Type: new Abstract: Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens. Consequently, gradient noise from summary generation leaks into tool-decision tokens, causing cross-segment credit misattribution and brittle optimization. In this work, we propose SLCA-GRPO, a framework incorporating Segment-Locked Credit As

阅读 arXiv 人工智能 原文 ↗