DCAI
← 返回全部动态
arXiv 人工智能规则精选09月25日 12:00

From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents

arXiv:2609.29051v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has become a popular recipe for post-training LLM agents. It supervises the agent model at the token level with a stronger teacher view of the same model, obtained by conditioning on privileged information (PI). In this work, we show that in multi-turn agents, this paradigm teaches the student to act with confidence but without the information behind it. The trained agent behaves as if it had privileged information it never observed, and its performance falls well short of plain RL, in the worst case below the untrained base model. Therefore, we propose Privile

阅读 arXiv 人工智能 原文 ↗