DCAI
← 返回全部动态
arXiv 机器学习规则精选09月25日 12:00

LastOPD: Taming Collapse in Latent On-Policy Distillation

arXiv:2609.28845v1 Announce Type: new Abstract: On-policy distillation (OPD) corrects a student on the responses it writes, but its signal is the teacher's next-token distribution: it tells the student what the teacher says but misses how it thinks. Latent supervision promises the missing part by aligning the student's latent states to the teacher's. Recent methods such as OPRD bring this signal into on-policy distillation. However, we observe two failures of this recipe when distilling Qwen3-4B and Qwen3-8B into Qwen3-1.7B-Base. Early gain, late collapse: latent supervision alone lifts MATH-500 accuracy from 25 to 46 in 10 steps, but subsequ

阅读 arXiv 机器学习 原文 ↗