DCAI
← 返回全部动态
arXiv 自然语言处理规则精选09月24日 12:00

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

arXiv:2609.27373v1 Announce Type: new Abstract: Recurrent language models repeatedly apply shared network blocks to refine latent representations, but standard inference recomputes global attention at every recurrent step. We study attention dynamics across recurrent depth and find that attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests a two-stage structure: early steps discover a sparse working set of relevant context, while later steps refine representations over largely the same routing support. Motivated by this structure, we introduce WISE (Working-set Inference wi

阅读 arXiv 自然语言处理 原文 ↗