Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models
arXiv:2609.27373v1 Announce Type: new Abstract: Recurrent language models repeatedly apply shared network blocks to refine latent representations, but standard inference recomputes global attention at every recurrent step. We study attention dynamics across recurrent depth and find that attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests a two-stage structure: early steps discover a sparse working set of relevant context, while later steps refine representations over largely the same routing support. Motivated by this structure, we introduce WISE (Working-set Inference wi
阅读 arXiv 自然语言处理 原文 ↗