DCAI
← 返回全部动态
arXiv 机器学习规则精选09月25日 12:00

Thinking Leakage: A Causal Audit of NoThink Post-Training in Hybrid Reasoning Models

arXiv:2609.28682v1 Announce Type: new Abstract: Post-training hybrid reasoning models in NoThink mode has attracted growing interest as a way to improve performance while keeping inference fast. However, these gains may draw on thinking behavior already accessible through the base model's Think mode. We formulate this thinking leakage in a causal mediation framework and audit its contribution using bidirectional interventions along a simple base-derived activation direction. Across three models and three post-training methods on competition math benchmarks, we find that leakage is real, causal, and substantial: behavioral and representational

阅读 arXiv 机器学习 原文 ↗