DCAI
← 返回全部动态
arXiv 自然语言处理规则精选09月25日 12:00

Grammatical "grandmother neurons" are rare in LLMs

arXiv:2609.29328v1 Announce Type: new Abstract: Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or "probes") are widely used for this task, they face significant methodological criticism: training auxiliary classifiers introduces capacity confounds and calibration issues, often making it difficult to distinguish the model's intrinsic representations from the probe's ability to learn the task. To address these limitations, we introduce a probe-free framework for localizing linguistic selectivity at the individual neuron level

阅读 arXiv 自然语言处理 原文 ↗