DCAI
← 返回全部动态
arXiv 自然语言处理规则精选09月25日 12:00

Likelihood Ranking doesn't Scale Like Prompting in LLMs

arXiv:2609.29390v1 Announce Type: new Abstract: LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs. Across 95 decoder-only models, ranging from 0.1B to 104B parameters, and 10 MCQA datasets, we find a

阅读 arXiv 自然语言处理 原文 ↗