DCAI
← 返回全部动态
arXiv 自然语言处理规则精选09月24日 12:00

Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models

arXiv:2609.27510v1 Announce Type: new Abstract: Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is particularly challenging for base models, whose limited instruction-following ability complicates task-based assessment. We introduce Uncheatable Eval, a dynamic benchmark that regularly collects newly published text to evaluate base language models and reduce the risk of data contamination. Drawing on the relationship between a model's predictive ability and its ability to co

阅读 arXiv 自然语言处理 原文 ↗