Research2026-05-01
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
Source: Arxiv CS.AI
arXiv:2604.27405v1 Announce Type: cross Abstract: We adapted the Reliable Change Index (RCI; Jacobson and Truax, 1991) from clinical psychology to item-level LLM version comparison on 2,000 MMLU-Pro items (K=10 samples at T=0.7). Two within-family pairs were tested: Llama 3 to 3.1 (+1.6 points) and...
arxivpapers