RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
Under cross-benchmark transfer evaluation across twelve models completed by all four readouts, difference-in-means attains the highest model-level mean on ten models, while logistic regression wins the most capability-model cells.