The Information Machine
Following·since 31 Aug 2026·Day 2·2 sources

Evaluation Setup, Not Just Model Quality, Drives LLM Benchmark Rankings

The gist

LLM benchmark rankings are widely used to compare models; these findings suggest that a single evaluation configuration can misrepresent relative model capability. If ranking positions reflect methodology choices more than model differences, comparisons based on standard leaderboards may be unreliable.

The full picture

A study holding 3,679 questions and 12 models fixed found that changing only evaluation parameters, such as prompt format, option order, and scoring method, produced large swings in benchmark rankings. Gemma 4-31B scored anywhere from 31% to 89% depending solely on evaluation setup, and four of the twelve models reached rank 1 under at least one valid configuration. The study found that 95.7% of the average gap between neighboring models stems from questions whose answers change with the evaluation setup, and that the scoring method, whether models generate a free-form answer or select the highest-likelihood option, is the single biggest driver of ranking instability. A separate framework called BenchMIRT examines what LLM benchmarks are actually measuring versus what they claim to measure, raising construct validity questions.

How it developed
1 September 2026

BenchMIRT framework published on Hugging Face, examining what LLM benchmarks actually measure versus what they claim to measure.

31 August 2026

Study shows that changing only evaluation parameters across 12 models and 3,679 questions causes large swings in benchmark rankings, with Gemma 4-31B ranging from 31% to 89%.

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free