One combined rank.
Every accepted mean@N protocol shares the ranking. Repeat count remains explicit and is not normalized away.
How ranking worksCombined ranking 243 configurations 189 retained questions
Rein et al. introduced GPQA (opens in a new tab), a 448-question graduate-level science benchmark; GPQA Diamond (opens in a new tab) is the paper's 198-question highest-quality subset. GPQA Diamond Clean (opens in a new tab) is this independent project's corrected overlay for Diamond, retaining 189 questions after nine audited exclusions—not an official update from the GPQA authors. The leaderboard credits Epoch AI's public GPQA Diamond evaluations (opens in a new tab) where used, while derived, recovery, and project-native evidence remain explicitly labelled.
| Rank Rank movement Compares this configuration's standard competition rank on Clean 189 with its paired rank on Original 198.
Ranks use exact scores. Filters never recompute them. | Model | Clean score | Clean change | Benchmark source Benchmark source classes How each score entered the combined ranking.
|
|---|---|---|---|---|
| 1↑3 |
|
Clean score97.4% | Clean change+3.4 ppfrom 93.9% | Benchmark sourceDerived uniform |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 1↑3 |
|
Clean score97.4% | Clean change+3.4 ppfrom 93.9% | Benchmark sourceNative transport composite |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 3— |
|
Clean score97.2% | Clean change+3.2 ppfrom 94.0% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 4↓3 |
|
Clean score97.0% | Clean change+2.5 ppfrom 94.5% | Benchmark sourceTargeted recovery |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 5↑1 |
|
Clean score96.8% | Clean change+3.4 ppfrom 93.4% | Benchmark sourceComplete native primary |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 6↑1 |
|
Clean score96.7% | Clean change+3.4 ppfrom 93.2% | Benchmark sourceTargeted recovery |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 7↓5 |
|
Clean score96.3% | Clean change+1.9 ppfrom 94.4% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 8↑2 |
|
Clean score96.1% | Clean change+3.3 ppfrom 92.8% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 9↑3 |
|
Clean score95.8% | Clean change+3.3 ppfrom 92.4% | Benchmark sourceComplete native primary |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 9↓1 |
|
Clean score95.8% | Clean change+2.8 ppfrom 92.9% | Benchmark sourceComplete native primary |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 11— |
|
Clean score95.5% | Clean change+2.9 ppfrom 92.6% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 12— |
|
Clean score95.2% | Clean change+2.8 ppfrom 92.4% | Benchmark sourceComplete native primary |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 12↓4 |
|
Clean score95.2% | Clean change+2.3 ppfrom 92.9% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 14↑2 |
|
Clean score95.0% | Clean change+3.6 ppfrom 91.4% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 15↓3 |
|
Clean score94.7% | Clean change+2.3 ppfrom 92.4% | Benchmark sourceComplete native primary |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 15↑2 |
|
Clean score94.7% | Clean change+3.8 ppfrom 90.9% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 15— |
|
Clean score94.7% | Clean change+2.8 ppfrom 91.9% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 18↑4 |
|
Clean score93.8% | Clean change+3.2 ppfrom 90.7% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 19↓2 |
|
Clean score93.8% | Clean change+2.9 ppfrom 90.9% | Benchmark sourceDerived uniform |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 20↓3 |
|
Clean score93.7% | Clean change+2.7 ppfrom 90.9% | Benchmark sourceComplete native primary |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 20↓3 |
|
Clean score93.7% | Clean change+2.7 ppfrom 90.9% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 20↑6 |
|
Clean score93.7% | Clean change+3.8 ppfrom 89.9% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 20↑4 |
|
Clean score93.7% | Clean change+3.2 ppfrom 90.4% | Benchmark sourceComplete native primary |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 20↑5 |
|
Clean score93.7% | Clean change+3.5 ppfrom 90.2% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
| 20↑6 |
|
Clean score93.7% | Clean change+3.8 ppfrom 89.9% | Benchmark sourceStrict |
Evidence detailExpand this record to load its answer-safe evidence detail. | ||||
No configurations match these filters.
What changed
Cleaning did not make models more capable. It removed questions that could not fairly distinguish a correct answer, then recomputed every accepted configuration on the same retained membership.
The broader answer-key audit
This leaderboard is the GPQA Diamond results layer of a broader calibrated audit of GPQA, MMLU-Pro, and MMMU-Pro. The paper explains how questionable items were found, how over-calls were stripped away, and why a single disagreement pass can miss benchmark defects.
Latest published paper10.5281/zenodo.21613590 (opens in a new tab)
Every accepted mean@N protocol shares the ranking. Repeat count remains explicit and is not normalized away.
How ranking works5 accepted evidence pathways are reconciled into 243 canonical configurations without hiding their source class.
Use the evidence key