GPQA Diamond Clean leaderboard

Combined ranking 243 configurations 189 retained questions

Rein et al. introduced GPQA (opens in a new tab), a 448-question graduate-level science benchmark; GPQA Diamond (opens in a new tab) is the paper's 198-question highest-quality subset. GPQA Diamond Clean (opens in a new tab) is this independent project's corrected overlay for Diamond, retaining 189 questions after nine audited exclusions—not an official update from the GPQA authors. The leaderboard credits Epoch AI's public GPQA Diamond evaluations (opens in a new tab) where used, while derived, recovery, and project-native evidence remain explicitly labelled.

Showing the first 25 of 243 configurations while the full leaderboard loads.

CSV
Combined GPQA Diamond Clean ranking
Rank
Rank movement

Compares this configuration's standard competition rank on Clean 189 with its paired rank on Original 198.

Improved
Fell
Unchanged

Ranks use exact scores. Filters never recompute them.

ModelClean scoreClean changeBenchmark source
Benchmark source classes

How each score entered the combined ranking.

  • Strict

    Complete public evaluation evidence validated under the current contract.

  • Compatibility

    Complete legacy evidence translated through a versioned compatibility validator.

  • Derived uniform

    A globally complete repeat subset reconstructed consistently for every question.

  • Targeted recovery

    Original attempts plus only the exact failed source slots recovered under a declared route.

  • Complete native primary

    A complete independent all-198 run on one declared primary route.

  • Native transport composite

    A complete independent source run plus documented exact-route transport recovery.

Full admission method →
1↑3 GPT 5.5 Pro pre-releaseOpenAI · mean@7 · xhighDerived uniform Clean score97.4% Clean change+3.4 ppfrom 93.9%
Benchmark sourceDerived uniform
1↑3 Gemini 3.6 FlashGoogle · mean@1 · highNative transport composite Clean score97.4% Clean change+3.4 ppfrom 93.9%
Benchmark sourceNative transport composite
3 GPT 5.5 pre-releaseOpenAI · mean@8 · xhighStrict Clean score97.2% Clean change+3.2 ppfrom 94.0%
Benchmark sourceStrict
4↓3 Gemini 3.1 Pro PreviewGoogle · mean@8 · defaultTargeted recovery Clean score97.0% Clean change+2.5 ppfrom 94.5%
Benchmark sourceTargeted recovery
5↑1 GPT 5.6 SolOpenAI · mean@1 · maxComplete native primary Clean score96.8% Clean change+3.4 ppfrom 93.4%
Benchmark sourceComplete native primary
6↑1 GPT 5.4OpenAI · mean@8 · xhighTargeted recovery Clean score96.7% Clean change+3.4 ppfrom 93.2%
Benchmark sourceTargeted recovery
7↓5 Gemini 3.1 Pro PreviewGoogle · mean@1 · highStrict Clean score96.3% Clean change+1.9 ppfrom 94.4%
Benchmark sourceStrict
8↑2 Gemini 3.5 FlashGoogle · mean@8 · highStrict Clean score96.1% Clean change+3.3 ppfrom 92.8%
Benchmark sourceStrict
9↑3 Grok 4.6xAI · mean@1 · highComplete native primary Clean score95.8% Clean change+3.3 ppfrom 92.4%
Benchmark sourceComplete native primary
9↓1 DeepSeek V4 FlashDeepSeek · mean@1 · maxComplete native primary Clean score95.8% Clean change+2.8 ppfrom 92.9%
Benchmark sourceComplete native primary
11 Gemini 3 Pro PreviewGoogle · mean@8 · defaultStrict Clean score95.5% Clean change+2.9 ppfrom 92.6%
Benchmark sourceStrict
12 DeepSeek V4 ProDeepSeek · mean@1 · maxComplete native primary Clean score95.2% Clean change+2.8 ppfrom 92.4%
Benchmark sourceComplete native primary
12↓4 Claude Opus 5Anthropic · mean@1 · noneStrict Clean score95.2% Clean change+2.3 ppfrom 92.9%
Benchmark sourceStrict
14↑2 GPT 5.2OpenAI · mean@8 · xhighStrict Clean score95.0% Clean change+3.6 ppfrom 91.4%
Benchmark sourceStrict
15↓3 GPT 5.6 LunaOpenAI · mean@1 · maxComplete native primary Clean score94.7% Clean change+2.3 ppfrom 92.4%
Benchmark sourceComplete native primary
15↑2 Qwen3.7 MaxQwen · mean@1 · maxStrict Clean score94.7% Clean change+3.8 ppfrom 90.9%
Benchmark sourceStrict
15 Kimi K3Moonshot AI · mean@1 · highStrict Clean score94.7% Clean change+2.8 ppfrom 91.9%
Benchmark sourceStrict
18↑4 GPT 5.5OpenAI · mean@8 · lowStrict Clean score93.8% Clean change+3.2 ppfrom 90.7%
Benchmark sourceStrict
19↓2 Kimi K2.6Moonshot AI · mean@6 · defaultDerived uniform Clean score93.8% Clean change+2.9 ppfrom 90.9%
Benchmark sourceDerived uniform
20↓3 GPT 5.6 TerraOpenAI · mean@1 · maxComplete native primary Clean score93.7% Clean change+2.7 ppfrom 90.9%
Benchmark sourceComplete native primary
20↓3 MiniMax M3MiniMax · mean@1 · defaultStrict Clean score93.7% Clean change+2.7 ppfrom 90.9%
Benchmark sourceStrict
20↑6 GPT 5.4OpenAI · mean@1 · highStrict Clean score93.7% Clean change+3.8 ppfrom 89.9%
Benchmark sourceStrict
20↑4 DeepSeek V4 FlashDeepSeek · mean@1 · highComplete native primary Clean score93.7% Clean change+3.2 ppfrom 90.4%
Benchmark sourceComplete native primary
20↑5 Claude Opus 4.7Anthropic · mean@8 · xhighStrict Clean score93.7% Clean change+3.5 ppfrom 90.2%
Benchmark sourceStrict
20↑6 GPT 5.6 SolOpenAI · mean@1 · lowStrict Clean score93.7% Clean change+3.8 ppfrom 89.9%
Benchmark sourceStrict

What changed

Nine defective questions moved scores and ranks.

Cleaning did not make models more capable. It removed questions that could not fairly distinguish a correct answer, then recomputed every accepted configuration on the same retained membership.

Scores that rose
235 / 243
Crossed 95%
14
Observed ranks changed
168
Unresolved evidence
0

The broader answer-key audit

When the Answer Key Is Wrong

This leaderboard is the GPQA Diamond results layer of a broader calibrated audit of GPQA, MMLU-Pro, and MMMU-Pro. The paper explains how questionable items were found, how over-calls were stripped away, and why a single disagreement pass can miss benchmark defects.

Adam Allcock · Preprint · July 24, 2026

Latest published paper10.5281/zenodo.21613590 (opens in a new tab)

01

One combined rank.

Every accepted mean@N protocol shares the ranking. Repeat count remains explicit and is not normalized away.

How ranking works
02

Every score carries evidence.

5 accepted evidence pathways are reconciled into 243 canonical configurations without hiding their source class.

Use the evidence key