# GPQA Diamond Clean Results 0.1.0

Status: **published**
Published: **2026-08-26**
Data through: **2026-08-17**

This immutable, answer-safe bundle contains 243 accepted evaluation configurations rescored on GPQA-Diamond-Clean 1.2.2. The default leaderboard combines fully accepted `mean@N` protocols in one descriptive rank. Repeat count remains visible and the ranking is not normalized for inference compute.

The source top-50 coverage snapshot is 38/50, with 12 exact-configuration evidence gaps documented in `coverage-gaps.json`.

## Files

- `release-manifest.json`: release, input, contract, count, toolchain, and artifact hashes.
- `leaderboard.json` and `leaderboard.csv`: exact paired Original-198 and Clean-189 scores and standard competition rank.
- `score-changes.json`: paired changes, ranks, thresholds, and separately labelled cohort summaries.
- `uncertainty.json`: 20,000-replicate synchronized Record-ID bootstrap intervals, encoded as decimal strings.
- `coverage-gaps.json`: top-50 evidence diagnoses and release decisions.
- `category-summary.json`: aggregate-only domain and subcategory accuracy with disclosure floors.
- `configuration-manifests.json`: stable configuration, run, evidence, protocol, and reconciliation metadata.
- `methodology.json`: machine-readable scientific and disclosure contracts.
- `schemas/`: self-contained JSON Schemas.
- `checksums.sha256`: hashes for the manifest and every payload file.

Verify from this directory:

```sh
shasum -a 256 -c checksums.sha256
```

`checksums.sha256` necessarily excludes itself. `release-manifest.json` inventories every payload except itself and the checksum file.

## Interpretation

`Original 198` and `Clean 189` are reconstructed from the same accepted attempt universe. The delta never subtracts a published roster score from an unrelated reconstructed result. `mean@N` is arithmetic mean across N attempts, not best-of-N, pass@N, or majority voting.

Observed standard competition rank is retained even when bootstrap rank intervals overlap. Close frontier positions should not be described as decisive when the intervals are broad.

For canonical evidence selection, every hash-pinned accepted input is assigned the frozen release timestamp as an administrative, deterministic selection value. It is not a claim about when an evaluation ran, when evidence was first observed, or when it was actually admitted. Configuration dates separately identify validated source dates where present and an explicit frozen-release-date fallback where the answer-safe native source has no run timestamp.

## Public boundary

This bundle contains aggregate result data only. It contains no GPQA task text, options, answer keys, Record IDs, model completions, reasoning, raw evaluation archives, private evidence locations, or model-by-question matrix.

## Attribution

Please cite the original GPQA benchmark, GPQA-Diamond-Clean, the answer-key audit, and the evidence sources. The corrected benchmark is available from [GitHub](https://github.com/adamallcock/gpqa-diamond-clean) and [Hugging Face](https://huggingface.co/datasets/adamallcock/gpqa-diamond-clean). The original benchmark paper is [Rein et al.](https://arxiv.org/abs/2311.12022).

This is an independent research project. Inclusion is not a provider submission, certification, or endorsement.
