AI benchmark results and model performance
Individual benchmark scores plotted by date.
| Organisation | Model | Reported | Top Score | Info | Self Reported | Source |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 24 Jul 2026 | 59.80% | Length-adjusted score; adaptive thinking at max effort; no tools or custom system prompt; Claude Opus 4.8 grader; average of 5 trials | Yes | Source |