See something missing, incorrect, or out of date? Tell us about this page.
AI benchmark results and model performance
Individual benchmark scores plotted by date.
| Organisation | Model | Reported | Top Score | Info | Self Reported | Source |
|---|---|---|---|---|---|---|
| Qwen 3.8 27B | 14 Aug 2026 | 84.30% | - | Yes | Source | |
| Muse Glimmer 30B | 10 Aug 2026 | 65.90% | High reasoning; 361-task split; mean reward over four attempts | Yes | Source |