AI benchmark results and model performance
Individual benchmark scores plotted by date.
| Organisation | Model | Reported | Top Score | Info | Self Reported | Source |
|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 10 Sept 2026 | 31.20% | Instruct model; reasoning_effort=100; DeepSeek Harness Minimal; 1M context; Pass@1 | Yes | Source |