Docs
Search
Ctrl K
Models
Chat
Compare
Providers
Apps
Rankings
Sign Up
Sign Up
Humanity's Last Exam with Tools Benchmark Leaderboard | Phaseo
Humanity's Last Exam with Tools
Humanity's Last Exam with Tools
AI benchmark results and model performance
Summary
▼
Summary
▼
Type: percentage
Reasoning
View benchmark source
Recorded Results
2
Average Score
60.35%
Scores Over Time
Individual benchmark scores plotted by date.
Models Using This Benchmark
Organisation
Model
Reported
Top Score
Info
Self Reported
Source
Moonshot
Kimi K3
16 Jul 2026
56%
HLE-Full with tools.
Yes
Source
Anthropic
Claude Opus 5
24 Jul 2026
64.70%
Web search, web fetch, programmatic tool calling, and code execution; thinking set to auto; 1M-token cap; contamination controls enabled
Yes
Source
Score Range
56% - 64.70%
Leading Model (lowest score)
56% - Kimi K3