TopLLM
← Back to leaderboards
BENCHMARK RECORD

Humanity's Last Exam

Source-linked result details from the TopLLM database.

BENCHMARK RECORD · 2026-09-15

Humanity's Last Exam

What it is

Published Best Overall (Humanity's Last Exam) scores.

What it tests

Published Best Overall (Humanity's Last Exam) scores.

How to read it

Scores are benchmark-specific and missing values are not estimated.

25 published entries65.0 top reported scoreoverall evaluation domain

Published leaders

Each bar is scaled only within this benchmark.

Back to all benchmarks →
01Claude Fable 5.165.0
02Claude Mythos 5.165.0Unresolved model record
03Claude Opus 564.7Unresolved model record
04Claude Mythos 564.505Claude Opus 4.857.9
06Claude Sonnet 557.4Unresolved model record
07GPT-6 Astra57.2
08Kimi K356.0Unresolved model record
09GLM 5.3 Flash55.310GLM 5.254.7
11Kimi K2.654.0Unresolved model record
12DeepSeek V4 Flash51.613DeepSeek V4 Pro48.2
14GPT-5.6 Sol47.2Unresolved model record
15Gemini 3 Pro45.816Kimi K2 Thinking44.917Gemini 3.1 Pro44.4
18GPT-5.5 Pro43.1Unresolved model record
19GPT-5.541.420Gemini 3.5 Flash40.221Claude Opus 4.640.0
22GPT-535.2Unresolved model record
23Kimi K2.530.1Unresolved model record
24Grok 425.4
25gpt-oss-120b14.9Unresolved model record

Scores are benchmark-specific. Missing values remain N/A and are never inferred from another evaluation.

Compare in the benchmark matrix →