TopLLM
← Back to leaderboards
BENCHMARK RECORD

SWE-Bench Verified

Source-linked result details from the TopLLM database.

BENCHMARK RECORD · 2026-10-11

SWE-Bench Verified

What it is

Published Best in Agentic Coding (SWE Bench) scores.

What it tests

Published Best in Agentic Coding (SWE Bench) scores.

How to read it

Scores are benchmark-specific and missing values are not estimated.

10 published entries96.2 top reported scorecoding evaluation domain

Published leaders

Each bar is scaled only within this benchmark.

Back to all benchmarks →
01GPT-5.6 Sol96.202Claude Mythos 595.5
03Claude Fable 595.0Unresolved model record
04GPT-5.6 Luna93.0Unresolved model record
05Claude Opus 4.888.6Unresolved model record
06DeepSeek V4 Pro80.607MiniMax M380.508Kimi K2.680.209DeepSeek V4 Flash79.010Kimi K2.576.8

Scores are benchmark-specific. Missing values remain N/A and are never inferred from another evaluation.

Compare in the benchmark matrix →