Coding Agent Leaderboard

Performance, efficiency, coverage, and reliability across coding-agent benchmarks. Each result is one model + harness run on one benchmark.

66benchmark results
10models
6harnesses
5benchmarks
88%token coverage
98%cost coverage
98%timing coverage

Cross-benchmark ordering uses within-benchmark percentiles rather than averaging incompatible raw score scales. Missing metrics remain missing and reduce coverage; they are never converted to zero.

Rankings

All ranking charts are shown on one page. Use the shared display controls below, then choose a benchmark for each metric. Tables are capped to a scrollable height so the visualizations stay primary.

Color by
Chart order
Color palette
Image background

Score

Score (%). Higher is better.

Benchmark

Token usage

Total tokens per task. Lower is better.

Benchmark

Cost

Cost per task (USD). Lower is better.

Benchmark

Response time

Total time per task (seconds). Lower is better.

Benchmark

Reliability

Execution error rate (%). Lower is better.

Benchmark

Tokens per successful task

Tokens per successful task. Lower is better.

Benchmark

Cost per successful task

Cost per successful task (USD). Lower is better.

Benchmark

Time per successful task

Time per successful task (seconds). Lower is better.

Benchmark

Ranking data

One table for the page, placed after all charts. It includes the score, resource, timing, reliability, and per-success values for the selected benchmark.

Table benchmark
Sort table by
Table order