ModelDial Radar

AI coding model benchmark

Compare current evidence for real model, effort, and route combinations under one protocol.

Updated Sep 16, 02:28 PM22 models / 51 configurationsChecking for updates…

Current ranking

Overall shows each model once using its highest measured score. Expand a model to inspect every published effort level.

22 models / 51 configurations
Filters

Highest-scoring configurations · Overall descending

RankModel / highest-scoring configurationAction
0195.796.096.095.087m 38s$7.12Complete cost
0289.083.091.095.079m 7s$26.15Complete cost
0381.066.087.095.092m 12s$7.99Complete cost
0479.575.080.085.0144m 16s$2.51Complete cost
0578.084.073.075.092m 25s$3.11Complete cost
0675.960.088.085.0106m 52s$16.33Complete cost
0772.565.085.070.0158m 17s$3.28Complete cost
0869.069.073.065.0218m 41s$1.63Complete cost
0968.573.076.055.093m 40s$2.57Complete cost
1068.460.078.070.0171m 52s$1.80Complete cost
1167.165.087.050.0134m 31s$0.42Complete cost
1265.148.088.065.038m 46s$0.42Complete cost
1364.254.072.070.070m 58s$0.66Complete cost
1462.059.063.065.096m 24s$1.11Complete cost
1561.058.086.040.035m 55s≥$0.13Partial cost
1660.954.076.055.099m 15s$0.41Complete cost
1758.560.075.040.0122m 53s≥$0.08Partial cost
1856.757.043.070.088m 43s$0.32Complete cost
1856.760.069.040.0133m 3s≥$3.41Partial cost
2056.163.078.025.079m 36s$0.34Complete cost
2151.743.065.050.056m 33s≥$0.62Partial cost
2250.053.071.025.040m 54s$0.13Complete cost
01
5 configs87m 38s$7.12
02
3 configs79m 7s$26.15
03
3 configs92m 12s$7.99
04
2 configs144m 16s$2.51
05
5 configs92m 25s$3.11
06
3 configs106m 52s$16.33
07
158m 17s$3.28
08
218m 41s$1.63
09
5 configs93m 40s$2.57
10
2 configs171m 52s$1.80
11
134m 31s$0.42
12
2 configs38m 46s$0.42
13
70m 58s$0.66
14
2 configs96m 24s$1.11
15
35m 55s≥$0.13
16
2 configs99m 15s$0.41
17
2 configs122m 53s≥$0.08
18
2 configs88m 43s$0.32
18
133m 3s≥$3.41
20
5 configs79m 36s$0.34
21
56m 33s≥$0.62
22
40m 54s$0.13

The overall ranking uses the latest confirmed results. Costs are estimated from public prices and are meant for comparison only.

How to read this ranking

The ranking uses the latest results. Compare history only when the test stays the same.

Backend, frontend, and knowledge results may update at different times. Overall uses the latest confirmed result for each area. To track changes over time, compare scores from the same test and scoring rules.

One lower score does not mean the model got worse.

A single dip can be normal variation. If scores keep falling across several runs of the same test, the model may actually be getting worse.