01
5 configs87m 38s$7.12
ModelDial Radar
Compare current evidence for real model, effort, and route combinations under one protocol.
Overall shows each model once using its highest measured score. Expand a model to inspect every published effort level.
Highest-scoring configurations · Overall descending
| Rank | Model / highest-scoring configuration | Action | ||||||
|---|---|---|---|---|---|---|---|---|
| 01 | 95.7 | 96.0 | 96.0 | 95.0 | 87m 38s | $7.12Complete cost | ||
| 02 | 89.0 | 83.0 | 91.0 | 95.0 | 79m 7s | $26.15Complete cost | ||
| 03 | 81.0 | 66.0 | 87.0 | 95.0 | 92m 12s | $7.99Complete cost | ||
| 04 | 79.5 | 75.0 | 80.0 | 85.0 | 144m 16s | $2.51Complete cost | ||
| 05 | 78.0 | 84.0 | 73.0 | 75.0 | 92m 25s | $3.11Complete cost | ||
| 06 | 75.9 | 60.0 | 88.0 | 85.0 | 106m 52s | $16.33Complete cost | ||
| 07 | 72.5 | 65.0 | 85.0 | 70.0 | 158m 17s | $3.28Complete cost | ||
| 08 | 69.0 | 69.0 | 73.0 | 65.0 | 218m 41s | $1.63Complete cost | ||
| 09 | 68.5 | 73.0 | 76.0 | 55.0 | 93m 40s | $2.57Complete cost | ||
| 10 | 68.4 | 60.0 | 78.0 | 70.0 | 171m 52s | $1.80Complete cost | ||
| 11 | 67.1 | 65.0 | 87.0 | 50.0 | 134m 31s | $0.42Complete cost | ||
| 12 | 65.1 | 48.0 | 88.0 | 65.0 | 38m 46s | $0.42Complete cost | ||
| 13 | 64.2 | 54.0 | 72.0 | 70.0 | 70m 58s | $0.66Complete cost | ||
| 14 | 62.0 | 59.0 | 63.0 | 65.0 | 96m 24s | $1.11Complete cost | ||
| 15 | 61.0 | 58.0 | 86.0 | 40.0 | 35m 55s | ≥$0.13Partial cost | ||
| 16 | 60.9 | 54.0 | 76.0 | 55.0 | 99m 15s | $0.41Complete cost | ||
| 17 | 58.5 | 60.0 | 75.0 | 40.0 | 122m 53s | ≥$0.08Partial cost | ||
| 18 | 56.7 | 57.0 | 43.0 | 70.0 | 88m 43s | $0.32Complete cost | ||
| 18 | 56.7 | 60.0 | 69.0 | 40.0 | 133m 3s | ≥$3.41Partial cost | ||
| 20 | 56.1 | 63.0 | 78.0 | 25.0 | 79m 36s | $0.34Complete cost | ||
| 21 | 51.7 | 43.0 | 65.0 | 50.0 | 56m 33s | ≥$0.62Partial cost | ||
| 22 | 50.0 | 53.0 | 71.0 | 25.0 | 40m 54s | $0.13Complete cost |
How to read this ranking
Backend, frontend, and knowledge results may update at different times. Overall uses the latest confirmed result for each area. To track changes over time, compare scores from the same test and scoring rules.
A single dip can be normal variation. If scores keep falling across several runs of the same test, the model may actually be getting worse.