93.8Overall
Model benchmark rankings
Compare scores, test time and reference cost across models and reasoning levels.
Tested on our private datasetsWhy we test this way
Updated 9/30/2026
85.8Overall
89.1Overall
Each model shows its best score here. Expand to see other effort levels.33 models · 83 configurations
About the scores
Backend 40% · Frontend 30% · Reasoning 30%, from the same tested configuration. Each model shows its highest-scoring tested effort for the selected category.
Time and cost follow the selected category. They describe these evaluations, not the speed or price of a typical request.
Explore the tasks →About the method →ⓘ
Ranking preferences and filters
Quality first favors scores; Balance score and time weighs both; Evaluation time first favors shorter task completion times. Each model shows the effort selected by that preference.
These are benchmark completion times, not everyday chat response speeds.
| Rank by | Model / Tested effort | Backend | Frontend | Reasoning | Test time | Reference cost | Compare | |
|---|---|---|---|---|---|---|---|---|
| 01 | gpt-6-astraNEW | 95.3 | 95 | 96 | 95 | 83m 56s | $7.01 | |
| 02 | gpt-6.1-solNEW | 93.8 | 95 | 96 | 90 | 62m 9s | $1.69 | |
| 03 | 89.1 | 81 | 89 | 100 | 33m 57s | $3.67 | ||
| 04 | 89.0 | 83 | 91 | 95 | 79m 7s | $26.15 | ||
| 05 | 85.8 | 84 | 89 | 85 | 48m 32s | $2.80 | ||
| 06 | grok-4.7NEW | 84.3 | 87 | 80 | 85 | 196m 18s | $5.25 | |
| 07 | gpt-6-solNEW | 81.5 | 77 | 94 | 75 | 39m 22s | $1.10 | |
| 08 | 81.0 | 66 | 87 | 95 | 92m 12s | $7.99 | ||
| 09 | grok-4.6NEW | 79.5 | 75 | 80 | 85 | 144m 16s | $2.51 | |
| 10 | gpt-5.6-solNEW | 77.6 | 83 | 73 | 75 | 87m 21s | $3.25 | |
| 11 | 75.9 | 60 | 88 | 85 | 106m 52s | $16.33 | ||
| 12 | k3NEW | 72.5 | 65 | 85 | 70 | 158m 17s | $3.28 | |
| 13 | 71.4 | 66 | 85 | 65 | 336m 3s | $0.66 | ||
| 14 | 70.5 | 78 | 76 | 55 | 133m 2s | $1.94 | ||
| 15 | glm-5.3NEW | 70.0 | 64 | 78 | 70 | 156m 42s | $1.91 | |
| 16 | space-bunnyNEW | 69.3 | 57 | 80 | 75 | 62m 35s | $0.000 | |
| 17 | hy4-previewNEW | 69.0 | 69 | 73 | 65 | 218m 41s | $1.63 | |
| 18 | 68.7 | 57 | 88 | 65 | 35m 40s | $0.41 | ||
| 19 | 68.0 | 71 | 57 | 75 | 54m 7s | N/ANo cost data | ||
| 20 | 67.1 | 65 | 87 | 50 | 134m 31s | $0.42 | ||
| 21 | 66.2 | 56 | 71 | 75 | 98m 22s | $1.24 | ||
| 22 | k2.8 previewNEW | 64.2 | 54 | 72 | 70 | 70m 58s | $0.66 | |
| 23 | 64.0 | 64 | 63 | 65 | 93m 15s | $1.15 | ||
| 24 | 61.8 | 60 | 86 | 40 | 25m 21s | $0.13Partial cost | ||
| 24 | gpt-6-lunaNEW | 61.8 | 57 | 80 | 50 | 63m 52s | $0.061 | |
| 26 | 60.5 | 53 | 76 | 55 | 98m 25s | $0.42 | ||
| 27 | 56.7 | 57 | 43 | 70 | 88m 43s | $0.32 | ||
| 27 | Qwen 3.8 MaxNEW | 56.7 | 60 | 69 | 40 | 133m 3s | $3.41Partial cost | |
| 29 | 56.4 | 54 | 71 | 45 | 386m 14s | $0.28 | ||
| 30 | gpt-5.6-lunaNEW | 55.3 | 61 | 78 | 25 | 86m 14s | $0.35 | |
| 31 | 52.4 | 59 | 71 | 25 | 24m 19s | $0.13 | ||
| 32 | 51.7 | 43 | 75 | 40 | 96m 41s | $0.074Partial cost | ||
| 32 | MiniMax-M3NEW | 51.7 | 43 | 65 | 50 | 56m 33s | $0.62Partial cost |
Each model uses its highest-scoring effort here. Equal scores share a rank.
Backend 40% · Frontend 30% · Reasoning 30%, from the same tested configuration. Evaluation scope and sources →

