Model benchmark rankings

Compare scores, test time and reference cost across models and reasoning levels.

Tested on our private datasetsWhy we test this way

Updated 9/30/2026

NewWorth a look
93.8Overall
89.1Overall
Each model shows its best score here. Expand to see other effort levels.33 models · 83 configurations
About the scores

Backend 40% · Frontend 30% · Reasoning 30%, from the same tested configuration. Each model shows its highest-scoring tested effort for the selected category.

Time and cost follow the selected category. They describe these evaluations, not the speed or price of a typical request.

Explore the tasks →About the method →
Rank by By score
ⓘ
Ranking preferences and filters

Quality first favors scores; Balance score and time weighs both; Evaluation time first favors shorter task completion times. Each model shows the effort selected by that preference.

These are benchmark completion times, not everyday chat response speeds.

All providers
Rank byModel / Tested effortBackendFrontendReasoningTest timeReference costCompare
01
gpt-6-astra
xhigh
95.395969583m 56s$7.01
02
gpt-6.1-solNEW
max
93.895969062m 9s$1.69
03
claude-opus-5-5NEW
max
89.1818910033m 57s$3.67
04
claude-fable-5-1
max
89.083919579m 7s$26.15
05
claude-sonnet-5-5NEW
xhighAnthropic
85.884898548m 32s$2.80
06
grok-4.7
high
84.3878085196m 18s$5.25
07
gpt-6-solNEW
xhigh
81.577947539m 22s$1.10
08
claude-opus-5
max
81.066879592m 12s$7.99
09
grok-4.6
xhigh
79.5758085144m 16s$2.51
10
gpt-5.6-sol
max
77.683737587m 21s$3.25
11
claude-opus-4-8
max
75.9608885106m 52s$16.33
12
k3
highMoonshot
72.5658570158m 17s$3.28
13
mimo-v2.6-proNEW
defaultXiaomi MiMo
71.4668565336m 3s$0.66
14
gpt-5.6-terra
max
70.5787655133m 2s$1.94
15
glm-5.3
max
70.0647870156m 42s$1.91
16
space-bunnyNEW
defaultOther
69.357807562m 35s$0.000
17
hy4-preview
highTencent Hunyuan
69.0697365218m 41s$1.63
18
deepseek-v4.1-flash
max
68.757886535m 40s$0.41
19
MiniMax-M3.1-Flash-PreviewNEW
max
68.071577554m 7sN/ANo cost data
20
Qwen 3.8 Flash
defaultQwen
67.1658750134m 31s$0.42
21
step-5-previewNEW
defaultStepFun
66.256717598m 22s$1.24
22
k2.8 preview
high
64.254727070m 58s$0.66
23
deepseek-v4-pro
max
64.064636593m 15s$1.15
24
Gemini 3.8 Flash High
defaultGoogle
61.860864025m 21s$0.13Partial cost
24
gpt-6-lunaNEW
xhigh
61.857805063m 52s$0.061
26
deepseek-v4-flash
high
60.553765598m 25s$0.42
27
deepseek-v4-flash-vision-exp
high
56.757437088m 43s$0.32
27
Qwen 3.8 Max
defaultQwen
56.7606940133m 3s$3.41Partial cost
29
mimo-v2.6-flashNEW
defaultXiaomi MiMo
56.4547145386m 14s$0.28
30
gpt-5.6-luna
max
55.361782586m 14s$0.35
31
Gemini 3.7 Flash High
defaultGoogle
52.459712524m 19s$0.13
32
glm-5.3-flash
max
51.743754096m 41s$0.074Partial cost
32
MiniMax-M3
highMiniMax
51.743655056m 33s$0.62Partial cost
Each model uses its highest-scoring effort here. Equal scores share a rank.
Backend 40% · Frontend 30% · Reasoning 30%, from the same tested configuration. Evaluation scope and sources →