Model benchmark rankings

Compare scores, test time and reference cost across models and reasoning levels.

Tested on our private datasetsWhy we test this way

Data snapshot · 10/1/2026

Each model shows its best score here. Expand to see other effort levels.33 models · 81 configurations
About the scores

Backend 40% · Frontend 30% · Reasoning 30%, from the same tested configuration. Each model shows its highest-scoring tested effort for the selected category.

Time and cost follow the selected category. They describe these evaluations, not the speed or price of a typical request.

Explore the tasks →About the method →
Rank by By score
ⓘ
Ranking preferences and filters

Quality first favors scores; Balance score and time weighs both; Evaluation time first favors shorter task completion times. Each model shows the effort selected by that preference.

These are benchmark completion times, not everyday chat response speeds.

All providers
Rank byModel / Tested effortBackendFrontendReasoningTest timeReference costCompare
01
gpt-6-astra
max
95.597949593m 26s$8.35
02
gpt-6.1-solNEW
max
93.895969062m 9s$1.69
03
claude-opus-5-5NEW
max
90.3848910033m 57s$3.67
04
claude-fable-5-1
max
89.885919579m 7s$26.15
05
claude-sonnet-5-5NEW
xhighAnthropic
86.686898548m 32s$2.80
06
grok-4.7
high
84.3878085196m 18s$5.25
07
gpt-6-solNEW
xhigh
82.780947539m 22s$1.10
08
claude-opus-5
max
82.670879592m 12s$7.99
09
grok-4.6
xhigh
80.7788085144m 16s$2.51
10
gpt-5.6-sol
max
78.485737587m 21s$3.25
11
claude-opus-4-8
xhigh
77.5618790102m 10s$9.18
12
k3
highMoonshot
74.5708570158m 17s$3.28
13
mimo-v2.6-proNEW
defaultXiaomi MiMo
73.0708565336m 3s$0.66
14
gpt-5.6-terra
max
71.7817655133m 2s$1.94
15
hy4-preview
highTencent Hunyuan
71.0747365218m 41s$1.63
16
space-bunnyNEW
defaultOther
70.961807562m 35s$0.000
17
deepseek-v4.1-flash
max
70.762886535m 40s$0.41
18
glm-5.3
max
70.4657870157m 1s$1.83
19
MiniMax-M3.1-Flash-PreviewNEW
max
69.675577554m 7sN/ANo cost data
20
Qwen 3.8 Flash
defaultQwen
68.3688750134m 31s$0.42
21
step-5-previewNEW
defaultStepFun
68.261717598m 22s$1.24
22
deepseek-v4-pro
max
65.267636593m 15s$1.15
23
k2.8 preview
high
64.254727070m 58s$0.66
24
gpt-6-lunaNEW
xhigh
63.862805063m 52s$0.061
25
Gemini 3.8 Flash High
defaultGoogle
63.464864025m 21s$0.13Partial cost
26
deepseek-v4-flash
high
62.558765598m 25s$0.42
27
deepseek-v4-flash-vision-exp
high
58.361437088m 43s$0.32
27
Qwen 3.8 Max
defaultQwen
58.3646940133m 3s$3.41Partial cost
29
mimo-v2.6-flashNEW
defaultXiaomi MiMo
58.0587145386m 14s$0.28
30
gpt-5.6-luna
max
56.965782586m 14s$0.35
31
glm-5.3-flash
max
56.1547540100m 0s$0.078Partial cost
32
Gemini 3.7 Flash High
defaultGoogle
54.063712524m 19s$0.13
33
MiniMax-M3
highMiniMax
52.946655056m 33s$0.62Partial cost
Each model uses its highest-scoring effort here. Equal scores share a rank.
Backend 40% · Frontend 30% · Reasoning 30%, from the same tested configuration. Evaluation scope and sources →