93.8Overall
Model benchmark rankings
Compare scores, test time and reference cost across models and reasoning levels.
Tested on our private datasetsWhy we test this way
Data snapshot · 10/1/2026
86.6Overall
90.3Overall
Each model shows its best score here. Expand to see other effort levels.33 models · 81 configurations
About the scores
Backend 40% · Frontend 30% · Reasoning 30%, from the same tested configuration. Each model shows its highest-scoring tested effort for the selected category.
Time and cost follow the selected category. They describe these evaluations, not the speed or price of a typical request.
Explore the tasks →About the method →ⓘ
Ranking preferences and filters
Quality first favors scores; Balance score and time weighs both; Evaluation time first favors shorter task completion times. Each model shows the effort selected by that preference.
These are benchmark completion times, not everyday chat response speeds.
| Rank by | Model / Tested effort | Backend | Frontend | Reasoning | Test time | Reference cost | Compare | |
|---|---|---|---|---|---|---|---|---|
| 01 | gpt-6-astraNEW | 95.5 | 97 | 94 | 95 | 93m 26s | $8.35 | |
| 02 | gpt-6.1-solNEW | 93.8 | 95 | 96 | 90 | 62m 9s | $1.69 | |
| 03 | 90.3 | 84 | 89 | 100 | 33m 57s | $3.67 | ||
| 04 | 89.8 | 85 | 91 | 95 | 79m 7s | $26.15 | ||
| 05 | 86.6 | 86 | 89 | 85 | 48m 32s | $2.80 | ||
| 06 | grok-4.7NEW | 84.3 | 87 | 80 | 85 | 196m 18s | $5.25 | |
| 07 | gpt-6-solNEW | 82.7 | 80 | 94 | 75 | 39m 22s | $1.10 | |
| 08 | 82.6 | 70 | 87 | 95 | 92m 12s | $7.99 | ||
| 09 | grok-4.6NEW | 80.7 | 78 | 80 | 85 | 144m 16s | $2.51 | |
| 10 | gpt-5.6-solNEW | 78.4 | 85 | 73 | 75 | 87m 21s | $3.25 | |
| 11 | 77.5 | 61 | 87 | 90 | 102m 10s | $9.18 | ||
| 12 | k3NEW | 74.5 | 70 | 85 | 70 | 158m 17s | $3.28 | |
| 13 | 73.0 | 70 | 85 | 65 | 336m 3s | $0.66 | ||
| 14 | 71.7 | 81 | 76 | 55 | 133m 2s | $1.94 | ||
| 15 | hy4-previewNEW | 71.0 | 74 | 73 | 65 | 218m 41s | $1.63 | |
| 16 | space-bunnyNEW | 70.9 | 61 | 80 | 75 | 62m 35s | $0.000 | |
| 17 | 70.7 | 62 | 88 | 65 | 35m 40s | $0.41 | ||
| 18 | glm-5.3NEW | 70.4 | 65 | 78 | 70 | 157m 1s | $1.83 | |
| 19 | 69.6 | 75 | 57 | 75 | 54m 7s | N/ANo cost data | ||
| 20 | 68.3 | 68 | 87 | 50 | 134m 31s | $0.42 | ||
| 21 | 68.2 | 61 | 71 | 75 | 98m 22s | $1.24 | ||
| 22 | 65.2 | 67 | 63 | 65 | 93m 15s | $1.15 | ||
| 23 | k2.8 previewNEW | 64.2 | 54 | 72 | 70 | 70m 58s | $0.66 | |
| 24 | gpt-6-lunaNEW | 63.8 | 62 | 80 | 50 | 63m 52s | $0.061 | |
| 25 | 63.4 | 64 | 86 | 40 | 25m 21s | $0.13Partial cost | ||
| 26 | 62.5 | 58 | 76 | 55 | 98m 25s | $0.42 | ||
| 27 | 58.3 | 61 | 43 | 70 | 88m 43s | $0.32 | ||
| 27 | Qwen 3.8 MaxNEW | 58.3 | 64 | 69 | 40 | 133m 3s | $3.41Partial cost | |
| 29 | 58.0 | 58 | 71 | 45 | 386m 14s | $0.28 | ||
| 30 | gpt-5.6-lunaNEW | 56.9 | 65 | 78 | 25 | 86m 14s | $0.35 | |
| 31 | 56.1 | 54 | 75 | 40 | 100m 0s | $0.078Partial cost | ||
| 32 | 54.0 | 63 | 71 | 25 | 24m 19s | $0.13 | ||
| 33 | MiniMax-M3NEW | 52.9 | 46 | 65 | 50 | 56m 33s | $0.62Partial cost |
Each model uses its highest-scoring effort here. Equal scores share a rank.
Other category scores
Backend 40% · Frontend 30% · Reasoning 30%, from the same tested configuration. Evaluation scope and sources →

