gpt-6-sol
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
gpt-6-sol · xhigh
- Backend & testing
- 80
- Frontend & interaction
- 94
- Knowledge & reasoning
- 75
- Evaluation time
- 39m 22s
- Reference cost
- $1.10
Totals for all three evaluations at this configuration; cost estimated from API usage.
Overall rank 7 / 33 models · Best tested effort
Backend tested · 9/23/2026
Published summary scores use the same effort. Axes may be evaluated on different dates; a backend date refers only to backend testing.
View task coverage and scoring →What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
- Overall score
- +7.0 pts 75.7 → 82.7
- Reference cost
- +$0.29 $0.81 → $1.10
- Evaluation time
- +24.0% 31m 45s → 39m 22s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| low | 66.9 | 54 | 81 | 70 | 33m 0s | $0.59 | |
| medium | 73.0 | 70 | 85 | 65 | 35m 34s | $0.77 | |
| high | 75.7 | 76 | 81 | 70 | 31m 45s | $0.81 | |
| xhigh | 82.7 | 80 | 94 | 75 | 39m 22s | $1.10 | |
| max | 80.0 | 80 | 90 | 70 | 35m 40s | $1.14 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Data snapshot published 10/1/2026