GPT-5.6
Compare tested reasoning efforts on private datasets. Scores and costs come from their recorded evaluations.
gpt-5.6-sol · max
- Backend & testing
- 85
- Frontend & interaction
- 73
- Knowledge & reasoning
- 75
- Evaluation time
- 87m 21s
- Reference cost
- $3.25
Totals for all three evaluations at this configuration; cost estimated from API usage.
Overall rank 10 / 33 models · Best tested effort
Backend tested · 9/21/2026
Published summary scores use the same effort. Axes may be evaluated on different dates; a backend date refers only to backend testing.
View task coverage and scoring →What changes with more effort
Compare scores, cost and time at neighboring tested efforts.
- Overall score
- +8.2 pts 70.2 → 78.4
- Reference cost
- +$0.69 $2.56 → $3.25
- Evaluation time
- +22.5% 71m 19s → 87m 21s
Results by reasoning effort
Select two or three configurations to compare scores and costs.
Swipe to see all results →
| Model / effort | Overall | Backend & testing | Frontend & interaction | Knowledge & reasoning | Evaluation time | Reference cost | Compare |
|---|---|---|---|---|---|---|---|
| gpt-5.6-sollow | 69.0 | 69 | 78 | 60 | 41m 49s | $1.53 | |
| gpt-5.6-terralow | 36.5 | 53 | 21 | 30 | 43m 42s | $0.71 | |
| gpt-5.6-lunalow | 19.7 | 32 | 18 | 5 | 52m 31s | $0.067 | |
| gpt-5.6-solmedium | 75.1 | 73 | 83 | 70 | 59m 12s | $1.72 | |
| gpt-5.6-terramedium | 49.6 | 49 | 50 | 50 | 46m 2s | $0.75 | |
| gpt-5.6-lunamedium | 31.8 | 51 | 13 | 25 | 18m 22s | $0.084 | |
| gpt-5.6-solhigh | 77.7 | 78 | 90 | 65 | 58m 24s | $2.05 | |
| gpt-5.6-terrahigh | 59.6 | 65 | 67 | 45 | 53m 5s | $1.03 | |
| gpt-5.6-lunahigh | 39.7 | 49 | 52 | 15 | 32m 18s | $0.14 | |
| gpt-5.6-solxhigh | 70.2 | 78 | 65 | 65 | 71m 19s | $2.56 | |
| gpt-5.6-terraxhigh | 58.9 | 67 | 72 | 35 | 67m 17s | $1.54 | |
| gpt-5.6-lunaxhigh | 54.8 | 56 | 83 | 25 | 75m 23s | $0.16 | |
| gpt-5.6-solmax | 78.4 | 85 | 73 | 75 | 87m 21s | $3.25 | |
| gpt-5.6-terramax | 71.7 | 81 | 76 | 55 | 133m 2s | $1.94 | |
| gpt-5.6-lunamax | 56.9 | 65 | 78 | 25 | 86m 14s | $0.35 |
How to read these results
The summary uses the same configuration as the highest overall score. A complete overall score requires all three axes; missing results appear as dashes.
Data snapshot published 10/1/2026