Compare the new model and tested efforts first
For example, compare GPT-6.1 Sol with GPT-6 Sol, then compare High and Max on the new model. Use the published axes that matter to your work, rather than interpreting different model identities as a time trend.
Keep evaluation completion dates separate from snapshot publication time. Results retained in one publication may have been measured on different dates.
Use historical records only when investigating a regression
Expand a model in the rankings to compare its tested efforts, or use the comparison page for task and subject scores. To study changes over time, retrieve historical results from the publication index on the API page.
- Look for repeated movement across protocol-matched batches, not one isolated low score.
- Check whether the change is concentrated in one question, one reasoning level, one route, or the whole model family.
- Treat stable quality with rising latency differently from a score decline accompanied by failures.
Verify locally before switching
Public Radar contains completed first-party results. It cannot diagnose failures, authentication changes, or routing conditions on your own account.
If the decline repeats, rerun the affected configuration locally, inspect failures and route identity, and compare a nearby effort level or alternate route. Change the default only after the replacement clears the same quality guardrail.