Model performance monitoring

Compare model upgrades and check for regressions

Start with published results for the new model, its predecessor and tested efforts. A regression is a separate question that needs repeated observations of the same configuration.

Short answer

Compare task scores, evaluation time and reference cost for a new model and its predecessor. ModelDial evaluates new releases promptly and retains existing scores; it does not promise routine retesting. Without repeated comparable records, a lasting regression cannot be determined.

Compare the new model and tested efforts first

For example, compare GPT-6.1 Sol with GPT-6 Sol, then compare High and Max on the new model. Use the published axes that matter to your work, rather than interpreting different model identities as a time trend.

Keep evaluation completion dates separate from snapshot publication time. Results retained in one publication may have been measured on different dates.

Use historical records only when investigating a regression

Expand a model in the rankings to compare its tested efforts, or use the comparison page for task and subject scores. To study changes over time, retrieve historical results from the publication index on the API page.

  • Look for repeated movement across protocol-matched batches, not one isolated low score.
  • Check whether the change is concentrated in one question, one reasoning level, one route, or the whole model family.
  • Treat stable quality with rising latency differently from a score decline accompanied by failures.

Verify locally before switching

Public Radar contains completed first-party results. It cannot diagnose failures, authentication changes, or routing conditions on your own account.

If the decline repeats, rerun the affected configuration locally, inspect failures and route identity, and compare a nearby effort level or alternate route. Change the default only after the replacement clears the same quality guardrail.