Live model readiness

The best modelis the one ready now.

Test the model, effort, and route you can actually reach before you hand them hours of work.

Hover the capsule. Click to open the evidence.

The compact capsule shows the recommended model and effort.

A live route at a glance.

Autoplay
What is testedModel, effort, account, endpoint
What comes backScores, runtime, cost, stability
Decision depthRecommendation / Radar / Compare
Where it staysLocal app and local evidence

You do not have to benchmark anything.

Use the shared Radar as your starting point. Run locally only when your own account, route, or endpoint needs a separate answer.

Provider usage
None to view
Reference refresh
Every six hours
Latest publication
Aug 3, 08:17 UTC
CurrentModelDial first-party
  1. 01gpt-5.6-solxhighRecommendedScore83Time32m 18sReference cost$1.65
  2. 02gpt-5.6-terramaxScore82Time55m 27sReference cost$0.9051
  3. 03gpt-5.6-solmediumScore80Time20m 23sReference cost$0.7228
15 configurations in the published batchView the full ranking

Bring the route only you can see.

Reuse a local sign-in, connect a provider API, or add a compatible endpoint. Every account and route keeps its own evidence.

Local sign-ins
Codex, Claude Code, Grok Build
Provider APIs
OpenAI, DeepSeek, Grok API, OpenRouter
Custom endpoints
Base URL, API format, and key
ModelDial Connections in the native macOS app

5fixed tasks
per profile

Change the scope. Keep the test identical.

Compare two routes, rebuild the whole field, or choose a temporary set. Every selected profile receives the same evidence contract.

Quick Comparison

Current against candidate. The shortest path to the next decision.

Full Scan

Rebuild the local field when models, accounts, or endpoints change.

Custom selection

Choose a temporary scope without changing what stays enabled.

Every completed result stays yours.

Resume what is missing, retry only failures, and refresh when the evidence gets old.

CompletedRetained in the run
MissingReturned to the queue
FailedRetried on their own
Old evidenceRefreshed on schedule

Speed only matters after quality holds.

Runtime, reference cost, and recent stability enter the decision only after a candidate clears your quality guardrail.

Quality guardrailMust hold
RuntimeThen compare
Reference costThen compare
Recent stabilityConfirm the choice

Local by default. Public only by choice.

Your Mac

Keys, configuration, scan history, and recommendations remain local.

Your provider

Requests go directly through the route you chose.

Public Radar

Only ModelDial's first-party reference runs are published.

Capability benchmarks askCan the model solve a fixed task?

ModelDial asksIs my exact route ready for the next work block?

Frequently asked questions

Can I use ModelDial without running my own scan?

Yes. The built-in Radar uses none of your provider quota and is refreshed from a first-party reference run every six hours. Treat it as the default starting point. Run locally only when you need separate evidence for your own account, route, or endpoint.

When is a local scan worth running?

Run locally when your account, endpoint, or custom model may behave differently from the shared reference. Start with Quick Comparison or a custom selection. Use Full Scan only when you need to rebuild the complete local ranking. Personal results stay on your Mac.

How much Codex usage does a Full Scan consume?

A local scan is optional, and its size depends on the scope you choose. For example, a Full Scan with 15 model-and-effort profiles runs five fixed tasks per profile, for 75 planned evaluations. Based on our current reference batch, a comparable mix represents roughly $15-17 in API-equivalent token usage. This is a reference-cost estimate, not a fixed subscription allowance. Actual usage varies with the selected models, reasoning effort, cache use, output length, and retries.

What exactly does ModelDial compare?

Each configuration is an exact combination of model, reasoning effort, account, and endpoint. Quick Comparison evaluates the current and recommended configurations, Full Scan evaluates every enabled configuration, and a custom selection applies only to that run. Each selected configuration receives the same fixed five-task evaluation.

How is the 100-point score calculated?

The current coding-fast-v4.10 pack contains five equally weighted tasks. Each task exposes 20 independently scored behaviors or failure modes, for 100 scoring points in total. A point is earned only when the submitted tests or counterexamples distinguish the reference behavior from the corresponding faulty behavior. The total measures evidence coverage, not a pass probability or a claim that the model will succeed on the same percentage of real repositories.

What do the five tasks test?

Task 1 tests black-box contract and boundary analysis across atomic saves, snapshots, deterministic archives, and replay. Task 2 tests counterexample design for retry selection, readiness, capacity, ordering, and wake-up logic. Task 3 tests adversarial CI-plan auditing across dependency propagation, fallback coverage, scheduling, cost, and ranking. Task 4 tests transaction state-machine reasoning across rejection, wave snapshots, conflicts, retries, and tombstones. Task 5 tests cache regression design across identity, invalidation, forced and warm scans, issue reporting, and state preservation.

Does a scan inspect or modify my repository?

No. ModelDial evaluates a fixed task set rather than your current repository. Local Codex evaluations run in an ephemeral, read-only workspace. API routes receive only the fixed evaluation request through the provider endpoint you configured.

How does ModelDial choose a recommendation?

ModelDial compares complete, compatible results produced with the same question pack and scoring protocol. Your selected goal determines how eligible alternatives are ranked, while an explicit quality guardrail limits the tradeoff. If evidence is incomplete or no alternative provides a meaningful benefit, ModelDial keeps the current configuration or asks for another test.

What happens if a task fails or hits a rate limit?

Completed evidence is retained. You can retry only the failed evaluations instead of restarting the entire batch. Retries may consume additional provider quota or incur additional API charges.

Can I add a custom model or endpoint?

Yes. Add custom endpoints and their model candidates in Settings, or select a custom configuration set for a single run. ModelDial keeps each account and endpoint distinct, so evidence from different routes is never merged into one result.

Download ModelDial for macOS.

The signed build will appear here after release checks. Until then, the public Radar is ready to explore.

Release in progressOpen Radar
View on GitHub