Quick Comparison
Current against candidate. The shortest path to the next decision.
Live model readiness
Test the model, effort, and route you can actually reach before you hand them hours of work.
Hover the capsule. Click to open the evidence.
The compact capsule shows the recommended model and effort.
A live route at a glance.
Use the shared Radar as your starting point. Run locally only when your own account, route, or endpoint needs a separate answer.
Reuse a local sign-in, connect a provider API, or add a compatible endpoint. Every account and route keeps its own evidence.

5fixed tasks
per profile
Compare two routes, rebuild the whole field, or choose a temporary set. Every selected profile receives the same evidence contract.
Current against candidate. The shortest path to the next decision.
Rebuild the local field when models, accounts, or endpoints change.
Choose a temporary scope without changing what stays enabled.
Resume what is missing, retry only failures, and refresh when the evidence gets old.
Runtime, reference cost, and recent stability enter the decision only after a candidate clears your quality guardrail.
Keys, configuration, scan history, and recommendations remain local.
Requests go directly through the route you chose.
Only ModelDial's first-party reference runs are published.
Capability benchmarks askCan the model solve a fixed task?
ModelDial asksIs my exact route ready for the next work block?
Yes. The built-in Radar uses none of your provider quota and is refreshed from a first-party reference run every six hours. Treat it as the default starting point. Run locally only when you need separate evidence for your own account, route, or endpoint.
Run locally when your account, endpoint, or custom model may behave differently from the shared reference. Start with Quick Comparison or a custom selection. Use Full Scan only when you need to rebuild the complete local ranking. Personal results stay on your Mac.
A local scan is optional, and its size depends on the scope you choose. For example, a Full Scan with 15 model-and-effort profiles runs five fixed tasks per profile, for 75 planned evaluations. Based on our current reference batch, a comparable mix represents roughly $15-17 in API-equivalent token usage. This is a reference-cost estimate, not a fixed subscription allowance. Actual usage varies with the selected models, reasoning effort, cache use, output length, and retries.
Each configuration is an exact combination of model, reasoning effort, account, and endpoint. Quick Comparison evaluates the current and recommended configurations, Full Scan evaluates every enabled configuration, and a custom selection applies only to that run. Each selected configuration receives the same fixed five-task evaluation.
The current coding-fast-v4.10 pack contains five equally weighted tasks. Each task exposes 20 independently scored behaviors or failure modes, for 100 scoring points in total. A point is earned only when the submitted tests or counterexamples distinguish the reference behavior from the corresponding faulty behavior. The total measures evidence coverage, not a pass probability or a claim that the model will succeed on the same percentage of real repositories.
Task 1 tests black-box contract and boundary analysis across atomic saves, snapshots, deterministic archives, and replay. Task 2 tests counterexample design for retry selection, readiness, capacity, ordering, and wake-up logic. Task 3 tests adversarial CI-plan auditing across dependency propagation, fallback coverage, scheduling, cost, and ranking. Task 4 tests transaction state-machine reasoning across rejection, wave snapshots, conflicts, retries, and tombstones. Task 5 tests cache regression design across identity, invalidation, forced and warm scans, issue reporting, and state preservation.
No. ModelDial evaluates a fixed task set rather than your current repository. Local Codex evaluations run in an ephemeral, read-only workspace. API routes receive only the fixed evaluation request through the provider endpoint you configured.
ModelDial compares complete, compatible results produced with the same question pack and scoring protocol. Your selected goal determines how eligible alternatives are ranked, while an explicit quality guardrail limits the tradeoff. If evidence is incomplete or no alternative provides a meaningful benefit, ModelDial keeps the current configuration or asks for another test.
Completed evidence is retained. You can retry only the failed evaluations instead of restarting the entire batch. Retries may consume additional provider quota or incur additional API charges.
Yes. Add custom endpoints and their model candidates in Settings, or select a custom configuration set for a single run. ModelDial keeps each account and endpoint distinct, so evidence from different routes is never merged into one result.
The signed build will appear here after release checks. Until then, the public Radar is ready to explore.