ModelDial · Integrations
API & Skill
Use ModelDial results in your own tools. Install the Skill for an AI assistant, or read the published data through the API.
Install the Skill
Send the prompt below to a tool that supports Agent Skills, or install from the terminal. The Skill reads public results and requires no account or API key.
Install the ModelDial Skill from https://modeldial.com/modeldial-skill/SKILL.md. Review the linked files before writing to the skills directory.Install from the terminal
bash <(curl -fsSL https://modeldial.com/modeldial-skill/install.sh) --target agentsInstall into the directory your agent already uses
The installer checks the files and updates an existing ModelDial Skill if needed.
Start a new conversation
Most tools discover Skills when a session starts. Reopen the tool or begin a fresh task after installation.
Verify with one question
Compare the published results for GPT-6.1 Sol and GPT-6 Sol, including their effort levels, axis scores, evaluation time and reference cost. Include source links and identify missing data.
The answer should includeThe model and effort level, test date or batch, supporting scores, and links to the ModelDial results.
If the data cannot be readThe Skill instructs the agent to report the failure instead of answering from older results or memory.
Data endpoints
Use compact JSON for current rankings. For audits, choose the Backend & Testing index or the three-capability publication index.
- GET / JSON
Current compact ranking
/api/v1/radar/latest.jsonDefault overall order with all three axis scores and weights, plus the compatible backend ranking.
- GET / JSON
Agent configuration profiles
/api/v1/radar/agent-profile.jsonBalanced, quality, and value role assignments from published decision tags; not a paired-agent benchmark.
- GET / JSON
Latest comparable changes
/api/v1/radar/changes.jsonBackend rank, score, time, and recommendation changes after a protocol compatibility check; cost deltas appear only when the pricing snapshot matches.
- GET / JSON
Backend & Testing publication index
/api/v1/radar/index.jsonBackend batch identity, publication time, protocol fields, hashes, and complete snapshot links.
- GET / JSON
Capability publication index
/data/benchmark-snapshots/index.jsonLatest verified batch identity and archive link for each capability and the weighted overall ranking.
- GET / JSON
Complete Backend & Testing snapshot
/data/reference-snapshots/latest.jsonQuestion-level Backend & Testing results and provenance fields for audit work.
Current results
Example from current results
Published after verification
Example questions and interpretation limits
Ask for the decision you need
Ask about a new model, its predecessor or tested efforts. The Skill retrieves published scores and sources; historical backend changes are available separately when comparable records exist.
- 01Current ranking
Which configuration leads the current Radar, and why?
- 02New model vs predecessor
How do GPT-6.1 Sol and its predecessor GPT-6 Sol compare at their published efforts?
- 03Your configuration
I use GPT-6.1 Sol XHigh through the official login. Which alternatives have similar scores and shorter evaluation times?
- 04Agent configuration
How should I configure my main and worker agents from the latest ModelDial results? Start with the balanced profile, then show quality and value alternatives.
Answer rules and data sources
The Skill instructs the assistant to cite published results and identify missing data. Check the sources in its answer.
Separate score differences from historical trends
Current model comparisons use published scores. Historical backend movement requires matching protocols and is not an overall-ranking trend.
Start from your configuration
If the model, effort, or route is missing, the Agent asks before turning the top row into personal advice.
Treat a drop as a signal
One lower comparable score can justify another check, but not a claim that a provider changed the model.
Keep acceptance with the main agent
The worker gets bounded execution. Requirements, integration, and final acceptance stay with the main agent.
What the Skill will not claim
Use the published ranking
Publisher order is the ranking. The Skill does not rebuild it from score or provider names.
Separate score changes from their causes
A lower comparable score is a signal, not proof that a provider changed or degraded a model.
Confirm your settings
Without your exact configuration, the Skill reports the field or asks one short follow-up question.
Distinguish estimates from charges
Reference cost describes the controlled batch, not your invoice or actual usage charge.
Explain how role suggestions were derived
Role profiles use configurations evaluated independently across the latest published axis results, not a paired-agent benchmark.
Licensing and attribution
Skill instructions use the MIT License. Published Radar data uses CC BY 4.0 and should identify ModelDial and the source batch when practical.