First-party model readiness radar
Models regress. Don't guess.ModelDial tells you when to switch.
ModelDial verifies on a schedule and publishes new evidence after it passes validation. The overall ranking combines the latest verified result from each capability.
Overall first. Three abilities underneath.
Overall score weights Backend & Testing 50%, Frontend & Interaction 30%, and Knowledge & Reasoning 20%. Only complete configurations enter the ranking.
- Overall score
- Pending all three axes
Overall score
A weighted result: 50% Backend & Testing, 30% Frontend & Interaction, and 20% Knowledge & Reasoning
Published after verification
See the change, not just the rank
Each Backend & Testing publication is checked against the previous comparable result, separating rank movement, score changes, and recommendation changes from ordinary reshuffling.
Capability leaderboards askHow capable is this model in general?
ModelDial rerunsWhich exact configuration is ready for my next work block?
Public evidence narrows the field. Local scans verify your route.
The public Radar compares three capabilities. A local scan applies the Backend & Testing method to your chosen route and keeps keys, history, and evidence on your Mac.
Comparable evidence
Verified before publication
Must hold
Requests go directly through the route you chose.
Keys, configuration, scan history, and recommendations remain local.
Read ModelDial with confidence
The three capabilities, 100-point scores, update state, reference cost, and the boundary between public and local evidence.
How is ModelDial different from other model benchmarks?
Most benchmarks ask how capable a model is in general. ModelDial publishes ModelDial-run, verified evidence for exact model, effort, and route configurations across three separate capabilities. Use the public Radar as a shortlist, then verify your own account or route locally instead of treating one ranking as universal advice.
What do the three capability scores measure?
Backend & Testing covers code implementation, debugging, and test design. Frontend & Interaction covers interface implementation, interaction details, and visual results. Knowledge & Reasoning covers information understanding, synthesis, and complex problem solving. Each capability is ranked separately.
What does a 100-point score mean, and when is the overall score available?
Each capability uses a protocol-bound 100-point score for that evaluation axis. It is not the chance of succeeding on real work and should not be compared across unrelated benchmarks. The overall score weights Backend & Testing 50%, Frontend & Interaction 30%, and Knowledge & Reasoning 20%, and is calculated only after all three capabilities have verified results for the same configuration. Missing capabilities remain pending and are never counted as zero.
Why is a capability pending, or still showing an earlier result?
ModelDial publishes a capability only after its evidence passes verification. A capability with no verified result stays pending. If a newer batch is not ready, the page keeps the prior verified result instead of inventing a score or resetting it to 100. The timestamp is publication time rather than job start time.
Is Radar's reference cost my actual bill?
No. Radar's reference cost is the approximate cost of a comparable reference run, not your actual bill. Your actual bill depends on your account, caching, and custom route.
What is the difference between Public Radar and a local scan?
Public Radar uses ModelDial-controlled runs across three capabilities and uses no quota from your provider account. A local scan applies the Backend & Testing method to your own account, route, or endpoint only when you start it. Keys, configuration, scan history, and recommendations remain on your Mac. Website analytics uses a first-party random browser cookie, expiring after 180 days, and honors GPC and DNT.