AI coding models, tested

Models regress. Don't guess.ModelDial tells you when to switch.

Use test results to choose a model and reasoning level. If needed, test your own connection locally on your Mac.

ModelDial Radar

Top five model scores

Explore all results
Published

Overall score

A weighted result: 40% Backend & Testing, 30% Frontend & Interaction, and 30% Knowledge & Reasoning

  1. 01OpenAIgpt-6-astraXHigh
    95.7/ 100
    Backend
    96.0
    Frontend
    96.0
    Reasoning
    95.0
  2. 02Anthropicclaude-fable-5-1Max
    89.0/ 100
    Backend
    83.0
    Frontend
    91.0
    Reasoning
    95.0
  3. 03Anthropicclaude-opus-5Max
    81.0/ 100
    Backend
    66.0
    Frontend
    87.0
    Reasoning
    95.0
  4. 04xAIgrok-4.6XHigh
    79.5/ 100
    Backend
    75.0
    Frontend
    80.0
    Reasoning
    85.0
  5. 05OpenAIgpt-5.6-solMax
    78.0/ 100
    Backend
    84.0
    Frontend
    73.0
    Reasoning
    75.0

Each model shown at its highest-scoring tested level

Tested configurations
51 configurations
Latest verification
Sep 16, 2026, 2:28 PM

Current overall leadergpt-6-astraXHigh·All three capabilities verified

How we measure

One overall score. Three capabilities.

Overall score weights Backend & Testing 40%, Frontend & Interaction 30%, and Knowledge & Reasoning 30%. Only complete configurations enter the ranking.

Backend & Testing

40%

Code, contracts, debugging, and test design

Frontend & Interaction

30%

Interface implementation and visual behavior

Knowledge & Reasoning

30%

Research depth, synthesis, and hard reasoning

Each Backend & Testing publication is checked against the previous comparable result, separating rank movement, score changes, and recommendation changes from ordinary reshuffling.

Read the evaluation method

Local verification

Compare publicly. Test locally if needed.

Use the public rankings to compare three capabilities. If you want to check your own account or connection, run the Backend & Testing evaluation on your Mac. Keys and test history stay local.

Consistent comparison

Test results
Verified before publication

Quality guardrail

Must hold

Your provider

Requests go directly through the route you chose.

Your Mac

Keys, configuration, scan history, and recommendations remain local.

Read the evaluation method

Read ModelDial with confidence

Learn how scores work, when rankings update, and when a local test may help.

How is ModelDial different from other model benchmarks?

ModelDial tests models at different reasoning levels and shows their scores, completion times, and reference costs. Use the public rankings to shortlist models. If you want to check whether your own account or connection performs differently, you can run a local test.

What do the three capability scores measure?

Backend & Testing covers code implementation, debugging, and test design. Frontend & Interaction covers interface implementation, interaction details, and visual results. Knowledge & Reasoning covers information understanding, synthesis, and complex problem solving. Each capability is ranked separately.

What does a 100-point score mean, and when is the overall score available?

Each capability has its own scoring rules, with a maximum of 100 points. Scores are not real-world success rates or directly comparable with unrelated benchmarks. The overall score weights Backend & Testing at 40%, Frontend & Interaction at 30%, and Knowledge & Reasoning at 30%. A model and effort level enter the overall ranking only when all three results are verified. Missing scores are never counted as zero.

Why is a capability pending, or still showing an earlier result?

Results appear after they have been checked. Until then, the page shows the previous verified score, or a pending label if there is no earlier result. The timestamp shows when a result was published, not when the test started.

Is Radar's reference cost my actual bill?

No. Radar's reference cost is the approximate cost of a comparable reference run, not your actual bill. Your actual bill depends on your account, caching, and custom route.

What is the difference between Public Radar and a local scan?

Public Radar uses ModelDial-controlled runs across three capabilities and uses no quota from your provider account. A local scan applies the Backend & Testing method to your own account, route, or endpoint only when you start it. Keys, configuration, scan history, and recommendations remain on your Mac. Website analytics uses a first-party random browser cookie, expiring after 180 days, and honors GPC and DNT.