Model selection

There is no best coding model without a configuration

A model name leaves out two variables that often change the result: reasoning effort and route. Compare the configuration you can actually run, not the brand in isolation.

Decision rule

Set a quality floor first. Among configurations that clear it, compare successful elapsed time and reference cost. Before changing your default, rerun the candidate through the account and route you will use in practice.

Compare model, effort, and route together

GPT-5.6 Sol at Low and GPT-5.6 Sol at Max are separate configurations. The same is true when one request uses an official login and another uses a custom endpoint. A ranking that drops those fields may look simpler, but it no longer describes the thing you will run.

Use one published batch when making the first comparison. It keeps the comparison standard consistent, so a score difference has a known frame.

Choose the quality floor before looking at speed

Decide what failure you cannot accept. A configuration with the fastest total time is not useful if it misses the capability your work depends on. Read the five question scores, not only the total, when one capability matters more than the others.

  • For a main coding agent, start with the minimum acceptable total and check for a weak individual question.
  • For a worker used on bounded tasks, a lower effort may be reasonable after it clears the same task-specific floor.
  • Do not turn a one-point lead in one batch into a permanent winner. Check repeated comparable batches when the margin is small.

Read elapsed time and cost as observed evidence

Elapsed time includes the behavior of the measured configuration and route. It is evidence from that run, not a vendor-wide speed guarantee.

A missing reference cost means the pricing or usage coverage was not sufficient for that configuration. It does not mean the request was free. Compare cost only when both sides have coverage under the same pricing snapshot.

Switch only after a route-level check

The public Radar answers which first-party configurations performed better under its protocol. It cannot see your repository, regional network, account limits, prompt shape, or endpoint proxy.

Use the Radar to select one candidate. Then run that candidate locally against the route you plan to keep. A switch is justified when the candidate clears the quality floor and the operational gain still exists on your setup.