Guides and evidence

Answers for choosing, running, and changing an AI coding model

Each page starts from a real decision: which configuration to try, what an effort label means, why a stream ended early, or what the latest measured result can and cannot prove.

What belongs here
  • A direct answer with an explicit evidence boundary.
  • Current ModelDial data where the question depends on a benchmark.
  • No model claim that cannot be traced to a configuration and batch.

Decision and reliability guides

Guide

How to detect AI model regressions without overreacting

Separate one noisy batch from a repeated protocol-matched decline, then verify the affected route locally.

Open page
Guide

There is no best coding model without a configuration

A model name leaves out two variables that often change the result: reasoning effort and route. Compare the configuration you can actually run, not the brand in isolation.

Open page
Guide

Low, High, and Max are settings, not guarantees

Reasoning effort is useful only when its model and route stay attached. The label is not a portable unit of intelligence across providers.

Open page
Guide

Streaming improves visibility. It does not remove timeouts.

A stream can show that generation is still moving, but the client, proxy, and provider still keep their own deadlines. Treat a stream without its terminal event as incomplete.

Open page
Guide

A cheaper model is not cheaper if it misses the quality floor

Quality, elapsed time, and reference cost answer different questions. Keep their denominators visible before turning them into one recommendation.

Open page

Current Backend & Testing evidence

Model evidence

GPT-5.6 Sol, Terra, and Luna

Compare the three GPT-5.6 variants inside one published batch instead of treating Sol, Terra, and Luna as interchangeable names.

Open page
Model evidence

GPT-5.6 Sol

Compare GPT-5.6 Sol effort levels without collapsing them into one model score.

Open page
Model evidence

Qwen 3.8

Read the measured Qwen 3.8 Max configuration without treating missing cost or an embedded Max label as a universal effort scale.

Open page
Model evidence

DeepSeek V4

Keep Flash, Pro, and their effort levels separate while comparing them inside one published batch.

Open page
Model evidence

Kimi K3

Inspect the measured Kimi K3 configuration and keep its effort and custom-endpoint route attached to the result.

Open page
Model evidence

GLM-5.3

Compare measured GLM-5.3 effort levels without turning one custom endpoint into a universal provider result.

Open page

Direct configuration comparisons

Comparison

GPT-5.6 Sol vs Qwen 3.8

A same-batch comparison of the highest-scoring measured configuration from each family, with route and missing-cost boundaries left visible.

Open page
Comparison

Qwen 3.8 vs Kimi K3

A same-batch comparison of the current measured configurations, with their route and missing-cost boundaries kept visible.

Open page
Comparison

Kimi K3 vs GLM-5.3

When totals are close, the five-question profile and repeated batches matter more than declaring a winner from one point.

Open page