Guides and evidence

Answers for choosing, running, and changing an AI coding model

Each page starts from a real decision: which configuration to try, what an effort label means, why a stream ended early, or what the latest measured result can and cannot prove.

Where to start
  • Choosing a model? Start with the selection guide and comparisons.
  • Balancing speed and cost? Learn how reasoning levels differ.
  • Seeing timeouts or score changes? Read the troubleshooting guides.

Decision and reliability guides

Guide

How to detect AI model regressions without overreacting

Separate one noisy batch from a repeated protocol-matched decline, then verify the affected route locally.

Open page
Guide

There is no best coding model without a configuration

A model name leaves out two variables that often change the result: reasoning effort and route. Compare the configuration you can actually run, not the brand in isolation.

Open page
Guide

Low, High, and Max are settings, not guarantees

Reasoning effort is useful only when its model and route stay attached. The label is not a portable unit of intelligence across providers.

Open page
Guide

Streaming improves visibility. It does not remove timeouts.

A stream can show that generation is still moving, but the client, proxy, and provider still keep their own deadlines. Treat a stream without its terminal event as incomplete.

Open page
Guide

A cheaper model is not cheaper if it misses the quality floor

Quality, elapsed time, and reference cost answer different questions. Keep their denominators visible before turning them into one recommendation.

Open page

Current Backend & Testing evidence

Model evidence

GPT-5.6 Sol, Terra, and Luna

Compare Sol, Terra, and Luna in the same test batch, including the score and completion time at each effort level.

Open page
Model evidence

GPT-5.6 Sol

Compare GPT-5.6 Sol scores, completion times, and reference costs to see whether a higher effort level is worth it.

Open page
Model evidence

Qwen 3.8

Read the measured Qwen 3.8 Max configuration without treating missing cost or an embedded Max label as a universal effort scale.

Open page
Model evidence

DeepSeek V4

Compare DeepSeek V4 Flash and Pro at each tested effort level, using results from the same batch.

Open page
Model evidence

Kimi K3

See Kimi K3 coding scores and question-level results, along with the effort level and connection used for the test.

Open page
Model evidence

GLM-5.3

Compare GLM-5.3 High and Max to see how the score and completion time change at a higher effort level.

Open page

Direct configuration comparisons

Comparison

GPT-5.6 Sol vs Qwen 3.8

Compare the highest-scoring GPT-5.6 Sol and Qwen 3.8 effort levels from the same test batch.

Open page
Comparison

Qwen 3.8 vs Kimi K3

Compare the highest-scoring Qwen 3.8 and Kimi K3 effort levels from the same test batch.

Open page
Comparison

Kimi K3 vs GLM-5.3

When totals are close, the five-question profile and repeated batches matter more than declaring a winner from one point.

Open page