Model results

GPT-5.6 Sol, Terra, and Luna

Compare Sol, Terra, and Luna in the same test batch, including the score and completion time at each effort level.

Current batch

GPT-5.6 Sol / Max has the highest measured score at 84/100. GPT-5.6 Luna / Low is the fastest measured configuration at 7m 40s. Choose based on whether quality or speed matters more for your task.

Evaluation details
Published
Sep 16, 2026, 06:28 AM UTC
Question pack
coding-fast-v4.12
Tested
15 measured configurations
Batch
evaluation-497e806be8b8470345169c1aad6752d74a0c1ce9e195a9eb7fcc02b628701c77

Measured configurations

ConfigurationRankScoreElapsedReference costRoute
GPT-5.6 Sol / Max#684/10034m 23s$1.57Custom endpoint
GPT-5.6 Sol / XHigh#975/10035m 6s$1.22Custom endpoint
GPT-5.6 Sol / High#1173/10013m 36s$0.66Custom endpoint
GPT-5.6 Terra / Max#1273/10036m 23s$1.34Custom endpoint
GPT-5.6 Sol / Medium#1866/10028m 4s$0.74Custom endpoint
GPT-5.6 Sol / Low#1965/10010m 43s$0.53Custom endpoint
GPT-5.6 Luna / Max#2463/10041m 48s$0.16Custom endpoint
GPT-5.6 Terra / High#3555/10027m 10s$0.42Custom endpoint
GPT-5.6 Terra / Medium#3853/1009m 16s$0.25Custom endpoint
GPT-5.6 Luna / XHigh#4053/10031m 15s$0.12Custom endpoint
GPT-5.6 Terra / XHigh#4251/10015m 5s$0.50Custom endpoint
GPT-5.6 Terra / Low#4644/10023m 23s$0.26Custom endpoint
GPT-5.6 Luna / Medium#4939/10010m 29s$0.0334Custom endpoint
GPT-5.6 Luna / High#5039/10017m 35s$0.0623Custom endpoint
GPT-5.6 Luna / Low#5133/1007m 40s$0.0237Custom endpoint

Five-question profile

QuestionGPT-5.6 Sol / MaxGPT-5.6 Sol / XHighGPT-5.6 Sol / HighGPT-5.6 Terra / MaxGPT-5.6 Sol / MediumGPT-5.6 Sol / LowGPT-5.6 Luna / MaxGPT-5.6 Terra / HighGPT-5.6 Terra / MediumGPT-5.6 Luna / XHighGPT-5.6 Terra / XHighGPT-5.6 Terra / LowGPT-5.6 Luna / MediumGPT-5.6 Luna / HighGPT-5.6 Luna / Low
Black-box Regression Audit14/2014/2016/2015/2012/2016/207/2012/208/2010/209/207/206/207/2010/20
Retry Planner Counterexamples18/2019/2017/2015/2016/2015/2019/2015/2015/2018/2015/2014/2010/2013/2012/20
CI Adversarial Audit15/2012/209/2011/2013/2012/2011/2011/208/209/2012/205/2011/2010/201/20
Transaction Regression Design17/2017/2019/2017/2013/2013/2017/2012/2013/2016/2012/209/2012/200/2010/20
Cache Propagation Certificate20/2013/2012/2015/2012/209/209/205/209/200/203/209/200/209/200/20

What the same-batch spread shows

The table keeps each available variant and effort level as a separate configuration. This makes it possible to compare Sol, Terra, and Luna like for like without assigning one family-wide score.

Compare the same effort first, then check whether a higher-scoring configuration still meets your elapsed-time requirement. Reference cost remains undecided when usage pricing coverage is absent.