Skip to content

Instantly share code, notes, and snippets.

@magnus919
Created August 2, 2026 03:01
Show Gist options
  • Select an option

  • Save magnus919/52d9b0554be656e4e18b6fec94edfab1 to your computer and use it in GitHub Desktop.

Select an option

Save magnus919/52d9b0554be656e4e18b6fec94edfab1 to your computer and use it in GitHub Desktop.
GPT-5.6 Sol, Terra, Luna, and DeepSeek V4 Flash: benchmarked cost-effectiveness

GPT-5.6 Sol, Terra, and Luna: which one buys the most completed work?

Research date: 2026-08-01

Executive answer

The strongest currently available public evidence does not support using Terra as the default solely for cost-effectiveness. Artificial Analysis, which evaluated all three GPT-5.6 tiers across reasoning efforts, reports that Luna and Sol are always on its intelligence-versus-cost Pareto frontier ahead of Terra. In its words: for any Terra effort level, a Luna or Sol configuration is either more intelligent at no added cost or equally intelligent at lower cost.

That is not a blanket instruction to replace Terra with Luna-max. The main trade-off is latency: in a direct current comparison, Luna-max scores 51 on Artificial Analysis’s Intelligence Index versus 46 for Terra-medium, while costing less and generating tokens faster, but its time to first token is 126.28 seconds versus Terra-medium’s 1.60 seconds. For an interactive task where that first response matters, Terra-medium can still be a rational latency choice. For quality-per-dollar and throughput, the evidence favors Luna or Sol.

DeepSeek V4 Flash 0731 is now a credible fourth option, not merely a preview rumor. DeepSeek’s official 2026-07-31 changelog calls its V4-Flash API release a public beta. Artificial Analysis measures its max-effort reasoning configuration at 50, one point below Luna-max, with an estimated $0.03 Intelligence-Index cost per task versus Luna-max’s $0.07 on the same comparison page. That is compelling but still immature evidence: it is a brand-new public beta, text-only in Artificial Analysis’s comparison, and its most impressive coding/agent results are vendor-published rather than independently reproduced.

The decision being tested

Claim: Terra is not cost-effective, so use Luna at a higher reasoning level instead.

Verdict: Supported for aggregate quality-per-dollar; complicated for interactive latency and unvalidated on a specific workload.

The evidence supports retiring Terra as the automatic economic middle ground. It does not prove that Luna-max is the best setting for every prompt, agent loop, tool configuration, or latency target.

What the models cost now

OpenAI’s official current API rates, effective from 2026-07-30:

Model Input / 1M tokens Cached input / 1M Output / 1M Positioning
GPT-5.6 Sol $5.00 $0.50 $30.00 Frontier tier for complex professional work
GPT-5.6 Terra $2.00 $0.20 $12.00 Balance of intelligence and cost
GPT-5.6 Luna $0.20 $0.02 $1.20 Cost-sensitive, high-volume work

Token price is not task cost. Higher reasoning can consume more output/reasoning tokens, and a cheaper model can cost more per completed task if it takes more attempts, produces a weaker plan, or requires repair. The useful metric is:

total cost of all attempts, checks, and repair work / successful completed task

Best like-for-like public evidence: Artificial Analysis

Artificial Analysis’s Intelligence Index v4.1 is a weighted English-language, text-focused composite of nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. Its weights are agents 34%, coding 24%, scientific reasoning 24%, and general capability 18%.

Family comparison at max reasoning

The following is Artificial Analysis’s published snapshot. Its token prices in the original article predate OpenAI’s 2026-07-30 Terra/Luna price cut, so these task-cost figures should be read as measured-snapshot comparisons, not today’s price quote.

Configuration Intelligence Index Coding Agent Index Published cost per Intelligence-Index task Interpretation
Sol (max) 59 80 $1.04 Highest capability in the family
Terra (max) 55 77 $0.55 Cheaper, but not on AA’s cost-intelligence frontier across efforts
Luna (max) 51 75 $0.21 Lower absolute score, much lower measured task cost

A purely mechanical score-per-dollar calculation from that historical snapshot is 56.73 index points per dollar for Sol-max, 100.00 for Terra-max, and 242.86 for Luna-max. Do not treat Intelligence Index points as a linear unit of business value. The calculation merely illustrates why token prices alone underestimate Luna’s economic advantage.

The key cross-effort comparison

Artificial Analysis currently compares Luna-max directly with Terra-medium:

Metric Luna (max) Terra (medium) Better for
Intelligence Index 51 46 Luna-max
Listed price per 1M tokens in comparison $0.17 $1.74 Luna-max
Output speed 175 tokens/s 99 tokens/s Luna-max
Time to first token 126.28 s 1.60 s Terra-medium
Context window 1M 1M Tie

This directly refutes the simplistic claim that “more reasoning on Luna must always be worse in responsiveness.” Luna-max is faster once generation begins, but it is much slower to start because it spends longer reasoning. It supports a more precise rule:

  • Use Luna-max where a slow start is acceptable and quality-per-dollar matters.
  • Use Terra-medium only when fast first response has independent value that outweighs Luna’s score/cost/throughput advantage.
  • Use Sol where the cost of a materially worse plan, diagnosis, or judgment exceeds the extra model cost.

Why the price cut strengthens, but does not prove, the conclusion

OpenAI cut Terra’s price by 20% and Luna’s by 80% on 2026-07-30. If the token use in Artificial Analysis’s published max-effort runs stayed exactly the same, a mechanical adjustment would move its published task-cost estimates from $0.55 to about $0.44 for Terra-max and from $0.21 to about $0.042 for Luna-max. Those are not newly measured benchmark costs, only a transparent rate-change calculation. They show the cut widened the economic gap in Luna’s favor.

What OpenAI itself reports

OpenAI’s claims are useful but vendor-reported:

  • On Agents’ Last Exam, Sol at medium reasoning reportedly beats Claude Fable 5 by 11.4 points at roughly one-quarter the estimated cost. OpenAI says Terra and Luna outperform Fable 5 at around one-sixteenth the estimated cost, but does not name their reasoning levels in that sentence.
  • On the Artificial Analysis Coding Agent Index, Sol-max reportedly scores 80. OpenAI says Terra and Luna each operate at roughly one-quarter the estimated cost of the cited competing configuration, with about half as many output tokens and roughly one-third the time.
  • OpenAI proposes a workflow split that matches the available evidence: Sol to resolve uncertainty and define the plan; Luna to implement well-specified changes, test, and evaluate results.

These claims align with the independent Artificial Analysis direction, but they are not substitutes for a controlled test on a real workload.

DeepSeek V4 Flash 0731: real, newly released, and worth testing

The premise about a very recent release is substantially correct. DeepSeek’s official changelog, dated 2026-07-31, says the official DeepSeek-V4-Flash API is now in public beta. It says 0731 retains the preview architecture and size, and is only re-post-trained.

Evidence table

Configuration Intelligence Index Cost per Index task What is independently comparable
DeepSeek V4 Flash 0731, reasoning max 50 $0.03 Artificial Analysis comparison
GPT-5.6 Luna, max 51 $0.07 Same comparison page
GPT-5.6 Terra, medium 46 $0.16 Artificial Analysis comparison
GPT-5.6 Sol, max 59 $1.86 in the current comparison ranking Same Artificial Analysis measurement family, but different config/cost basis than the launch snapshot

The first three rows support a limited, important finding: V4 Flash 0731-max is one Intelligence Index point behind Luna-max but costs about 60% less per measured Index task. Its cheap cached input, $0.0028 per million tokens versus $0.14 uncached input, is a major driver.

DeepSeek’s own 2026-07-31 release note reports these max-effort, DeepSeek-Harness-minimal results: Terminal Bench 2.1 82.7, NL2Repo 54.2, Cybergym 76.7, DeepSWE 54.4, Toolathlon verified 70.3, Agents’ Last Exam 25.2, and Automation Bench (Public) 25.1. These figures are useful discovery evidence, but DeepSeek controls the harness and report. They are not an independent head-to-head basis for deciding that it beats any GPT-5.6 configuration.

Artificial Analysis independently reports V4 Flash 0731 at 50, up from the April V4 Flash’s 40, including GDPval-AA v2 Elo 1559, Terminal-Bench 2.1 at 79%, τ³-Banking at 31%, SciCode at 50%, Humanity’s Last Exam at 37%, AA-LCR at 66%, and GPQA Diamond at 91%. That makes it a serious budget contender, not proof that it is operationally interchangeable with OpenAI’s family.

Sweet spot: a decision, not one universal model

Work type First choice Why Escalate when
High-volume, bounded, text-heavy execution where startup delay is acceptable Luna at high/max Strongest GPT-5.6 quality-per-dollar evidence; Luna-max beats Terra-medium on AA’s aggregate score The task requires materially better planning/judgment or repeats/failures erase savings
Interactive work needing immediate first response Terra at medium, or Luna at lower effort after testing Terra-medium’s measured TTFT is far lower than Luna-max’s Quality problems or expensive repair appear
Ambiguous, high-stakes planning, hard debugging, architecture, or consequential synthesis Sol, with effort matched to stakes Highest family capability; lower downstream risk of a bad plan The task becomes well specified and independently checkable
Cheap external challenger for text-only, controlled agent/coding trials DeepSeek V4 Flash 0731 max 50 Index points, roughly 60% less measured cost per task than Luna-max Image input, reliability, compliance, tool/harness behavior, or real-work error rate matter more than benchmark cost

What public benchmarks cannot settle

No source found provides a complete, independently replicated matrix of Sol/Terra/Luna at low, medium, high, xhigh, and max on the same tasks with raw token logs and repair loops. That missing matrix is the actual answer to “what does it cost to get the same task done?”

The public evidence is strong enough to change the default hypothesis from “Terra is the economic middle” to “Terra needs a latency-specific justification.” It is not strong enough to deploy a universal routing policy without local measurement.

A decision-grade test to run next

Build a small task set from real, non-sensitive work, stratified into:

  1. bounded implementation and tests,
  2. repository diagnosis,
  3. research/synthesis with source checks,
  4. interactive short-turn requests,
  5. a deliberately ambiguous planning task.

Run each task at least three times with fixed prompts, tools, maximum wall time, and acceptance criteria. Compare Sol-medium, Terra-medium, Luna-high, Luna-max, and V4 Flash 0731 max. Record:

Metric Why it matters
Task success / acceptance rate The denominator for cost per completed task
Total billable input, cache-read, cache-write, reasoning, and output tokens Actual spend, not list-price intuition
Tool calls and retries Agent-loop costs that token comparisons hide
Human repair minutes A cheap wrong result is not cheap
Time to first token and time to verified completion Separates interactive latency from throughput
Failure mode Shows where a routing rule is unsafe

Calculate all-in model cost / accepted outcomes and separately report human repair time. Keep the models that form the actual Pareto frontier for the work, and remove any routing rule that is consistently dominated.

Evidence quality and limits

  • High confidence: Current OpenAI prices and model positioning, because they are in OpenAI documentation; current DeepSeek 0731 public-beta status and vendor benchmark settings, because they are in DeepSeek’s changelog.
  • High confidence within its declared scope: Artificial Analysis’s published score/cost comparisons and method. It is the strongest independent cross-model source found.
  • Medium confidence for operational routing: Transfer from public benchmarks to a particular agent stack or task mix. Harness, prompts, tool configuration, retry policy, caches, and human review can change the cost per accepted task.
  • Low confidence: Any claim that V4 Flash 0731 is already generally “better” than GPT-5.6. The evidence shows a close budget comparison with Luna-max and a clear advantage over Terra-medium in this one aggregate Index; it does not establish universal superiority.

Sources

  1. OpenAI, GPT-5.6 model family, accessed 2026-08-01.
  2. OpenAI, Advancing the price-performance frontier with GPT-5.6, accessed 2026-08-01.
  3. OpenAI, GPT-5.6 Sol model documentation, accessed 2026-08-01.
  4. OpenAI, GPT-5.6 Terra model documentation, accessed 2026-08-01.
  5. OpenAI, GPT-5.6 Luna model documentation, accessed 2026-08-01.
  6. OpenAI, Reasoning models guide, accessed 2026-08-01.
  7. Artificial Analysis, GPT-5.6 benchmarks across Intelligence, Speed and Cost, accessed 2026-08-01.
  8. Artificial Analysis, Luna-max versus Terra-medium, accessed 2026-08-01.
  9. Artificial Analysis, Intelligence benchmarking methodology, accessed 2026-08-01.
  10. DeepSeek, API changelog, 2026-07-31 V4 Flash update, accessed 2026-08-01.
  11. Artificial Analysis, DeepSeek V4 Flash 0731 scores 50, accessed 2026-08-01.
  12. Artificial Analysis, V4 Flash 0731-max versus Luna-max, accessed 2026-08-01.
  13. Artificial Analysis, V4 Flash 0731-max versus Terra-medium, accessed 2026-08-01.

Research log: GPT-5.6 family cost-effectiveness

Research date: 2026-08-01 Question: Which GPT-5.6 family configuration offers the best verified capability, latency, and cost trade-off, and does new DeepSeek V4 Flash 0731 change the choice?

Search history

Query / retrieval path Result Decision
Official OpenAI GPT-5.6 family, model pages, reasoning guide, and price-cut announcement Current prices, positioning, public vendor benchmark claims Retained as primary vendor documentation
Artificial Analysis GPT-5.6 launch article and model comparison pages Cross-tier scores, task-cost snapshots, Pareto-frontier claim, methodology Retained as strongest independent comparison source
DeepSeek official API changelog and V4 announcement 0731 public-beta status, harness/configuration, vendor benchmark claims Retained as primary vendor documentation
Artificial Analysis V4 Flash 0731 article and comparison pages Independent 50 Index score, cost/task comparison to Luna and Terra Retained
CodeRabbit Sol/Terra coding blog Practical agent runs, but no Luna and internally inconsistent Terra pass-rate figures Rejected from conclusions; not used as a quantitative source
BenchLM V4 Flash versus Terra page Transparently labels sparse/incomplete coverage and does not support an overall winner Rejected from conclusions; useful only as a warning against flattening uneven benchmark sets
LiveBench / classic SWE-bench / complete cross-effort score matrix No current, configuration-matched public data located Recorded as an evidence gap

Retained source ledger

Source Material evidence preserved in Why it matters
OpenAI GPT-5.6 family announcement Overview sections “What OpenAI itself reports” and “Sources” Vendor-reported family benchmarks and availability
OpenAI price-performance announcement “What the models cost now” and post-cut caveat Official July 30 price change and model-routing example
OpenAI Sol, Terra, Luna model documentation “What the models cost now” Exact current API rates and tier positioning
Artificial Analysis GPT-5.6 article “Best like-for-like public evidence” Max-effort cross-tier score/cost snapshot and Pareto claim
Artificial Analysis Luna-max vs Terra-medium “The key cross-effort comparison” Direct configuration-matched score, cost, speed, and TTFT values
Artificial Analysis methodology “Best like-for-like public evidence” Index composition, weighting, and scope limits
DeepSeek official changelog “DeepSeek V4 Flash 0731” July 31 public beta and vendor test configuration/results
Artificial Analysis V4 Flash 0731 article “DeepSeek V4 Flash 0731” Independent score, price, and per-benchmark improvements
Artificial Analysis V4 Flash vs Luna/Terra comparisons “DeepSeek V4 Flash 0731” Direct score/cost comparisons

Evidence gaps explicitly not papered over

  1. No complete independent score-and-token-log matrix of Sol/Terra/Luna at every reasoning level.
  2. No public controlled experiment measuring retries, tool calls, human repair, and accepted outputs on one real workflow.
  3. No independent reproduction of DeepSeek’s July 31 agent/coding benchmark claims with an OpenAI-equivalent harness.
  4. Artificial Analysis’s published $1.04/$0.55/$0.21 max-effort task-cost snapshot predates OpenAI’s July 30 price cut. Any rate-adjusted figures in the overview are labeled mechanical estimates, not remeasured results.

Publication decision

The overview makes no claim that a single model wins universally. It states a source-supported default: Terra needs a latency-specific justification; Luna and Sol dominate Terra on Artificial Analysis’s published intelligence-cost frontier. It recommends a controlled local acceptance-cost test before a routing-policy change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment