Research date: 2026-08-01
The strongest currently available public evidence does not support using Terra as the default solely for cost-effectiveness. Artificial Analysis, which evaluated all three GPT-5.6 tiers across reasoning efforts, reports that Luna and Sol are always on its intelligence-versus-cost Pareto frontier ahead of Terra. In its words: for any Terra effort level, a Luna or Sol configuration is either more intelligent at no added cost or equally intelligent at lower cost.
That is not a blanket instruction to replace Terra with Luna-max. The main trade-off is latency: in a direct current comparison, Luna-max scores 51 on Artificial Analysis’s Intelligence Index versus 46 for Terra-medium, while costing less and generating tokens faster, but its time to first token is 126.28 seconds versus Terra-medium’s 1.60 seconds. For an interactive task where that first response matters, Terra-medium can still be a rational latency choice. For quality-per-dollar and throughput, the evidence favors Luna or Sol.
DeepSeek V4 Flash 0731 is now a credible fourth option, not merely a preview rumor. DeepSeek’s official 2026-07-31 changelog calls its V4-Flash API release a public beta. Artificial Analysis measures its max-effort reasoning configuration at 50, one point below Luna-max, with an estimated $0.03 Intelligence-Index cost per task versus Luna-max’s $0.07 on the same comparison page. That is compelling but still immature evidence: it is a brand-new public beta, text-only in Artificial Analysis’s comparison, and its most impressive coding/agent results are vendor-published rather than independently reproduced.
Claim: Terra is not cost-effective, so use Luna at a higher reasoning level instead.
Verdict: Supported for aggregate quality-per-dollar; complicated for interactive latency and unvalidated on a specific workload.
The evidence supports retiring Terra as the automatic economic middle ground. It does not prove that Luna-max is the best setting for every prompt, agent loop, tool configuration, or latency target.
OpenAI’s official current API rates, effective from 2026-07-30:
| Model | Input / 1M tokens | Cached input / 1M | Output / 1M | Positioning |
|---|---|---|---|---|
| GPT-5.6 Sol | $5.00 | $0.50 | $30.00 | Frontier tier for complex professional work |
| GPT-5.6 Terra | $2.00 | $0.20 | $12.00 | Balance of intelligence and cost |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | Cost-sensitive, high-volume work |
Token price is not task cost. Higher reasoning can consume more output/reasoning tokens, and a cheaper model can cost more per completed task if it takes more attempts, produces a weaker plan, or requires repair. The useful metric is:
total cost of all attempts, checks, and repair work / successful completed task
Artificial Analysis’s Intelligence Index v4.1 is a weighted English-language, text-focused composite of nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. Its weights are agents 34%, coding 24%, scientific reasoning 24%, and general capability 18%.
The following is Artificial Analysis’s published snapshot. Its token prices in the original article predate OpenAI’s 2026-07-30 Terra/Luna price cut, so these task-cost figures should be read as measured-snapshot comparisons, not today’s price quote.
| Configuration | Intelligence Index | Coding Agent Index | Published cost per Intelligence-Index task | Interpretation |
|---|---|---|---|---|
| Sol (max) | 59 | 80 | $1.04 | Highest capability in the family |
| Terra (max) | 55 | 77 | $0.55 | Cheaper, but not on AA’s cost-intelligence frontier across efforts |
| Luna (max) | 51 | 75 | $0.21 | Lower absolute score, much lower measured task cost |
A purely mechanical score-per-dollar calculation from that historical snapshot is 56.73 index points per dollar for Sol-max, 100.00 for Terra-max, and 242.86 for Luna-max. Do not treat Intelligence Index points as a linear unit of business value. The calculation merely illustrates why token prices alone underestimate Luna’s economic advantage.
Artificial Analysis currently compares Luna-max directly with Terra-medium:
| Metric | Luna (max) | Terra (medium) | Better for |
|---|---|---|---|
| Intelligence Index | 51 | 46 | Luna-max |
| Listed price per 1M tokens in comparison | $0.17 | $1.74 | Luna-max |
| Output speed | 175 tokens/s | 99 tokens/s | Luna-max |
| Time to first token | 126.28 s | 1.60 s | Terra-medium |
| Context window | 1M | 1M | Tie |
This directly refutes the simplistic claim that “more reasoning on Luna must always be worse in responsiveness.” Luna-max is faster once generation begins, but it is much slower to start because it spends longer reasoning. It supports a more precise rule:
- Use Luna-max where a slow start is acceptable and quality-per-dollar matters.
- Use Terra-medium only when fast first response has independent value that outweighs Luna’s score/cost/throughput advantage.
- Use Sol where the cost of a materially worse plan, diagnosis, or judgment exceeds the extra model cost.
OpenAI cut Terra’s price by 20% and Luna’s by 80% on 2026-07-30. If the token use in Artificial Analysis’s published max-effort runs stayed exactly the same, a mechanical adjustment would move its published task-cost estimates from $0.55 to about $0.44 for Terra-max and from $0.21 to about $0.042 for Luna-max. Those are not newly measured benchmark costs, only a transparent rate-change calculation. They show the cut widened the economic gap in Luna’s favor.
OpenAI’s claims are useful but vendor-reported:
- On Agents’ Last Exam, Sol at medium reasoning reportedly beats Claude Fable 5 by 11.4 points at roughly one-quarter the estimated cost. OpenAI says Terra and Luna outperform Fable 5 at around one-sixteenth the estimated cost, but does not name their reasoning levels in that sentence.
- On the Artificial Analysis Coding Agent Index, Sol-max reportedly scores 80. OpenAI says Terra and Luna each operate at roughly one-quarter the estimated cost of the cited competing configuration, with about half as many output tokens and roughly one-third the time.
- OpenAI proposes a workflow split that matches the available evidence: Sol to resolve uncertainty and define the plan; Luna to implement well-specified changes, test, and evaluate results.
These claims align with the independent Artificial Analysis direction, but they are not substitutes for a controlled test on a real workload.
The premise about a very recent release is substantially correct. DeepSeek’s official changelog, dated 2026-07-31, says the official DeepSeek-V4-Flash API is now in public beta. It says 0731 retains the preview architecture and size, and is only re-post-trained.
| Configuration | Intelligence Index | Cost per Index task | What is independently comparable |
|---|---|---|---|
| DeepSeek V4 Flash 0731, reasoning max | 50 | $0.03 | Artificial Analysis comparison |
| GPT-5.6 Luna, max | 51 | $0.07 | Same comparison page |
| GPT-5.6 Terra, medium | 46 | $0.16 | Artificial Analysis comparison |
| GPT-5.6 Sol, max | 59 | $1.86 in the current comparison ranking | Same Artificial Analysis measurement family, but different config/cost basis than the launch snapshot |
The first three rows support a limited, important finding: V4 Flash 0731-max is one Intelligence Index point behind Luna-max but costs about 60% less per measured Index task. Its cheap cached input, $0.0028 per million tokens versus $0.14 uncached input, is a major driver.
DeepSeek’s own 2026-07-31 release note reports these max-effort, DeepSeek-Harness-minimal results: Terminal Bench 2.1 82.7, NL2Repo 54.2, Cybergym 76.7, DeepSWE 54.4, Toolathlon verified 70.3, Agents’ Last Exam 25.2, and Automation Bench (Public) 25.1. These figures are useful discovery evidence, but DeepSeek controls the harness and report. They are not an independent head-to-head basis for deciding that it beats any GPT-5.6 configuration.
Artificial Analysis independently reports V4 Flash 0731 at 50, up from the April V4 Flash’s 40, including GDPval-AA v2 Elo 1559, Terminal-Bench 2.1 at 79%, τ³-Banking at 31%, SciCode at 50%, Humanity’s Last Exam at 37%, AA-LCR at 66%, and GPQA Diamond at 91%. That makes it a serious budget contender, not proof that it is operationally interchangeable with OpenAI’s family.
| Work type | First choice | Why | Escalate when |
|---|---|---|---|
| High-volume, bounded, text-heavy execution where startup delay is acceptable | Luna at high/max | Strongest GPT-5.6 quality-per-dollar evidence; Luna-max beats Terra-medium on AA’s aggregate score | The task requires materially better planning/judgment or repeats/failures erase savings |
| Interactive work needing immediate first response | Terra at medium, or Luna at lower effort after testing | Terra-medium’s measured TTFT is far lower than Luna-max’s | Quality problems or expensive repair appear |
| Ambiguous, high-stakes planning, hard debugging, architecture, or consequential synthesis | Sol, with effort matched to stakes | Highest family capability; lower downstream risk of a bad plan | The task becomes well specified and independently checkable |
| Cheap external challenger for text-only, controlled agent/coding trials | DeepSeek V4 Flash 0731 max | 50 Index points, roughly 60% less measured cost per task than Luna-max | Image input, reliability, compliance, tool/harness behavior, or real-work error rate matter more than benchmark cost |
No source found provides a complete, independently replicated matrix of Sol/Terra/Luna at low, medium, high, xhigh, and max on the same tasks with raw token logs and repair loops. That missing matrix is the actual answer to “what does it cost to get the same task done?”
The public evidence is strong enough to change the default hypothesis from “Terra is the economic middle” to “Terra needs a latency-specific justification.” It is not strong enough to deploy a universal routing policy without local measurement.
Build a small task set from real, non-sensitive work, stratified into:
- bounded implementation and tests,
- repository diagnosis,
- research/synthesis with source checks,
- interactive short-turn requests,
- a deliberately ambiguous planning task.
Run each task at least three times with fixed prompts, tools, maximum wall time, and acceptance criteria. Compare Sol-medium, Terra-medium, Luna-high, Luna-max, and V4 Flash 0731 max. Record:
| Metric | Why it matters |
|---|---|
| Task success / acceptance rate | The denominator for cost per completed task |
| Total billable input, cache-read, cache-write, reasoning, and output tokens | Actual spend, not list-price intuition |
| Tool calls and retries | Agent-loop costs that token comparisons hide |
| Human repair minutes | A cheap wrong result is not cheap |
| Time to first token and time to verified completion | Separates interactive latency from throughput |
| Failure mode | Shows where a routing rule is unsafe |
Calculate all-in model cost / accepted outcomes and separately report human repair time. Keep the models that form the actual Pareto frontier for the work, and remove any routing rule that is consistently dominated.
- High confidence: Current OpenAI prices and model positioning, because they are in OpenAI documentation; current DeepSeek 0731 public-beta status and vendor benchmark settings, because they are in DeepSeek’s changelog.
- High confidence within its declared scope: Artificial Analysis’s published score/cost comparisons and method. It is the strongest independent cross-model source found.
- Medium confidence for operational routing: Transfer from public benchmarks to a particular agent stack or task mix. Harness, prompts, tool configuration, retry policy, caches, and human review can change the cost per accepted task.
- Low confidence: Any claim that V4 Flash 0731 is already generally “better” than GPT-5.6. The evidence shows a close budget comparison with Luna-max and a clear advantage over Terra-medium in this one aggregate Index; it does not establish universal superiority.
- OpenAI, GPT-5.6 model family, accessed 2026-08-01.
- OpenAI, Advancing the price-performance frontier with GPT-5.6, accessed 2026-08-01.
- OpenAI, GPT-5.6 Sol model documentation, accessed 2026-08-01.
- OpenAI, GPT-5.6 Terra model documentation, accessed 2026-08-01.
- OpenAI, GPT-5.6 Luna model documentation, accessed 2026-08-01.
- OpenAI, Reasoning models guide, accessed 2026-08-01.
- Artificial Analysis, GPT-5.6 benchmarks across Intelligence, Speed and Cost, accessed 2026-08-01.
- Artificial Analysis, Luna-max versus Terra-medium, accessed 2026-08-01.
- Artificial Analysis, Intelligence benchmarking methodology, accessed 2026-08-01.
- DeepSeek, API changelog, 2026-07-31 V4 Flash update, accessed 2026-08-01.
- Artificial Analysis, DeepSeek V4 Flash 0731 scores 50, accessed 2026-08-01.
- Artificial Analysis, V4 Flash 0731-max versus Luna-max, accessed 2026-08-01.
- Artificial Analysis, V4 Flash 0731-max versus Terra-medium, accessed 2026-08-01.