Shop
VERTUVERTU

GPT-5.6 Sol, Terra or Luna? A Workload Router for Real Projects

[_AI_TOOLS_]

> date: PUBLISHED ON JUL 16, 2026> decoder: VERTU AI & INNOVATION DESK

An executive routing three different project cards to fast, balanced and deep-reasoning AI work lanes on a laptop

Why it matters

Choose GPT-5.6 Sol, Terra or Luna by reversibility, complexity, latency and review cost—not by using the largest model for every task.

OpenAI’s GPT-5.6 family introduces three durable capability tiers: Sol as the flagship, Terra as the balanced option and Luna as the fastest, lowest-cost tier. The tempting policy is to use Sol for everything important and Luna for everything cheap. That wastes time and money because task risk is not the same as task difficulty.

A short public summary can be difficult but easy to reverse. A simple-looking spreadsheet update can be high risk if it changes a board number. The better router considers four variables: consequence, ambiguity, latency and verification cost.

The workload router

Work type Default tier Reasoning effort Human control Escalate when
Classification, extraction, formatting Luna Low/standard Sample and reconcile Inputs are ambiguous or errors compound downstream
Routine research synthesis and first drafts Terra Medium Source review and edit Sources conflict or recommendation affects money/people
Complex analysis, architecture and multi-source judgement Sol High or max Define acceptance tests and approve consequences The task needs independent workstreams or unresolved evidence
Irreversible external action Any tier may prepare Appropriate to analysis Human approval before action Never delegate the approval boundary to model size
Repeated production workflow Luna/Terra after validation Lowest proven level Monitor drift and exceptions Error rate or input distribution changes

OpenAI’s GPT-5.6 launch describes Sol, Terra and Luna, model availability, effort settings and an ultra mode that coordinates multiple agents for demanding work. The earlier Sol preview provides additional capability and safeguard context. These are first-party claims and evaluations. Your router should be calibrated on your own tasks and acceptance criteria.

Luna: high-volume work with visible rules

Use Luna where the task can be specified tightly and checked cheaply:

  • label support tickets against an existing taxonomy;

  • extract dates, entities and amounts from standard documents;

  • reformat approved copy into channel variants;

  • produce a first-pass meeting digest;

  • check a page against a mechanical checklist;

  • route straightforward requests to a known owner.

The design principle is boundedness. Give Luna a schema, examples and a reject state. Do not force a guess when a field is missing. For a thousand-row job, test a stratified sample before processing the full set and reconcile totals afterwards.

Fast generation is not useful if humans must inspect every character. Prefer tasks where a database constraint, parser, comparison or spot-check can catch errors.

Terra: the default for ordinary knowledge work

Terra is the sensible starting point for work that needs synthesis but not the longest reasoning path:

  • compare several documented options;

  • turn interviews into themes and evidence;

  • draft a brief from supplied research;

  • prepare a meeting pack;

  • explain a technical issue to a non-specialist;

  • revise content to an established voice.

Give it source boundaries and an output contract. Ask it to separate facts, assumptions and recommendations. A balanced model can still hallucinate a citation or flatten a disagreement, so the reviewer should inspect the claims that carry the decision.

If the output repeatedly needs major reconstruction, do not merely increase prompt length. Escalate the task or break it into evidence collection, analysis and final communication.

Sol: use deeper reasoning where it changes the result

Sol earns its place when the work contains interacting constraints, long context or expensive mistakes:

  • architecture or security review across multiple systems;

  • financial or operational scenario analysis;

  • investigation where several plausible causes must be tested;

  • negotiation preparation with conflicting stakeholder goals;

  • a decision memo based on incomplete evidence;

  • creation and inspection of a complex deliverable.

Higher reasoning effort should buy a better decision, not a longer answer. Define success before the run: reconciled numbers, named assumptions, passed tests, evidence links and a clear stop condition.

Use max when one difficult reasoning path needs more time. Use multi-agent or ultra-style coordination when independent workstreams genuinely exist—research, modelling, counterargument and verification, for example. More agents can duplicate errors or expand scope if roles are vague.

The reversibility rule

Model selection does not grant authority. A small model can safely prepare a large email if a person approves it. A flagship model should not autonomously send a payment, delete records or publish a claim merely because it reasoned longer.

Classify actions:

Action class Examples Required control
Reversible draft Outline, analysis copy, local file Normal review
Controlled mutation CRM update, CMS draft, calendar change Preview, scoped identifiers and audit log
External commitment Send, publish, purchase, approve Explicit authorised approval
Destructive or regulated Delete, access control, legal/medical decision Specialist policy and independent verification

The model can recommend escalation. It should not redefine the boundary.

Build a measured router in one week

Select 20 representative tasks and record:

  1. Input size and ambiguity.

  2. Tier and effort used.

  3. Completion time.

  4. Human review minutes.

  5. Number and severity of corrections.

  6. Whether the output passed its acceptance test.

  7. Whether a larger tier materially improved the outcome.

Calculate total work cost, not token cost alone. A cheaper model that adds 20 minutes of executive review is expensive. A flagship model that produces a polished but unverified answer is also expensive.

Promote a task to a lower tier only after repeated passes. Re-test when templates, tools or source types change.

Continuity and archive discipline

Do not make the router dependent on one model being available. Keep prompts, source files and acceptance tests portable. If an incident interrupts the preferred tier, our 30-minute AI continuity plan preserves decisions and deadlines. Organise reusable context so retrieval does not become the bottleneck; the ChatGPT archive search guide provides a practical reset.

Verdict

Use Luna for bounded, high-volume transformations with cheap checks. Use Terra as the default for ordinary synthesis and drafting. Use Sol when ambiguity, interacting constraints or the cost of a weak answer justifies deeper reasoning. Keep consequential actions behind the same approval boundary regardless of tier.

The best model policy is not “always use the smartest”. It is “use the lowest tier that repeatedly passes the real acceptance test, and escalate before review cost or risk becomes hidden”.

More In AI Tools