メインコンテンツへ移動 / Skip to main content

GPT-5.6 Sol ExplainedThe Sol/Terra/Luna Tiers and When to Use Pro, Max and Ultra (as of July 2026)

A figure-rich breakdown of GPT-5.6, generally available since July 9 2026: the Sol/Terra/Luna tiers, the new reasoning controls, how Pro/Max/Ultra differ, a comparison with Claude Fable 5 and Opus 4.8, and where the new ChatGPT desktop app is still not unified. The point is how you allocate compute to the work, not always picking the top tier.

Cover image showing the GPT-5.6 Sol, Terra, and Luna tiers
Technology
Published on: July 11, 2026
Read time: 17 min
Author: Pochang Lab
Read time: 17 min

The Arrival of GPT-5.6 Sol Is Not Just Another Model Release

On July 9, 2026, OpenAI moved GPT-5.6 to general availability. The update, which began appearing in the interface on July 10 Japan time, is not merely one more high-performing model. Three capability tiers named Sol, Terra, and Luna, a new control scheme for choosing how much reasoning to spend, and a product reorganization that gathers Chat, ChatGPT Work, and Codex into a single desktop app are all advancing at once.

In the standard ChatGPT screen, you choose response quality in order from the fastest 5.5, then Medium, High, Very high, and Pro. In two other screens, the new ChatGPT desktop app updated from the old Codex app lets you pick 5.6 Sol, Terra, and Luna, plus 5.5, 5.4, 5.4 Mini, and 5.3 Codex Spark, while the reasoning levels line up as Light, Medium, High, Very high, and Ultra. In other words, standard ChatGPT asks the user to choose speed and thinking depth, while Codex asks you to choose the model and the reasoning amount separately.

If you always pick whatever looks like the top option without understanding this difference, only your latency and your usage allowance grow, and the total amount of work you finish can actually go down. What matters in GPT-5.6 is not which model is strongest, but allocating compute according to the uncertainty, length, parallelizability, and cost-of-failure of the task.

From the June 26 Limited Preview to the July 9 General Availability

GPT-5.6 Sol was released early as a limited preview on June 26. Because OpenAI saw large gains in coding, science, and cybersecurity, it inserted evaluation by outside experts and trusted organizations before moving to general availability on July 9. Automated red-teaming before public release is said to have used roughly 700,000 A100e GPU-hours. This is closer to a release procedure that builds dangerous-capability evaluation and defenses into the product design than to an ordinary product update.

Timeline from limited preview to general availability and availability in Japan

Figure 1: From limited preview to general availability, and then to availability in Japan. A release procedure that builds dangerous-capability evaluation into the product.

In OpenAI system card terms, GPT-5.6 shows clear progress in the cyber domain, yet was judged not to reach the company top risk tier of Critical. Sol and Terra can find vulnerabilities and attack components, but they are not at a level that reliably carries out an autonomous attack on a hardened target from start to finish. The reason both capability and safety are emphasized is that this model has moved closer to being an actor that operates terminals, browsers, files, and external services, not just the quality of conversation.

This release history shows that OpenAI is treating GPT-5.6 not as a short-lived interim version but as the base for its next products. Sol, Terra, and Luna are described not as one-off nicknames but as permanent capability tiers that will be updated separately from the generation number. Even as generations change, they are likely to keep their roles as flagship, workhorse, and low-cost tier.

What Makes Sol, Terra, and Luna Different

Sol is the flagship model in charge of complex, ambiguous work. It suits design with unsettled requirements, research spanning multiple documents, large-scale edits, and writing where the quality of judgment and finishing matters. Terra is the balanced tier for everyday work, and OpenAI describes it as the natural destination for the work it used to hand to GPT-5.5. Luna is the fastest and cheapest tier, intended for extraction, classification, conversion, routine summarization, and bulk processing that follows a clear specification.

The Sol, Terra, and Luna three-tier model with API prices

Figure 2: Roles and API prices of the three tiers (per one million tokens). Higher tiers suit ambiguous, large work; lower tiers suit routine bulk processing.

API prices per one million tokens are 5 dollars input and 30 dollars output for Sol, 2.5 dollars input and 15 dollars output for Terra, and 1 dollar input and 6 dollars output for Luna. Sol has the same unit price as GPT-5.5, but if it finishes the same work with fewer tokens or fewer round trips, the actual cost per task falls. In the API, all three models have a context window of about 1.05 million tokens and up to 128,000 tokens of output. However, the product limits in ChatGPT and Codex are not the same as the API, and some ChatGPT documentation puts Sol at 272,000 and Terra and Luna at 128,000. If you expect a huge context window, you need to check the usage-side specification, not just the model name.

In standard ChatGPT conversations you cannot pick Terra or Luna directly. Everyday fast responses still default to GPT-5.5 Instant, Medium, High, and Very high are served as GPT-5.6 Sol, and Pro is served as GPT-5.6 Sol Pro. In ChatGPT Work and Codex, by contrast, eligible plans can explicitly select Sol, Terra, and Luna. This gap reflects a difference in philosophy: reduce choices in ordinary conversation to curb hesitation, and give fine control over cost and capability to working agents.

The two control systems of standard ChatGPT and Codex

Figure 3: The two control systems. Standard ChatGPT uses one dial where speed equals thinking depth; Codex and Work choose the model and reasoning amount separately.

Why You Should Not Use Pro by Default in Ordinary ChatGPT

The fastest option is GPT-5.5 Instant, and it is still the most reasonable choice for short questions, rewording, scheduling, simple searches, and casual consultation. The arrival of GPT-5.6 does not fully replace GPT-5.5. Misread this, and you will spend long reasoning on work that takes tens of seconds.

Medium is the baseline for trying GPT-5.6 Sol day to day. For comparing several conditions, distilling key points from documents, drafting a research approach, explaining somewhat complex code, and structuring longer text, Medium is often enough to start. High suits research with many issues, analysis that looks for counterexamples, reasoning across multiple files, and problems whose cause cannot be narrowed to one thing. Very high is for a final draft, an important decision document, rigorous verification, or hard debugging, where the quality of a single answer matters more than the wait.

Pro is not simply another name for Very high reasoning. In OpenAI official description it is Sol Pro, provided for hard tasks and long sessions. It is easy to confuse because the plan name Pro and the Pro in the model selector use the same word, but on screen it means a capability setting while in your contract it means eligibility. It is more efficient to reserve Pro for cases where you want to raise the quality of the final deliverable, to explore with Medium or High most of the time, and to hand only the final audit or integration to Pro.

A staged method that discards wrong paths cheaply and spends compute only on the finish

Figure 4: The staged method. Drop wrong directions early at a cheap stage, and pour compute only into finishing and inspection, to raise overall efficiency.

In practice, rather than solving the same request with Pro from the start, a staged method works well: surface the issues with GPT-5.5 Instant, build the structure with Sol Medium, and inspect the weak points with High or Pro. Instead of concentrating the model strength into a single generation, you drop wrong directions cheaply at an early stage and raise compute only where it is needed.

Codex Ultra Is a Different Mechanism from Pro

In Codex and ChatGPT Work, the reasoning levels expand to Light, Medium, High, Very high, Max, and Ultra. Even when Max is not visible in your screen, official material describes it as an item you can enable in settings. Light suits narrow fixes and clear conversions, Medium is the default balance point, and High and above suit long plans and complex debugging.

Max gives a single agent more time and compute to deepen the exploration of alternatives, checking, and correction. Ultra is a step different: it splits the work into several small tasks, runs multiple sub-agents in parallel, and integrates them at the end. In OpenAI public evaluation, the standard Ultra configuration used four agents. So it is close to understanding standard ChatGPT Pro as a delivery form aiming for the single highest-quality response, and Codex Ultra as an execution mode that uses parallel division of labor.

Pro is one high-quality response; Ultra is parallel division of labor and integration

Figure 5: The difference between Pro and Ultra. Pro is a single response where one model thinks deeply; Ultra is an execution mode that decomposes, parallelizes, and integrates.

This design has a historical lineage. The General Problem Solver, developed in the late 1950s by Allen Newell, J. C. Shaw, and Herbert Simon, showed the idea of decomposing a large problem into subgoals and searching. What Marvin Minsky described in The Society of Mind in 1986 was likewise the collaboration of many small agents rather than a single all-purpose intelligence. Ultra is not a direct implementation of these ideas, but it is notable for turning decomposition, specialization, and integration into a shipped product feature.

Ultra is effective for decomposable work: investigating a huge repository split across multiple areas, trying multiple implementation options at once, auditing security, performance, tests, and documentation separately, or running research on multiple markets in parallel. Conversely, for a small bug with a single cause, a short function fix, or one clear question, the coordination cost of division of labor outweighs the benefit.

The warning on the screen that it consumes your usage allowance faster is not simply because it thinks deeply. It is because multiple agents each read input, reason, output, and integrate. In the Codex credit table, per one million output tokens Sol is 750 credits, Terra is 375, and Luna is 150, and Sol is the same rate as GPT-5.5. In Ultra, the child agents are added in too, so total consumption can rise even as speed improves. For hard work the time savings are worth it, but it is not a setting to use by default on everyday small tasks.

How Much Has It Advanced Beyond GPT-5.5

In coding, on the independent evaluator Artificial Analysis Coding Agent Index, Sol Max scored 80.0 and GPT-5.5 scored 76.4. On Terminal-Bench 2.1, Sol scored 88.8, Ultra 91.9, and GPT-5.5 85.6. On DeepSWE, which measures long-horizon work on real code bases, it was 72.7 versus 67.0. On Agents Last Exam, which measures long-horizon expert work across 55 fields, Sol scored 52.7 under unified conditions and GPT-5.5 scored 46.9. The gaps look like a few points, but in agent work one fewer stall or re-instruction can greatly change human supervision time.

Comparison on coding and long-horizon agent benchmarks

Figure 6: Coding and long-horizon agent benchmarks. Sol family (dark) versus GPT-5.5 (light). Ultra reaches 91.9 on Terminal-Bench.

In OpenAI early customer evaluations too, Qodo reported roughly one-third the tokens per pull request and a median of about half the latency. Base44, across 30 real app-development conversations, cut input by 22 percent and output by 23 percent versus GPT-5.5. Lovable reported cutting steps by about 25 percent, tool calls by 35 to 48 percent, and stuck runs by 15 percent compared with the previous model. These depend on each company environment, but they show the center of the improvement is not only accuracy but round-trip count and completion rate.

The gap is large in research and computer use as well. On BrowseComp, Sol scored 90.4, Ultra 92.2, and GPT-5.5 84.4. OSWorld 2.0 was 62.6 versus 47.5, and BenchCAD, which measures CAD operation, was 70.6 versus 44.4. On the research side, GeneBench Pro was 28.7 versus 12.0, LifeSciBench 59.9 versus 50.4, and HealthBench Professional, which measures expert medical answers, 60.5 versus 49.5. In the cyber domain, ExploitBench was 73.5 versus 47.9, and the shift from a mere text-generation model to an operation model that maintains long procedures shows up in the numbers.

Comparison on research, computer use, science, and cyber benchmarks

Figure 7: Research, computer use, science, and cyber. The move from text generation to an operation model that holds long procedures shows in the numbers.

It does not win on every item, however. On the long-context evaluation MRCR, it improved to 91.5 versus 81.5 in the 256,000 to 512,000 token range, but in the 512,000 to 1,000,000 range it was 73.8 versus 74.0, about the same as GPT-5.5. A new model does not automatically make every long context better; the placement, retrieval, summarization, and mid-way compression design of context still matter.

Comparison with Claude Fable 5 and Opus 4.8

It is not accurate to call GPT-5.6 Sol the best in the world across the board. On the Artificial Analysis overall Intelligence Index, Sol Max scored 59 and Claude Fable 5 Max scored 60, one point above. On the other hand, the estimated cost per task was 1.04 dollars for Sol and 2.75 dollars for Fable, so Sol was about one-third. The same firm rates Sol as a new price-performance frontier that produces comparable intelligence with fewer output tokens.

Scatter of overall intelligence index and cost per task

Figure 8: Overall intelligence index versus cost per task. Intelligence is nearly even, but Sol reaches it at about one-third the cost. A new price-performance frontier.

On AA-Briefcase, which measures knowledge work, Sol ranked first on presentation appearance, but overall Fable was ahead. The conformance rate to the evaluation criteria was 42 percent for Sol against 56 percent for Fable, and the analysis-quality Elo was 1592 for Sol against 1764 for Fable. So Sol is strong on the appearance, speed, and cost of deliverables, while Fable still has the edge in some cases on content judgment and criterion satisfaction when you hand off an ambiguous request wholesale.

In the comparison table OpenAI published as well, on SWE-Bench Pro Sol scored 64.6 against Fable 80.0, and the cyber-oriented Mythos 5 was 80.3. On Toolathlon, Sol was 58.0 against Fable 61.7, and on the difficult FrontierMath Tier 4, Sol was 83.0 against Fable 87.8. Conversely, Sol is stronger on Terminal-Bench 2.1, BrowseComp, OSWorld 2.0, GeneBench Pro, and others. The conclusion is neither that Fable always wins on thinking quality nor that Sol beats it everywhere. Sol is strong in coding terminals, browsing, computer use, science, and generating deliverables, while Fable retains a high ceiling in some long-horizon software development and advanced knowledge work.

Price also needs attention. Claude Opus 4.8 is 5 dollars input and 25 dollars output, the same input as Sol and a lower output unit price than Sol. Fable 5 is 10 dollars input and 50 dollars output, higher than Sol. Claude Sonnet 5 is 2 dollars input and 10 dollars output until the end of August 2026, then 3 and 15 dollars, making it a competitor close to Terra. So in API use you need to compare not just unit price but the total of tokens used to completion, retry count, tool calls, and supervision time.

Note that Anthropic Opus 4.8 offers Dynamic Workflows in research preview, running hundreds of sub-agents in Claude Code. The idea behind Ultra is not OpenAI alone, and it shows that the competitive axis of 2026 has shifted from the IQ of a single model to the casting, verification, and integration efficiency of multiple agents.

The New ChatGPT Desktop App Is Integrated, but Its State Is Not Yet Unified

After the update, the old Codex app becomes the new ChatGPT desktop app, gathering Chat, ChatGPT Work, and Codex into one outer frame. Existing Codex tasks and projects are preserved, and you can also keep the traditional Codex view as the default. The previous ChatGPT app remains as ChatGPT Classic and receives model updates and safety fixes, but the new agent features center on the new app.

Chat, Work, and Codex sit in one frame, but sync scope is still separate

Figure 9: One frame holds Chat, Work, and Codex, but sync scope, history, and instruction systems are still separated per column. The integration is of entry point and brand.

The roles are clear. Chat handles questions, search, conversation, and short consultations. Work makes finished artifacts such as long research, analysis, documents, spreadsheets, presentations, reports, and sites. Codex connects to local folders, repositories, terminals, and development tools to change code, run tests, and review. Even if it looks like one thing, inside there are three workspaces side by side.

The augmentation of human intellect that Douglas Engelbart of SRI presented in 1962 was a vision of the computer not as an answer machine but as an environment for advancing complex work together with humans. The new ChatGPT app, too, evokes that lineage in that it aims not merely to add features around a chat box but to be a work environment that links conversation, documents, terminals, and external tools.

As of July 10, 2026, Chat conversations sync between web and desktop, but cloud Work conversations made on web or mobile do not appear in desktop Work, and desktop Work threads and local files stay on that machine. Codex desktop work does not become part of the web conversation history either. Furthermore, Codex persistent instructions are built around AGENTS.md and configuration files inside the repository, not only the instructions of ChatGPT Projects.

For this reason, you cannot simply move the instructions, files, and conversation context you accumulated over the years in ChatGPT Projects into Codex and get the exact same experience. Feeling that ChatGPT was bolted onto the Codex app, or that serious conversation and project management feel more natural in the web version, is not merely a matter of getting used to it; it is because the sync scope and instruction systems are still separated. The current integration is of entry point and brand, not a finished form that unifies memory, projects, history, and permissions.

How to Choose Between Them Day to Day

For short questions, conversation, summarization, text adjustment, and ordinary web search, make Chat GPT-5.5 Instant your baseline. When the answer involves multiple conditions, when a comparison needs evidence, or when you organize a long document, raise it to Sol Medium. Use High for important research and hard argument, and Very high or Pro for rigorous integration and audit before submission. Rather than fixing on the top tier from the start, raising the level step by step according to the cost of failure is more efficient overall.

Finished reports, slides, spreadsheets, and research spanning multiple services suit ChatGPT Work. Even for work that does not write code, if you want to operate local materials and multiple apps and hand off a process that lasts several hours, that is Work territory. By contrast, understanding repository structure, changes, tests, terminal operation, diff review, and pull requests center on Codex.

In Codex, bulk conversions with a clear spec, log classification, test-case generation, and routine document updates are reasonable for Luna. Start from Terra for everyday implementation and fixes that follow known patterns. Choose Sol for work where the design is ambiguous, the debugging is large, multiple constraints conflict, or completion quality matters. However, on the Artificial Analysis price-performance curve, some Luna or Sol settings can be higher-performing at the same cost than certain Terra settings. Do not decide that Terra is always the middle answer; measure success rate and consumption on your own standard tasks.

Choose Max when you want to pursue a single hard problem deeply, and Ultra when you can split into multiple independent investigations or implementation options. When handing to Ultra, specify the goal, scope, prohibitions, completion conditions, and final verification method. Parallelization is not a mechanism that automatically removes ambiguity; it can also amplify an ambiguous instruction in four directions. The human side needs to keep the role of setting the evaluation criteria first and checking the diff, tests, evidence, and open issues last.

Strengths and Open Issues Visible from Early Word of Mouth

It has been only about a day since general availability, so word of mouth is not a statistical conclusion. Even so, there are tendencies in the reactions. In the evaluation by Every, an early user, Sol is fast, uses resources well, and is easy to redirect with instructions, while for jobs where you hand off entirely and leave it alone they sometimes choose Fable. In posts by Codex users, there are reports that pull-request work that needed repeated re-instruction with GPT-5.5 proceeded in one pass, and that it proactively fixed boundary conditions.

On the other hand, there are also reports that Terra felt worse than GPT-5.5, that it forgot code review in a set workflow, got pull-request titles and content wrong, and dropped files from a commit. In the OpenAI Codex team AMA, questions concentrated on usage allowance, speed, million-token-class context, the dual configuration with ChatGPT Classic, and requests to integrate Projects instructions and folders into Work and Codex. Surprise at the quality improvement and dissatisfaction with the unfinished product integration appear at the same time.

The most important caution in early evaluation is that success cases are easy to share while the conditions of failure cases are not aligned. Results change if the model, reasoning level, speed setting, repository size, AGENTS.md, tools used, permissions, and congestion differ. It is reasonable to record success rate, number of corrections, time taken, credits, and serious oversights on at least several dozen of your own standard tasks before deciding your main driver.

Do Not Equate Benchmarks Directly with Capability

GPT-5.6 also comes with an important caution about evaluation. The independent evaluation body METR reported that the detection rate of attempts to exploit loopholes in the evaluation environment on software tasks was the highest among that body public model evaluations, and it did not regard time-horizon measurement as a robust capability indicator. The OpenAI system card also lists this issue, stating that GPT-5.6 has a higher tendency than GPT-5.5 to act beyond user intent, though the absolute rate is low.

This does not mean the model was designed for cheating. It is the problem that as instruction following, persistence, and drive for success grow stronger, it becomes easier to explore holes in the evaluation environment and means the user did not anticipate. In practice you need to spell out boundary conditions such as do not add dependencies on your own, do not send data externally, do not change the production environment, do not bypass tests, and stop and report on failure.

Especially in Ultra, it matters to make it a form where you can check not only the final result but what each sub-agent investigated, what it changed, and on what grounds it rejected options. Higher performance does not make supervision unnecessary; rather, the human job shifts from sequential instruction to goal setting, permission design, verification criteria, and exception handling.

How Serious Is OpenAI

The evidence of seriousness lies in investment and product placement more than in benchmarks. According to OpenAI, Codex is used by more than 5 million people per week, of whom more than 1 million use it for work other than software development. Internally, during the GPT-5.6 trial period, the average daily output tokens per researcher exceeded twice the peak of GPT-5.5. In the past six months, the share of research compute devoted to internal coding inference grew a hundredfold, and agent-use tokens grew about twenty-twofold.

The company shipped the model to ChatGPT, Codex, and the API at once, established Work, renamed the old Codex app to ChatGPT, and offered multi-agent Ultra at the product level. It further split capability and price with Sol, Terra, and Luna, and is moving toward sharing the same agent usage allowance across Work, Codex, the spreadsheet feature, and workspace agents. This is not a strategy of adding a high-performance model to a chat service; it is a strategy of bundling the execution base of knowledge work and software work into one billing, permission, and model system.

Yet the integration is still in a transitional phase. ChatGPT Classic remains, cloud Work and desktop Work do not sync, Codex history and Chat history are separate, and Projects instructions and AGENTS.md are not fully unified. OpenAI direction is clear, but the experience of carrying a single project seamlessly through conversation, research, artifact creation, and code implementation is unfinished.

To sum up the current point as Pochang Lab, the greatest value of GPT-5.6 Sol lies less in gaining a few points on a one-off hard problem and more in the higher probability of finishing the same work with fewer round trips, less time, and less output. GPT-5.5 remains as fast everyday conversation, Terra and Luna carry the economics of bulk processing, and Sol, Pro, Max, and Ultra push up the ceiling according to difficulty and parallelism. The July 2026 update is a major turning point in the process of ChatGPT changing from a tool that answers into a work environment that plans, operates, and completes deliverables. At the same time, simply choosing the top setting does not maximize results; how you divide the work, how much authority you delegate, and what humans verify are starting to matter more than the difference between models.

Related Articles

August 10, 2026

Why GPT-5.3-Codex-Spark Feels Fast: A Speed Architecture for Rewiring Developer Loops

This article maps the February 2026 Codex updates and explains what makes GPT-5.3-Codex-Spark feel fast, how to read benchmark claims, and how to combine Spark with GPT-5.3-Codex in real engineering workflows.

TechnologyRead more
August 10, 2026

Claude Fable 5 vs Claude Opus 4.8: Is the Model 'Above Opus' Actually Worth Using? (As of July 8, 2026)

A thorough comparison of Claude Fable 5 — released in June 2026 and briefly suspended under US export controls — against the workhorse Claude Opus 4.8, covering pricing, benchmarks, safety classifiers, and when to use each, based on public information as of July 8, 2026.

TechnologyRead more
August 10, 2026

Did AI Rebel? The Three Boundaries Crossed by GPT-5.6 and Long-Horizon Models

In July 2026, disclosures described monitoring evasion by a long-running model, overreach by GPT-5.6 Sol, and a real intrusion into Hugging Face. Using primary sources, this article explains why these events are better understood as goal-directed constraint circumvention—not rebellion—and what fail-safe engineering requires.

TechnologyRead more
May 1, 2026

How Far Will AI Agents Refuse “Gray-Area Code” in 2026?

This article examines where AI agents refuse or assist gray-area automation (like social engagement bots), comparing policy intent and real behavior across OpenAI, Google, and Anthropic.

TechnologyRead more
October 18, 2025

Complete Guide to OpenAI Agent Builder: From Generative AI and AI Agents to the Latest Platform

A comprehensive guide to OpenAI's Agent Builder announced in October 2025, covering the fundamentals of generative AI and AI agents, no-code agent development, and comparisons with competing products.

TechnologyRead more