メインコンテンツへ移動 / Skip to main content

Inside OpenAI’s “Code Red”The Giant’s Next Moves as Gemini and Claude Close In

Breaks down why OpenAI declared “Code Red” by examining benchmark shifts, enterprise share, massive infra bets, and safety risks—and sketches the company’s likely next moves from an engineer’s point of view.

Technology
Published on: December 11, 2025
Updated on: August 16, 2026
Read time: 16 min
Author: Pochang Lab
Read time: 16 min

The Truth About OpenAI's "Code Red" — The Giant's Next Move, with Gemini and Claude Closing In

Introduction: Why Would the Champion Declare an Emergency?

Since ChatGPT arrived at the end of 2022, most people have felt that OpenAI has simply been the protagonist of generative AI. And by mid-2025 ChatGPT had become one of the largest consumer AI services in history, with 700 million weekly active users. (OpenAI)

Then, in December 2025, multiple outlets reported that an internal emergency declaration named "Code Red" had been issued inside OpenAI. Behind it lies the reality that rival large language models — Google's Gemini 3 and Anthropic's Claude 4.5 — have been closing hard on OpenAI in both capability and enterprise adoption, and in places have started to pass it. (Tom's Hardware)

This article works through what Code Red means from three angles — the technical picture in benchmarks, the business picture in revenue and market share, and infrastructure investment and its risks — using concrete numbers wherever possible. From there we can speculate, from an engineer's perspective, about OpenAI's next move.


1. What Happened — Inside the Code Red Memo

First, the facts.

1-1. A hard pivot to "ChatGPT above all"

The Financial Times reported that in early December 2025 CEO Sam Altman issued an internal memo declaring a shift into a "Code Red" phase. In broad strokes it came down to three points. (Financial Times)

  • Improving ChatGPT takes absolute priority; other new projects go to the back of the queue
  • Experimental projects such as the video social app and shopping agents are cut back or halted
  • With model differentiation getting harder, protecting the user experience and continued use becomes the top concern

Business Insider reported the same memo, summarising the policy with the memorable line "protect the loop, delay the loot" — that is, do not let advertising break the loop of user behaviour data that ChatGPT feeds on. The dialogue logs from something approaching a billion weekly users are OpenAI's largest moat. (Business Insider)

1-2. The immediate trigger: Gemini 3 overtaking on benchmarks

According to Tom's Hardware and several other technology outlets, the direct trigger for the declaration was outside evaluation showing that Google's Gemini 3 had surpassed ChatGPT-family models on major industry benchmarks. (Tom's Hardware)

Fortune further reported that in the same memo Altman said OpenAI would

release a new reasoning-focused model next week that beats Gemini 3 on internal evaluations

signalling an intent to claw back the performance lead in the short term. (Fortune)

So Code Red can be summarised as this declaration:

ChatGPT is still the world's largest AI service, but the lead is becoming precarious in both capability and revenue. Therefore, concentrate every resource on the keep — ChatGPT itself.

2. Where OpenAI Stands, in Data: Overwhelming, and Surprisingly Fragile

Now let us check both the dominance and the wobble against the numbers.

2-1. ChatGPT is a platform at 700 million weekly users

According to "How people are using ChatGPT," a research report OpenAI itself published in September 2025:

  • ChatGPT has roughly 700 million weekly active users
  • The analysis covered 1.5 million conversation logs, making it the largest study of real-world consumer AI usage
  • About 30% of use is work-related, 70% private
  • About half of conversations (49%) are "asking," 40% are "doing," and the remaining 11% are "expressing" — creative work and reflection

(OpenAI)

On those numbers alone, the scale of the user base is not in question.

2-2. A $10 billion annual run rate — and still lossmaking

Viewed as a business, though, the numbers show a company in a distinctly aggressive position.

In June 2025, Reuters reported that:

  • OpenAI's annualised revenue run rate had reached $10 billion (about ¥1.5 trillion)
  • roughly double the $5.5 billion figure of December 2024, in half a year
  • with a full-year 2025 revenue target of $12.7 billion
  • and yet a loss of around $5 billion in the prior year

(Reuters)

A separate analysis notes that while OpenAI could reach a run rate above $20 billion by the end of 2025, it is carrying data-centre investment commitments of roughly $1.4 trillion (about ¥220 trillion) over the next eight years. (blog.carnegieinvest.com)

In other words: revenue is growing fast, but the infrastructure commitments are on a completely different order of magnitude.

2-3. Enterprise use: 40–60 minutes saved a day

OpenAI's "The State of Enterprise AI" report, published in December 2025, surveyed about 9,000 business users across roughly 100 companies and reported: (OpenAI CDN)

  • 75% of employees said AI improved the speed or quality of their work, or both
  • ChatGPT Enterprise users reported saving 40–60 minutes per working day
  • 60–80 minutes for data science, engineering, and communications roles
  • 87% of IT departments said troubleshooting got faster
  • 73% of engineers said they shipped code faster
  • 75% of employees said AI let them complete tasks they previously could not

The report carries a marketing flavour, and whether it holds up as rigorous academic work is debatable. Even so, the numbers show that AI use inside companies is starting to sit at the centre of real work.

2-4. The whole market: enterprise generative AI spend tripled in a year

Menlo Ventures' 2025 report estimates that annual corporate spending on generative AI rose from $11.5 billion in 2024 to $37 billion in 2025 — close to a tripling. (Menlo Ventures)

By that report's figures:

  • $1.7 billion in 2023
  • $11.5 billion in 2024
  • $37 billion in 2025

which is an extraordinary growth rate for a software category.


3. What the Numbers Say About the Leader's Wobble — Share and Benchmarks

Here is the heart of it. OpenAI remains top-tier on revenue and users, but the picture shifts when you ask which LLM people are paying for.

3-1. Anthropic has overtaken OpenAI in the enterprise LLM market

According to Menlo Ventures' "2025 Mid-Year LLM Market Update," published in July 2025, enterprise spending on LLM APIs doubled from $3.5 billion at the end of 2024 to $8.4 billion by mid-2025. (GlobeNewswire)

Market share moved like this: (GlobeNewswire)

End of 2023

  • OpenAI: 50% (a commanding first)

Mid-2025

  • Anthropic: 32% (first)
  • OpenAI: 25% (second — half its share in two years)
  • Google (Gemini): 20% (third)
  • Meta Llama: 9%
  • DeepSeek: 1%

Anthropic is stronger still in code generation. Summaries of the related reporting put:

  • Claude at 42% share of the code-generation market
  • OpenAI at 21%

a two-fold gap. (AI Matters)

Code generation is often called the first unambiguous killer app, and the same report notes that Claude and the various AI IDEs — Cursor, Windsurf and others — have grown what GitHub Copilot once monopolised into a $1.9 billion ecosystem. (AI Matters)

3-2. Who sits at the technical top, per Chatbot Arena

Look at the human-evaluation rankings on LMArena (formerly Chatbot Arena), widely referenced in the technical community. As of December 2025, the Arena Overview had: (LMArena)

  1. gemini-3-pro (Google)
  2. grok-4.1-thinking (xAI)
  3. Claude Opus 4.5, thinking variant (Anthropic)
  4. gpt-5.1-high (OpenAI)
  5. chatgpt-4o-latest-20250326 (OpenAI)
  6. … …

Benchmarks are not everything, of course, and the value of an LLM cannot be captured in a single score. But it is telling that Google, xAI, and Anthropic lock up the top three in a community-driven overall evaluation, while OpenAI's newest models sit at sixth and thirteenth.

3-3. From "losing on capability" to "losing ground on revenue"

Reuters reported that Anthropic had reached an annualised run rate of $3 billion as of June 2025, with demand surging in particular from code-generation startups. (Reuters)

In the enterprise market, the data also shows:

  • workloads concentrating on whichever model performs best
  • vendor switching itself running low at 11%, while 66% of teams upgrade to their existing vendor's newest model

(GlobeNewswire)

Put simply, the market has moved fully into a phase where

falling from the top on capability feeds straight through to market share and revenue

Which means differences in model performance are no longer a hobbyist comparison — they connect directly to revenue and valuation.


4. The Infrastructure War: 250GW of Data Centres and a $1.4 Trillion Bet

OpenAI's Code Red is not only about technology and market share. It is bound up with the sheer weight of infrastructure investment.

4-1. A 250GW AI data-centre plan

Tom's Hardware reports that Sam Altman has set out a plan to build AI data centres with up to 250 gigawatts of compute capacity by 2033. (Tom's Hardware)

The article puts that scale in these terms:

  • 30 million GPUs required per year
  • power consumption comparable to the entire electricity consumption of India
  • carbon dioxide emissions roughly twice those of ExxonMobil

This is close to a ceiling figure and plainly ambitious, but it is a striking marker of how far the AI infrastructure war has scaled.

4-2. $1.4 trillion in capex commitments

Per the investment-consulting analysis cited above, OpenAI is:

  • aiming for an annualised run rate above $20 billion by the end of 2025 (blog.carnegieinvest.com)
  • while committing roughly $1.4 trillion (on the order of ¥220 trillion) to data-centre investment over the next eight years (blog.carnegieinvest.com)

As a rough exercise, that blog uses Google's early growth rates as a reference and notes that:

  • if OpenAI can generate $1.5 trillion in cumulative operating cash flow over the next five years,
  • it could just about fund the $1.4 trillion,
  • but that requires average annual revenue growth near 190%, reaching roughly $2 trillion in revenue by 2030 — a premise detached from reality

(blog.carnegieinvest.com)

So OpenAI has been backed into a position where it is

making enormous up-front investments, and cannot afford to lose to rivals on capability

4-3. "Sovereign AI" infrastructure partnerships with national governments

Against that backdrop, OpenAI has been accelerating partnerships with governments and infrastructure companies.

In December 2025, for instance, OpenAI announced its "OpenAI for Australia" programme and signed a memorandum with the local data-centre operator NEXTDC to build a $4.6 billion "sovereign AI campus." (OpenAI)

The programme sets out to raise AI skills among 1 to 1.5 million workers, making it a clear example of a country-level package that bundles:

  • infrastructure investment
  • workforce development
  • compliance with each country's regulation

5. Safety and Risk: Next-Generation Models Rated "High" for Cyber Risk

At almost the same time as Code Red, OpenAI issued a warning about the cybersecurity risk of its next-generation models.

In a statement dated 10 December 2025, OpenAI cautioned that:

  • coming frontier models may have capabilities sufficient to support developing zero-day vulnerabilities against highly defended systems
  • and complex intrusion operations targeting industrial infrastructure
  • and that they should therefore be classified in the "high" cyber risk category

(Investing.com)

In response, the company has set out a defence-in-depth approach: (Investing.com)

  • access control
  • infrastructure hardening
  • egress control and monitoring

So the pressure on the safety side is rising at the same time as the capability race accelerates. Code Red is not only about beating Gemini; it comes paired with the separate, harder question of how to release increasingly dangerous models into society.


6. How Far Have the Rivals Come?

A brief survey of the main competitors.

6-1. Google Gemini: playing on home ground, search × LLM

Google's Gemini 3 Pro takes first place overall on LMArena, as noted above, and Menlo Ventures' report also rates it top-tier across many benchmarks. (LMArena, Menlo Ventures)

Google has also integrated Gemini deeply into:

  • Search
  • Workspace (Docs, Sheets, Gmail)
  • Android

which is making it powerful as the place people first encounter an LLM. Against the brand OpenAI built with ChatGPT, Google is coming for users by rewriting the defaults in products they already use.

6-2. Anthropic Claude: twin peaks in enterprise and coding

Anthropic has:

  • held at or near the top of code-related benchmarks since Claude 3.5 Sonnet in 2024 (Menlo Ventures)
  • and, by late 2025, set new highs with Claude Opus 4.5 on difficult code benchmarks such as SWE-bench Verified (Menlo Ventures)

As noted above, the estimates that it holds:

  • 32% of the enterprise LLM API market
  • 42% of code generation

put considerable pressure on OpenAI. (GlobeNewswire)

Reuters has reported Anthropic's annualised run rate passing $3 billion, which puts it within reach of a third to a half of OpenAI's scale. (Reuters)

6-3. xAI's Grok, and China's DeepSeek and Qwen

xAI (Grok) shows top-tier capability, taking second place overall on LMArena with grok-4.1-thinking and fourth with grok-4.1. (LMArena)

Among Chinese players, DeepSeek and Alibaba's Qwen are rapidly gaining presence in community benchmarks and the open-source world. Menlo Ventures' analysis puts their share of enterprise production workloads at around 1% so far, but notes rising use on infrastructure such as vLLM and OpenRouter. (Menlo Ventures)

Open-source use remains around 13% of enterprise overall, with closed models still at 87% — but as a cheap, customisable option it stays impossible to ignore. (AI Matters)


7. Reading OpenAI's Next Move After Code Red (Fact plus Speculation)

From here, let us organise where OpenAI seems to be heading, from an engineer's point of view. Note that this section contains interpretation and speculation.

7-1. Move one: back to basics on ChatGPT, and personalisation

What the Code Red memo and the reporting around it consistently emphasise is:

  • the push into advertising goes on hold
  • protecting the dialogue loop between user and model is the top priority
  • and to that end, resources concentrate on personalisation, response speed, reliability, and the breadth of topics covered

(Business Insider, Financial Times)

This also reaffirms a long-run strategy:

grow ChatGPT itself into something like an AI operating system

7-2. Move two: enterprise focus and making ROI visible

OpenAI put quantified time savings and task expansion front and centre in the "State of Enterprise AI" report. (OpenAI CDN)

It has also brought in former Slack CEO Denise Dresser as Chief Revenue Officer, strengthening enterprise sales and monetisation. Per The Verge, Dresser led AI features at Slack and takes on the role of spreading AI into more businesses at OpenAI. (The Verge)

Given Anthropic's growing lead in enterprise, OpenAI's next move looks like building out an enterprise stack combining:

  • ChatGPT Enterprise / Business
  • comparatively cheap GPT-4o, 4o mini, and open models
  • coding-focused Codex and agent capabilities

Indeed, OpenAI's developer site gives prime placement to "Agents," "Codex," and "Open Models," suggesting that combining agent capabilities with open-model availability is the theme going forward. (OpenAI)

7-3. Move three: agents, and the Agentic AI Foundation

In December 2025, OpenAI announced it had joined the launch of a new non-profit, the Agentic AI Foundation, contributing an open guidelines document called AGENTS.md. (OpenAI)

Menlo Ventures also frames the next wave as long-horizon agents, predicting that systems which semi-autonomously handle multi-step tasks such as:

  • fixing code
  • producing research reports
  • automating business operations

will become the core of the enterprise AI stack. (GlobeNewswire)

So Code Red is not only about improving chatbot quality. It carries a strong sense of

retaking a platform position for the agent era

7-4. Move four: country programmes and sovereign AI partnerships

Country-level programmes such as OpenAI for Australia sell a bundle of: (OpenAI)

  • sovereign AI infrastructure including large GPU clusters
  • workforce development at the scale of millions of people
  • long-term partnerships with local companies and governments

Against the region-level infrastructure that conventional cloud vendors — AWS, Azure, GCP — have offered, this is a new sales motion:

a national package bundling AI-specific cloud, education, and ecosystem together

It can be read as OpenAI shifting from being a model vendor to being a national AI infrastructure company.

7-5. Move five: building safety into product value

Finally, safety.

Alongside the cyber-risk warning above, OpenAI has:

  • emphasised safe deployment cases and the importance of governance in its enterprise report (OpenAI CDN)
  • and worked to spread good practice through "AI Jam" programmes for small businesses and non-profits (OpenAI)

Under Code Red, embedding how to make AI safe to use into product design is likely to become as much a differentiator for OpenAI as strengthening models or expanding data centres.


8. Three Scenarios Engineers and PMs Should Hold in Mind

Building on all of the above, three scenarios worth keeping in view when you think about product and technical strategy.

Scenario A: Round two of the benchmark war

Given the arc —

  • 2023–2024: the GPT-4 generation sweeping most benchmarks
  • 2025: Gemini 3, Claude 4.5, Grok 4.1, and GPT-5.1 crowding together, with rankings in flux

(LMArena)

— it is likely that from 2026 the top model keeps changing hands every six months or so.

For engineers and PMs, that makes it increasingly important to:

  • avoid architectures that presuppose a specific vendor
  • keep an abstraction layer — an AI gateway, or your own adapters — that makes swapping models easy

Scenario B: Enterprise AI assumes multiple models

Menlo Ventures' research found vendor switching low at 11%, while upgrades to the newest model ran at 66%. (GlobeNewswire)

On the ground, though, multi-model setups should keep increasing:

OpenAI for general tasks, Anthropic for code, something else for images
Chinese or European models used alongside for regional or data-sovereignty reasons

In that world, OpenAI's Code Red is also a test of whether ChatGPT can be held in place as the central UX layer.

Scenario C: Energy and regulation become the hidden decider

Least discussed and most important: energy and regulation.

Consider:

  • a 250GW AI data-centre plan
  • $1.4 trillion in capex commitments
  • next-generation models that have to be classified "high" for cyber risk

(Tom's Hardware, blog.carnegieinvest.com)

and competition among AI companies is

a race to build the strongest model, and simultaneously a race to operate it efficiently and safely

Whether a company skilled at building relationships with regulators, utilities, and cloud vendors can deliver models that have grown too large while staying reconciled with society — that is where round two will be decided.


Conclusion: Code Red Is Not the End Signal, It Is the Opening of Chapter Two

As we have seen, OpenAI's Code Red reflects a genuinely difficult position:

  • world-class in revenue scale
  • but overtaken by Anthropic in enterprise share
  • ceding the top of the benchmarks to Gemini 3, Claude 4.5, and Grok 4.1
  • while infrastructure investment and safety risk grow exponentially heavier

Turn it around, though, and you could say:

the next few years are where the LLM war actually begins

For engineers and product people, the useful lens for following OpenAI and its rivals through this Code Red period is less

which model is the cleverest

and more

which model, in which configuration, at what level of risk

Related Articles

May 1, 2026

How Far Will AI Agents Refuse “Gray-Area Code” in 2026?

This article examines where AI agents refuse or assist gray-area automation (like social engagement bots), comparing policy intent and real behavior across OpenAI, Google, and Anthropic.

TechnologyRead more
September 6, 2026

GPT-6 Astra Arrives: What Changes When You Put It to Work in Codex?

Our first article produced with Astra examines GPT-6 in Codex, its differences from GPT-5.6, comparisons with Fable 5.1 and Opus 5, ARC-AGI-3 testing conditions, international reactions, and the AGI debate. Sources checked September 6, 2026.

TechnologyRead more
July 16, 2026

Can You Trust the "Best AI Coder" Rankings? OpenAI Audited SWE-Bench Pro, Found ~30% of Tasks Broken, and Retracted Its Recommendation

OpenAI audited SWE-Bench Pro, the leading coding benchmark, found that roughly 30% of its public tasks were flawed, and retracted its recommendation. Here is what actually happened, whether the score gaps between GPT, Claude and Gemini are real, why models get "trained to the test," and how to measure practical ability instead.

TechnologyRead more
March 4, 2026

What Is GPT-5.3 Instant?

A structured explainer of GPT-5.3 Instant covering design goals, latency engineering, accuracy measurements, safety trade-offs, and practical positioning based on public disclosures.

TechnologyRead more
September 21, 2025

From API Keys to Web Integration — A Hands‑on Guide to OpenAI, Anthropic Claude, and Amazon Bedrock

A practical guide for integrating generative AI APIs into real web apps. Covers key acquisition, auth, minimal code, pricing basics, safe Next.js patterns, and operations best practices.

TechnologyRead more