Back to Intel Feed
Research10 min Analysis

OpenAI Loses $33B a Year. Its Fix Is the Most Dangerous Thing for Your Job.

M
Meet

August 4, 2026

OpenAI Loses $33B a Year. Its Fix Is the Most Dangerous Thing for Your Job.

Executive Intelligence Summary

  • 01

    The thing protecting most jobs right now is not skill — it's price. AI is capable of far more than companies currently pay it to do.

  • 02

    OpenAI's own numbers show why: revenue up 92% to a $25B run rate, yet 2026 losses tracking to $33B, with roughly $14B of that burned purely on inference.

  • 03

    Jalapeño, OpenAI's Broadcom-built inference chip, is claimed to deliver ~50% lower cost per token than current Nvidia GPUs — cutting the price of every automated task in half.

  • 04

    Cost per task is already collapsing inside one vendor's own lineup: about $3.15 on Claude Fable 5, $1.86 on GPT-5.6 Sol, and roughly $0.07 on Luna. When a task costs 7 cents, nobody writes a business case to automate it.

There is a line from an OpenAI investor pitch that should be far more famous than it is:

"We have no current plans to make revenue. We have no idea how we may one day generate revenue. We have made a soft promise to investors that once we've built this generally intelligent system, we will ask it to figure out a way to generate an investment return for you."

For three years, that was not a joke. It was the business model.

And yet in 2026, OpenAI's revenue exploded by 92% to a run rate of $25 billion a year — while the company is on track to lose $33 billion. In January 2025, Sam Altman admitted they were losing money even on the $200/month Pro subscribers.

Read that again. In every normal business on earth, your best customers are your most profitable customers. In OpenAI's world, the best customers are the least profitable ones.

Wall Street looked at this and called it an AI bubble. The thesis is brutally simple: if the biggest and best player cannot make money, how will anyone else? And if nobody makes money, why is everyone pouring trillions into this?

But there was always one escape hatch. One way the bubble turns into a genuine gold rush: somebody has to collapse the cost of intelligence by 10 to 100 times.

In the last 90 days, that collapse started. And this is where it stops being a finance story and becomes a career story — because the only thing standing between most professionals and automation right now is not capability. It's price.

First, understand the actual problem: inference

The whole nightmare comes down to one word.

Think of a chef. Training a model is sending that chef to culinary school. You do it once, it costs a fortune, and the chef comes out with extraordinary skills. This is why OpenAI spent billions making the smartest models on earth.

Inference is what happens every time a customer orders a dish. The stove fires up, the food gets cooked, and it costs money — every single time. Every question you ask, every agent you run, spins up GPUs and burns cash.

Training happens once. Inference happens forever.

Now multiply that by roughly 900 million people placing billions of orders a week, and you get the number that defines the industry: OpenAI is set to burn around $14 billion on inference alone in 2026.

Why the best customers lose the most money

Take a heavy user — a software engineer who pays $20 a month and genuinely loves the product. She uses it to write code, debug, draft emails, plan trips, and think through decisions.

Analysts estimate a heavy user firing ~100 queries a day can burn through more than $30 of compute in a month.

She pays $20. She costs $30. OpenAI loses $10 on her every month — and the more she loves the product, the more she uses it, and the deeper the hole gets.

So why not just charge her $40? Because OpenAI is caught in two traps.

Trap 1: Competition is racing the price of intelligence to zero

The price collapse in this industry has no precedent in modern business. GPT-3-level intelligence cost roughly $60 per million tokens when it first shipped. By late 2024, that same level of capability cost about 6 cents — a ~1,000x collapse in around three years.

Google made Gemini free at the consumer tier. Anthropic cut prices dramatically in a single announcement. In this market, if OpenAI raises the price to $40, the customer switches to Claude or Gemini in thirty seconds.

You cannot raise the menu prices.

Trap 2: OpenAI doesn't own the kitchen

Jensen Huang describes AI as a five-layer cake: energy, chips, infrastructure, models, and applications. Look at who actually owns each layer.

LayerWho owns itOpenAI's position
EnergyPower plants and grid operatorsZero control, rising costs
ChipsNvidia, running ~75% gross marginsFor every $1 spent, ~75c is Nvidia's profit
InfrastructureMicrosoft and hyperscaler data centresRented capacity
ModelsOpenAIThe one layer it owned
ApplicationsStartups, enterprises, builders like youSomeone else captures the end value

Every ChatGPT answer runs on Nvidia's chips, in Microsoft's data centres, powered by someone else's electricity. OpenAI has been a restaurant that doesn't own the kitchen, the ovens, or the building — and can't raise its prices.

There is exactly one way out of that trap. Stop renting the kitchen. Build your own oven.

The 90-day comeback nobody was ready for

While everyone stared at the losses, OpenAI shipped four things in quick succession. Individually they read as product news. Together they are a single, coherent escape plan.

1. A model that matched the leader at a third of the cost

For most of 2026, the coding crown belonged to Anthropic. Claude Fable 5 was good enough that developers happily paid $10 per million input tokens, and Anthropic's ARR ran to roughly $47 billion against OpenAI's $25 billion.

Then on July 9, 2026, OpenAI shipped GPT-5.6 with Sol, Terra, and Luna.

Most coverage treated it as another version bump. The efficiency numbers say otherwise:

MetricClaude Fable 5GPT-5.6 SolGPT-5.6 Luna
Cost per representative task~$3.15~$1.86~$0.07
Coding indexBaseline leaderBeats Fable 5Volume tier
Agentic indexBaseline leaderTops Fable 5Sub-agent tier
Time to completionBaseline~60% lessFastest

Sol reaches near-frontier intelligence at roughly half the cost and 60% less time. Terra beats Claude Opus 4.8 on several axes. And Luna does a comparable task for about seven cents.

The market reaction told the real story. Anthropic had Fable 5's subscription access set to expire on July 7. It got extended to July 12. Four days after Sol landed, it was extended again to July 19. Read between the lines: it looks a lot like Anthropic had planned to pull its best model out of subscription tiers to protect margins — and then a competitor shipped something close enough in quality and dramatically cheaper, and the plan changed.

2. Codex turned questions into work

For two years the deal with AI was simple: you type a question, it gives you an answer.

Codex breaks that. It writes the code, builds the tool, ships the site, runs the multi-step job, and comes back when it's finished. Active users exploded by 500%. Steve Jobs called the computer a bicycle for the mind — Codex is closer to an autopilot for it, and OpenAI confirmed it stays included on every ChatGPT plan.

Marketers, PMs, lawyers, analysts, researchers: anyone can now commission software instead of requesting it.

3. GPT Work folded the desktop into one screen

Email, docs, files, and Codex in a single surface. The significance isn't the UI — it's that the agent now sits where the work actually lives, with the context it needs to act without being asked twice.

4. Jalapeño: OpenAI finally owns a second layer

Jalapeño is OpenAI's first custom silicon — an inference-only chip designed with Broadcom. Broadcom's CEO has said early tests show roughly 50% lower cost per token than current Nvidia GPUs, with performance on par with Nvidia's best Blackwell parts.

If that claim holds, the $14 billion inference bill gets cut close to half over the next few years.

So the model, the agent, and the chip are not three stories. They're one:

  • Sol creates demand with a product too good and too cheap to leave.
  • Codex blows up usage, turning every user from a question-asker into an agent-runner.
  • Jalapeño halves the cost of every single answer so the first two don't bankrupt the company.
  • Here's the part that concerns your career

    Every analysis of this story stops at "will OpenAI become profitable?" That's the wrong question for anyone who isn't an investor.

    The right question is: what happens to the automation decision when a task costs seven cents?

    Right now, most jobs are not protected by capability. They're protected by cost and friction. A manager doesn't automate a workflow because the model can't do it — they don't automate it because the migration is a project, the tokens add up, and the ROI case is annoying to write. That friction is the actual moat, and almost nobody admits it out loud.

    Cheap inference dissolves it. Consider what a 10x cost collapse does to the arithmetic:

    Task costWhat it takes to approve automation
    $3.15 per taskA business case, a budget line, a pilot, an exec sponsor
    $1.86 per taskA team lead's discretionary spend
    $0.07 per taskNothing. Below the noise floor of any budget

    At $3.15 a task, automating you is a project. At 7 cents, it's a default. Nobody convenes a meeting to approve a rounding error — and there is no procurement process to slow down, no sponsor to lobby, and no ROI review where a human advocate might speak up for the team.

    Then there is Codex, which changes the shape of what gets automated. Chat-era AI replaced *sentences* in your workflow. Agent-era AI replaces the *workflow*. That's why one Codex task can cost what a hundred ChatGPT questions cost — and why cost per token becomes the single most important number in the labour market.

    The trap in the comeback: Jevons paradox

    There is one more wrinkle, and it's the reason none of this stabilises.

    Everyone assumes cheaper AI means OpenAI saves money. It doesn't. When you make something 10x cheaper, people don't use less of it — they use insanely more of it. That's Jevons paradox, and it's been true of coal, electricity, bandwidth, and cloud storage.

    Which means the cost curve never gets to rest. OpenAI cuts inference costs, usage explodes past the savings, and the pressure to cut costs again returns immediately. The whole industry is locked into a permanent race to push the price of thinking toward zero.

    For your career, the implication is uncomfortable but clear: this is not a wave that crests. There is no version where the cost of automating knowledge work goes back up.

    Who is most exposed

  • Execution-layer knowledge workers. If your value is turning a defined brief into a deliverable — decks, reports, first-draft code, data pulls, QA passes — you are competing directly with the tier that just hit seven cents.
  • Junior and mid-level engineers. Codex doesn't need a ticket refined into submission. It needs a prompt and a repo.
  • Analysts and ops roles. Multi-step business processes were expensive to automate for exactly one reason: compute. That reason is being deleted.
  • Anyone whose moat is "it's not worth the cost to replace me." That moat is being priced out of existence on a quarterly cadence.
  • The defensive playbook

  • Price yourself against agents, honestly. Take one week of your output and estimate it as a stack of agent tasks. That number is what a finance team will eventually compute. Better you see it first.
  • Move from execution to specification and verification. Cheap inference makes doing the task nearly free and makes *choosing the right task* and *proving it was done correctly* far more valuable. Sit on the side of the trade that appreciates.
  • Learn orchestration economics. Routing volume work to Luna-class models and judgment work to Sol or Fable-class models is a management competency now. The person who designs that routing captures the savings instead of being the savings.
  • Compound the things that never get cheaper. Accountability, client relationships, regulatory sign-off, taste, and the political work of getting a decision made inside a real organisation. None of these have a cost-per-token curve.
  • The honest closing

    There is a genuinely optimistic reading here, and it deserves stating. Trillion-dollar companies are burning their balance sheets to demolish the barrier between an idea and a working thing. If you have ever wanted to build something and been stopped by cost, time, or not knowing how — that constraint is dissolving in real time. When thinking is nearly free and building is nearly free, an infinite canvas opens up for anyone willing to use it.

    But that canvas only rewards people who shift from operating to directing. The same price collapse that lets you build a company from a prompt lets your employer replace a function with one.

    Same curve. Two completely different outcomes. Which one you get depends on whether you spend the next twelve months proving you can execute — or proving you can decide.

    Coming soonFree for a year

    Applying is 12 minutes of typing you've already done.

    Job Autofill fills every application from your saved profile - about 10 hours back over a job hunt.

    Get early access

    Professional Defense Strategy

    6-Month Strategic Action Plan

    1

    Reprice your own role honestly: estimate what your weekly output would cost as agent tasks at $0.07-$1.86 each. If the gap between that number and your salary is your only moat, you have 12-18 months, not five years.

    2

    Move up the stack from execution to specification. Cheap inference destroys the value of doing the task and inflates the value of deciding which task is worth doing and verifying that it was done right.

    3

    Get fluent in orchestration economics — routing cheap models to volume work and frontier models to judgment work is becoming a management skill, not an engineering one. Be the person who designs that routing.

    4

    Own the layers agents can't touch: accountability, client trust, regulatory sign-off, and the messy political work of getting a decision approved inside an organisation. These are the only line items that don't get cheaper when tokens do.

    Frequently Asked Questions

    Why is OpenAI losing money if its revenue is growing 92%?

    Because revenue and cost scale together. OpenAI's costs are dominated by inference — the compute burned every time a user sends a query or runs an agent — rather than by one-off training runs. In 2026 the company is on track for roughly $14 billion in inference costs alone against a $25 billion revenue run rate, with total losses tracking to about $33 billion. Growth adds inference cost at close to the same rate it adds revenue, so scaling up does not automatically close the gap.

    What is the difference between training cost and inference cost?

    Training is the one-time cost of building a model — the culinary school tuition for the chef. Inference is the recurring cost of running it — firing up the stove for every single order. Training is enormous but finite. Inference is smaller per unit and effectively infinite, which is why it dominates the economics of any AI company serving hundreds of millions of users.

    How can OpenAI lose money on $200/month Pro subscribers?

    Subscriptions are flat-rate while usage is not. Sam Altman said in January 2025 that OpenAI was losing money on Pro subscribers because the heaviest users consume far more compute than their fee covers. Analysts estimate a heavy user running around 100 queries a day can burn more than $30 of compute a month on a $20 plan. Agentic tools make this worse: a single Codex task can cost what a hundred chat questions cost.

    What is Jalapeño, OpenAI's custom chip?

    Jalapeño is OpenAI's first custom silicon — an inference-only chip designed with Broadcom. Broadcom has said early tests show roughly 50% lower cost per token than current Nvidia GPUs, with performance comparable to Nvidia's best Blackwell parts. Strategically it moves OpenAI from renting one layer of the AI stack to owning two: models and chips.

    Is the AI bubble about to burst?

    The bear case is that if the largest and best-funded player cannot turn a profit, nobody downstream can either. The bull case is that the cost of intelligence collapses 10-100x, at which point today's losses become tomorrow's margins. The last 90 days favour the bull case: GPT-5.6 Sol matched frontier quality at roughly a third of the cost, and Jalapeño targets halving inference cost again. Jevons paradox complicates it — cheaper compute drives usage up faster than costs come down, so profitability keeps receding even as unit economics improve.

    Does cheaper AI make my job safer or more at risk?

    More at risk. For most professionals the barrier to automation today is cost and friction, not capability. When a task costs $3.15, automating it needs a budget line, a pilot, and an executive sponsor. At $0.07 it falls below the noise floor of any budget and gets adopted by default, with no ROI review where anyone might argue on your behalf. Falling prices remove the friction that has been quietly protecting a great many roles.

    What is Jevons paradox and why does it matter for careers?

    Jevons paradox is the observation that making a resource cheaper increases total consumption rather than reducing it — it held for coal, electricity, bandwidth, and cloud storage. Applied to AI, it means cheaper tokens produce far more AI usage, not the same usage at a lower bill. For careers, the implication is that this is not a wave that crests: there is no future state in which automating knowledge work becomes expensive again.

    Which jobs are most exposed to collapsing inference costs?

    Execution-layer knowledge work is the most exposed: turning a defined brief into a deliverable such as decks, reports, first-draft code, data pulls, and QA passes. Junior and mid-level engineers, operations and process analysts, and anyone whose job security rests on the argument that replacement is not worth the cost. Roles built on accountability, client trust, regulatory sign-off, and organisational judgement hold their value because none of those have a cost-per-token curve.

    Researcher Intelligence

    Automation Intensity

    96%

    Overall Substitution Risk

    Critical

    Grounding

    Primary Technical Launch

    Status

    Market Active

    "Market saturation for Research intelligence is expected to accelerate significantly following this window."

    #OpenAI#Jalapeño#Broadcom#Inference Cost#GPT-5.6 Sol#Codex#AI Bubble#Jevons Paradox#Career Security#Automation Economics
    Free Tool

    Land Your First Interview Call Faster.

    In the middle of job loss chaos, don't get lost in the pile. Resume Maximiser scores your resume in 30 seconds and shows you exactly why it is being filtered out.

    More Critical Intelligence