RSI Is the Platform
The unit economics of tokens versus the pipeline, and which one justifies the AI labs' valuations.
Listen: references
Listen: sec1
Listen: sec10
Listen: sec11
Listen: sec12
Listen: sec2
Listen: sec3
Listen: sec4
Listen: sec5
Listen: sec6
Listen: sec7
Listen: sec8
Listen: sec9
The unit economics of tokens versus the pipeline — and which one justifies the AI labs’ valuations
1. Nobody argues about AGI anymore
A year ago every AI conversation ended in AGI. Now the serious conversation is about RSI: recursive self-improvement.
That shift matters, because RSI is the first version of the story concrete enough to build a business on. AGI was a destination with no road. RSI is a road. Anthropic defines it as an AI system capable of fully autonomously designing and developing its own successor, and it is already partway there — by Anthropic’s own numbers, more than 80% of the code merged into its production codebase is written by Claude, and its engineers ship roughly 8x as much code per quarter as they did between 2021 and 2025 (When AI builds itself). OpenAI says its agents now produce about 3.1 agent-workdays per human workday (Research acceleration). Jack Clark puts a 60% chance on automated AI research by 2028 (Import AI 455).
So the question is no longer “will AGI arrive?” It is “what does the self-improving pipeline let you sell?”
2. The AGI economy was a dead end
Take AGI literally and there is no economy left to discuss. If a system can do all cognitive work, and then build robots to do the physical work, then all jobs are automated — and you get one of two fictional outcomes: a totalitarian world, or a paradise of abundance.
Interesting as philosophy. Useless as business analysis. DeepMind now treats AGI as a starting condition, not a finish line, with plural and uncertain pathways after it (From AGI to ASI). Kapoor and Narayanan are blunter: AGI is not a milestone at all; value arrives through diffusion, not a moment (AGI is not a milestone).
RSI is different, because it produces something you can price every month. This article does that pricing — first for the token business as it exists today, then for the platform business RSI would create, and finally the valuation question.
3. Even a frontier model can’t do the work on day one
Suppose you have the best model in the world and unlimited tokens. What can you actually do with it?
Consider an insurance company processing claims. The model cannot do the work from day one. It needs onboarding: the workflows, the systems, the edge cases, the history. It tries, fails, corrects. First run is rough; by the third to fifth run it is reliable — assuming nothing changes. Now do the same for a hospital doing triage, or a pharma team screening compounds: same pattern. The model is not the product. The onboarding is.
Today, that onboarding is mostly context engineering — modules talking to modules, a complex but navigable system, held together by people like me. It works, and it is expensive, and it is exactly the kind of work RSI absorbs.
4. Maps, not territory
The right mental model is map-making.
You face a new problem. The territory is unmapped. You draw a rough map — inaccurate, but better than nothing. You walk the terrain, hit surprises, improve the map. At some point further mapping stops paying: the value of another 1% of accuracy is smaller than the cost of getting it.
Every company already has such a map, at exactly the quality that was economically rational for humans to draw. Now the model enters, navigates with your map, and finds improvements — things your people never encountered, or encountered too rarely to document. Should you improve the map? Yes, because the model has lowered the cost of improving it. And once you start, it makes sense to redraw it properly: from a flat runbook to a rich, machine-usable map that holds more complexity than any human document ever could.
That is what RSI is: the map keeps getting better because the cost of better maps collapses.
5. The economics of one model — the token business
Let us do the actual arithmetic for one frontier model, sold the way models are sold today: by the token. We will use a GPT-6 “Astra”-class release, because it is the best-documented 2026 flagship: pretrained on more than 100,000 Nvidia GB200-class GPUs, priced at $10 input / $50 output per million tokens (Epoch AI).
What it costs to build one. No lab discloses per-model training cost, so every number is an estimate, and the estimate spread is wide — three to seven times, depending on whether you price compute at owned cost or cloud-rental rates, and whether you count failed runs.
The best-documented anchor is Epoch AI’s estimate for Grok 4: $490 million for the final training run alone (Epoch AI). For a 2026 flagship, the final run lands around $0.3–2.5 billion ($0.3–0.5B low, $0.7–1.1B base, $1.5–2.5B high). But the final run is only part of the program. Epoch finds final runs are roughly 9.6% of a lab’s R&D compute, implying a 4–10x multiplier once you add experiments, data, staff, safety, and the runs that failed. Fully loaded, one 2026 flagship costs roughly $1.5B (low) / $3–6B (base) / $8–15B (high) through launch (research summary).
What it costs to serve one. Here the numbers are better than most people assume. With modern batching, caching, and hardware, a frontier model’s true blended serving cost is around $1 per million tokens — SemiAnalysis puts Opus 4.7 at $0.99, after accounting for ~90% cache hits and a 300:1 input-to-output ratio (SemiAnalysis, via research). Smaller 70B-class models serve for $0.02–0.35/M.
What it earns. Sticker prices are $10/$50 for Astra, $5/$25 for Claude Opus 5, $2/$6 for Grok 4.6. Realized prices are lower after caching and committed-use discounts: Anthropic’s books imply about $5.6 per million tokens, OpenAI’s blended all-surface figure is around $3.4/M, and enterprise effective pricing fell about 6% in 2026 after 39% in late 2025 (revenue research). At a $5/M realized price and $1/M serve cost, the unit margin is 80% — and that is consistent with Anthropic telling investors its gross margin is above 80% before distribution and training costs (Reuters/FT via CNBC).
So the unit economics of a token are good. The problem is everything above the unit.
What has to be true to break even. A lab does not carry one model; it carries a fixed cost base of compute, staff, and research. Attribute $8B, $20B, or $40B per year to one flagship and divide by the $4/M contribution:
| Cost base / year | Break-even volume | Share of the entire paid API market (~100T tokens/day) |
|---|---|---|
| $8B | ~4.4T tokens/day | 4.4% |
| $20B | ~15.1T tokens/day | 15.1% |
| $40B | ~123T tokens/day | 123% |
That middle row is the whole story. One model, at a realistic 80% gross margin, must capture 15% of the entire paid token market to cover its cost base. The high row is arithmetically impossible: no model can sell 123% of the market.
We can make the treadmill explicit (serve cost assumed to fall with the model class: $0.50/M at $6, $1.00/M at $4–5, $1.50/M at $2–3):
| Realized price/M | $8B base | $20B base | $40B base |
|---|---|---|---|
| $6.00 | 5.0% | 11.0% | 20.9% |
| $5.00 | 6.8% | 15.1% | 28.8% |
| $4.00 | 9.1% | 20.1% | 38.4% |
| $3.00 | 13.7% | 30.1% | 57.5% |
| $2.00 | 54.8% | 120.5% | 230.1% |
And the floor is dropping. Open-weight models already route at $0.05–0.30 per million tokens effective; DeepSeek’s V4-Flash is priced at $0.14/$0.28 (revenue research). Run the same model at $0.20/M with a $0.15/M serve cost and break-even requires 1,205% of the paid market. That is what commoditization does to the arithmetic: it does not shrink margins, it makes the business impossible.
6. What the token numbers actually say
Three conclusions.
First, the token business is a scale game, not a margin game. The margin per token is fine; the volume needed to carry a frontier lab’s cost base is enormous, and it is growing.
Second, the treadmill is baked in. Training costs rise about 3.5x per year, while the price of a fixed level of capability falls 5–10x per year. To stand still, a lab must grow volume by roughly an order of magnitude annually. Volume is growing — global demand roughly 10x per year against supply growth of about 3.4x — which is why the business still works for the leader.
Third, only the best enterprise mix survives it. Anthropic, at a $65B run rate and an enterprise-first book, just posted its first positive adjusted operating quarter. OpenAI, at roughly $40B with a consumer-heavy mix, has not. Same technology, different position on the cost curve. The audited OpenAI gross margin for 2025 was 42.6%, and Epoch estimates the GPT-5 bundle at about 30% gross margin once R&D is counted — that is the difference between a viable business and a subsidy (Epoch AI).
The token business is not worthless. It is a treadmill, and only one or two labs can stay on it.
7. The platform business — sell the pipeline, not the weights
Now the alternative. Instead of selling tokens from a shared model, sell each customer a model of their own — personalized, continuously refreshed, running on the customer’s inference budget — and sell the pipeline that keeps improving it.
The industry has already started. Variant Fund calls these dedicated model foundries and borrows the TSMC playbook: unbundle training, aggregate demand, and pour the proceeds into improving the training process (Variant). Applied Compute argues the model, harness, and application have fused, so the defensible asset is the feedback loop (Applied Compute). Prime Intellect ships “Lab,” a platform whose product is the training loop, sold per token (Prime Intellect). Thinking Machines gave away its base model and monetizes Tinker — whose model once ran and evaluated its own fine-tuning job in 27 minutes (TechCrunch).
What a personalized model costs and sells for. The striking finding: the compute is trivial. A 35B RL specialist was trained for about $500 in 26 hours on rented compute. The real cost is data work, evals, and engineering: roughly $15–25K per model self-serve, $150–300K for a base foundry engagement, and $1–2M for a bespoke build (platform research).
Prices are public. Tinker charges $0.44 per million training tokens for an 8B model (a 2.5–5x markup over DIY cost); Prime Intellect charges $0.60/M for 9B and $4.00/M for 397B. Prime Intellect reached $100M+ ARR with 6,000+ customers in under a year — an ARPU around $17K, which tells you the self-serve tier is volume, not whale, business. Enterprise deals for private models run $1M+ per year.
What has to be true to break even. A platform carries a $5B one-time build and roughly $2B/year to maintain and improve the pipeline. Then:
| Tier | Price/model | Cost/model | Margin | Break-even customers |
|---|---|---|---|---|
| Self-serve (Tinker/Lab-like) | $20K | $10K | 50% | 200,000 |
| Base foundry (enterprise) | $1M | $400K | 60% | 3,333 |
| Bespoke (large enterprise) | $5M | $2M | 60% | 667 |
Read those two tables together and the strategic picture is clear. The token business needs double-digit share of the entire market to carry its cost base. The platform business needs 3,333 enterprise customers at $1M each — or 667 at $5M — and it gets there by selling to companies individually, not by winning a price war for a commodity.
The honest caveats: fine-tuning and RL are only about 1% of enterprise AI spend today ($4B of $37B), and only about 19% of the Global 2000 pays for AI at all (Menlo/a16z, via research). The platform thesis is a bet that RSI makes personalization so cheap and so valuable that this share grows — and that the labs capture it before the challengers do.
8. Which business justifies the valuation
Anthropic’s run rate passed $65B in mid-2026; OpenAI’s is around $40B. At a 15x revenue multiple, Anthropic’s $965B valuation needs about $64B — the current run rate already carries it. The question is the terminal state, and that is where the two businesses separate.
The token business can justify the current multiple for the leader, at current volumes and mix. It cannot easily justify the growth to Anthropic’s $190–200B 2028 revenue projection, because that projection requires either far more volume, higher prices, or both — against a deflationary price trend and a credible open-weight floor. Damodaran’s test makes the same point from the other side: a $5T valuation needs $5T of annual revenue, against roughly $250B today.
The platform business does not have the same market-share cliff. Its revenue scales with customers, its gross margin is 50–60% before RSI even lands, and the customer funds the training compute. If RSI drives the cost of personalization down — and the evidence says it will — those margins go up, not down, and the model count per enterprise goes from one to several.
So the valuations, read honestly, are a bet on the platform. The token business pays today’s bills. RSI is what justifies the terminal value. That is the strongest available version of “not a bubble”: a bubble in AGI-as-end-state can coexist with a real business in RSI-as-platform.
9. The two objections that actually matter
The strongest objection comes from Variant itself: if in-context learning is good enough to absorb most proprietary know-how, then renting tokens may be sufficient, and the foundry thesis shrinks (Variant). The second, from the other side, is the harness argument: if the enterprise owns the context, tools, and memory around the model, the model commoditizes and the enterprise captures the value (Own the Harness).
Both are real constraints and neither is fatal. Context is valuable and will stay valuable — the question is who can improve it fastest. A self-improving pipeline improves the map continuously, including the parts outside the weights. The customer’s context does not disappear; it becomes an input the platform processes better than the customer can alone. The bet is not that context stops mattering. It is that keeping your own map up to date becomes the more expensive option.
If that bet is wrong, the platform is a services business with good margins, not a compounding one. That is the honest downside, and it should be stated.
10. Three businesses, and the research they fund
Once you see the platform, the labs end up with three businesses:
- Training / RSI platform. Help companies build and refresh their own models.
- Inference. Run those models cheaply, reliably, everywhere.
- Domain integration. Partner deeply in medicine, materials, research — learn how the work is actually done, and reshape it so AI can do it natively.
On top sits continued frontier research toward AGI: new architectures, new capabilities. That work is high-risk and not directly profitable, and it is funded by the three businesses above. This is where the safety question resolves cleanly. Anthropic is right that RSI raises genuine control risks (CNBC). The answer is not to stop building the platform; it is to fund the risky research from the platform’s profits and ring-fence it, rather than betting the company on AGI arriving.
Inference deserves special mention, because it is the bridge. It is already 40–50% of lab revenue and more than 60% of compute spend, at gross margins around 70% for the best operators (research). It works with either business model, and it is the part that makes the platform practical: the customer’s models need somewhere to run.
11. Not a bubble, not AGI
The bull case says this is a genuine platform shift, and it is self-funding. Blackstone calls the spending “exactly the opposite of a bubble,” and Goldman’s caution is about an earnings bubble, not a valuation bubble (Goldman). The bear case is not that AI is fake; it is that the price is wrong, and that the cost base cannot be covered.
The arithmetic in this article says both are partly right. Tokens alone: good unit margins, brutal volume requirements, deflating prices — viable for one or two leaders, not a foundation for a $2T terminal value. Platform plus inference: lower revenue per unit, higher margins, no market-share cliff, and a cost structure that improves as RSI lands. Neither outcome requires AGI. Both are compatible with it.
That is why the AGI debate was the wrong frame all along. AGI is research. RSI is the product.
12. What this means
The smart money has already started moving from tokens to pipelines, foundries, and platforms. The startups are doing it. The incumbents have the compute, the customers, and the talent to do it better, if they choose the platform over the dream.
Bet on the lab that builds the pipeline everyone else uses to make their own models smarter — and the inference engine they run them on. Not on the one that promises AGI. Let the research be research, and fund it from the business that actually compounds.
References
- Anthropic, When AI builds itself (2026)
- OpenAI, Research acceleration (2026)
- Jack Clark, Import AI 455 (2026)
- Google DeepMind, From AGI to ASI (2026)
- Kapoor & Narayanan, AGI is not a milestone (2025)
- Epoch AI, Grok 4 training resources (2025)
- Epoch AI, Can AI companies become profitable? (2026)
- Epoch AI, LLM inference price trends (2025)
- Variant Fund, Towards Dedicated Model Foundries (2026)
- Applied Compute, Moats Need Models (2026)
- Prime Intellect, Introducing Lab (2026)
- TechCrunch, Thinking Machines and Inkling (2026)
- Narayanan & Kapur, Up the Stack (2026)
- Nate Jones, Own the Harness, Not the Model (2026)
- Galaxy, When Intelligence Commoditizes (2026)
- Reuters, Anthropic’s $190–200B 2028 forecast (2026)
- CNBC, Anthropic Q2 revenue and margin (2026)
- Goldman Sachs, The real AI risk is an earnings bubble (2026)
- CNBC, Existential concerns at Anthropic and OpenAI (2026)
The full model, with every input and the sensitivity tables, is
in projects/rsi-platform/model/econ_model.py.