You're Betting on the Layer That's About to Be Free
Find the line in your AI stack before your next vendor contract does.
Hey Transformation Leader,
Every few months, a headline tells you AI just got dramatically cheaper. A new model launches at a fraction of the previous price.
Executives read it as good news, because in almost every other part of business, falling input costs are good news. Cheaper steel helps the carmaker. Cheaper freight helps the retailer, and so cheaper intelligence should help everyone who buys it.
Except it doesn't work that way, and the companies finding this out right now are finding out the hard way.
Source: https://tokencost.app/blog/ai-price-index
Roughly 40 percent of AI startups have shut down within roughly 24 months of launching, and the trajectory points toward something close to 80 percent by the end of this year.
Nobody generally picked a bad model in these startups, but what killed them was believing that access itself was the advantage.
Last week I introduced Framework 01, the Intelligence Supply Chain, and its central law: value migrates toward the constraint, not toward the customer.
As I thought about that framework a bit more, I realized that law tells you where value ends up, but does not tell you how fast it gets there or why some layers of the stack empty out in months while others take years to show even a crack.
That is the missing mechanism, and it is the entire subject of Part 2 today.
Here is the short version, and then I will earn your trust with evidence.
Some layers of the Intelligence Supply Chain commoditize. They get cheaper, more interchangeable, and less defensible every quarter.
Other layers compound.
They get more valuable, more entrenched, and harder to replicate every quarter they operate.
The single highest leverage skill a leader can build right now is learning to tell these two categories apart before signing the next contract, not after reading the shareholder letter that explains why the AI investment didn't pay off.
Table of Contents:
The Cheaper It Gets, the More It Costs Someone
Two Datasets, Same Eighteen Months, Opposite Directions
The Line That Runs Through Every AI Stack
Same Model, Different Layer, Completely Different Outcome
The Honest Complications, and Why They Matter More Than the Headline
The Instrument: Scoring Your Own Stack
The Takeaway, and What Comes Next
Two Datasets, Same Eighteen Months, Opposite Directions
You do not have to take my word for any of this.
Watch two numbers move in the same window of time, and the mechanism explains itself.
The first number is the price of intelligence. Cost per million tokens has fallen roughly 300 times since 2023.
Source: https://benchlm.ai/llm-pricing-trends
Anthropic's own pricing tells the story cleanly. Claude Opus 3 launched in early 2024 at $15 per million input and output tokens.
Claude Sonnet 4.5 now delivers comparable quality at $3, an 80% drop in 18 months for equivalent capability.
Google cut Gemini 1.5 Flash's price 78% just three months after launch.
GPT-4o mini landed at a 92% discount to GPT-4's original price, and GPT-4.1 Nano pushed more than 99 percent below it.
DeepSeek V3 arrived from outside the usual frontier labs at roughly one-hundredth of GPT-4's launch cost.
It is an industry-wide pattern, repeating at every model generation, and the cycle time between generations is itself compressing, from about 18 months in 2023 down to roughly 12 months now.
The second number is the price of everything underneath the model.
Utility interconnection queues in the largest 2026 data center markets, Northern Virginia, Phoenix, Dallas, now run four to seven years.
A campus that joined the Northern Virginia queue this spring cannot expect utility power before 2030. PJM alone is sitting on more than 2,600 gigawatts of pending interconnection requests.
GPU lead times have stretched to 36 to 52 weeks, and the deeper bottleneck isn't even the chip itself; it is TSMC's CoWoS advanced packaging, the process that bonds high bandwidth memory to the die, which is fully booked through at least mid-2027.
Microsoft, Google, Meta, and Amazon have already placed multi-billion-dollar forward orders that consume most of Nvidia's allocation into 2027, which means any organization that didn't pre-commit compute in 2025 is no longer executing a plan.
It is reacting to scarcity.
Put those two numbers side by side, and you are not looking at two unrelated news stories. You are looking at one mechanism, captured twice, in the same eighteen months.
The model layer is collapsing in price because anyone can build it and everyone eventually will. The infrastructure layer is tightening in supply because almost nobody can build it and very few have tried in time.
Convergence and scarcity are not opposite forces. They are the same force, read from two different ends of the same stack.
The Line That Runs Through Every AI Stack
Some layers of your AI stack are racing toward free as we noted earlier.
These are your Commoditizing Layers.
Other layers do the opposite, and the longer you operate them, the harder they become to replicate, and the more value they accumulate, and that’s the Compounding Layers.
Between them sits a boundary that I call the Commoditization Line.
The single most useful thing you can do this quarter is figure out exactly where that line runs through your own organization’s AI stack because every dollar you spend on the wrong side of it is working against you.
Picture the Physical Intelligence stack from Part 1: power and grid at the bottom, then silicon, then data centers, then foundation models, then orchestration and fine-tuning, then applications, with proprietary data and workflow ownership sitting alongside as the layer that can buck the pull in either direction.
The Commoditization Line moves through that stack, and right now, in 2026, it sits roughly between foundation models and everything below them. Power, silicon, and owned infrastructure compound. Model access, prompt engineering, and thin application wrappers commoditize.
But do not mistake this for a single clean cut.
The research is explicit on this point, and so should you be. Within the model layer itself, open-weight models like Llama, Qwen, and GLM are commoditizing faster than frontier proprietary models, because there are more of them competing on the same benchmarks.
Within infrastructure, generic colocation space commoditizes while owned power generation and GPU allocation locked in before 2025 compound, because the owner captured a scarce position before the queue got to seven years.
The line does not run cleanly between named layers but runs inside them. Any executive who treats "foundation models" or "infrastructure" as internally uniform categories is already misreading their own stack.
There is a historical precedent for exactly this pattern, and it is sharper than the usual electricity and oil comparison. When IBM launched the original PC, it did not build the operating system or the microprocessor in house.
It licensed both, from Microsoft and from Intel. Within a handful of years, industry value migrated almost entirely away from the computer assemblers, IBM included, and toward those two upstream suppliers. Through most of the 1990s, the combined profits of Intel and Microsoft alone exceeded the total profit of the entire world PC assembly industry.
Only Dell made real money as an assembler.
Here is the detail that makes the story worth telling twice: Microsoft and Intel were not obviously the best technology available at the time. Zilog and Motorola built chips that many engineers considered superior.
Digital Research had an operating system widely seen as comparable. IBM did not lose the PC business by picking bad partners. IBM lost it by picking partners at the layer that was structurally destined to compound, the layer that set the standard interface everyone else had to build against, and then letting those partners own that interface completely.
The lesson was about knowing which room the value was going to end up in.
Same Model, Different Layer, Completely Different Outcome
If the framework is right, companies with the same commoditized models should end up in very different places because of what they built beneath or beside them, not because of the model they chose. They in fact do, and the contrast is striking.
Jasper AI reached a valuation of roughly 1.5 billion dollars as an AI writing assistant built on top of early GPT access. Its entire value proposition was a friendlier interface to a model anyone could call.
When ChatGPT launched in November 2022, that interface advantage evaporated overnight, because the model vendor itself shipped the same convenience directly to consumers for free.
Jasper cut its internal valuation, pivoted toward enterprise customers, and still watched revenue decline sharply over the following two years, because there was nothing underneath the interface that a competitor, or the model vendor, could not replicate by Tuesday.
Now look at Harvey, a legal AI platform built on the same category of underlying model access. Harvey trains custom agents on each law firm's own proprietary documents, inside that firm's security and compliance boundary, and firms are now running more than 25,000 such agents per company.
Once a firm has built that depth of custom workflow on Harvey's platform, switching to a competitor means rebuilding all of it from scratch. Harvey reached 300 million dollars in annual recurring revenue in May 2026, up from 195 million just five months earlier, and closed a 200 million dollar growth round at an 11 billion dollar valuation.
Or look at Cursor, the AI coding assistant that passed 2 billion dollars in annualized revenue by February 2026 and was bought by SpaceX last month.
Analysts are explicit that the valuation rests on workflow lock-in and developer trust, not on which underlying model Cursor calls, which is precisely why thin AI wrappers are dying while thick ones, the ones with genuine data or workflow depth, are trading at premium multiples well above ordinary SaaS.
All four companies Jasper, Harvey, Cursor, and the model vendors they all built on top of, had access to comparable commoditizing model capability at the time that mattered most. The difference in outcomes was whether each company built on top of it or failed to.
The Honest Complications, and Why They Matter More Than the Headline
Let me challenge my own assumptions and play a devil’s advocate.
The first complication to this model is the capital allocation story. It would be satisfying to say enterprises are recklessly overspending on the commoditizing application layer while starving the compounding infrastructure layer.
The actual 2025 spending data does not fully support that. Enterprise generative AI spend split nearly evenly, $90 billion to applications and $18 billion to infrastructure, and infrastructure already captures something like 45 percent of AI spending broadly defined.
The sharper, better evidenced claim comes from MIT's Project NANDA, which studied 300 public deployments and interviewed 150 enterprise leaders: 95 percent of enterprise generative AI pilots fail to deliver any measurable profit and loss return, and internal builds succeed at roughly 1/3rd the rate of purchasing from a specialized vendor or partner.
The real problem is not a dollar misallocation between layers. It is that most internal builds land one layer too high, competing on ground that commoditizes out from under them, without anyone first asking which layer the build actually sits on.
Your spend goes in the wrong layers.
The second complication is the data moat question, and here the honest answer is domain conditional rather than universal.
One camp argues that because foundation models train on public and increasingly synthetic data, most proprietary datasets are speed bumps rather than castle walls.
A competing camp points out, also correctly, that in specific regulated or high transaction volume domains, payments processors sitting on transaction data no model vendor can replicate, or legal and healthcare firms operating inside a trust and compliance perimeter, genuinely hold a moat no foundation model release will erase.
The honest synthesis is not "data moats are dead."
There you go, I said it.
It is that data moats are domain conditional, and knowing which domain you are in matters more than the general debate.
The third complication is switching cost. The clean version of the model layer is commoditizing; that implies switching vendors is trivial.
For an organization that has already sunk real cost into fine-tuning, retraining, and deep integration with one vendor's specific model behavior, it is not trivial at all, and no dataset currently quantifies exactly how not trivial.
Treat this as real friction, not a footnote.
Finally, a distinction that complicates the model versus infrastructure binary itself: Cursor's advantage is arguably distribution and developer trust as much as it is proprietary data.
Distribution, the depth of platform integration and adoption reach, may be a third kind of compounding asset that sits at neither the model layer nor the classic infrastructure layer. A framework that only looks for compounding value in data and hardware will miss it.
A framework that survives its own best objections is worth more than one that pretends they do not exist.
The Instrument: Scoring Your Own Stack
Here is where this stops being interesting and starts being usable Monday morning.
Take every AI capability your organization currently funds and score it, layer by layer, against three questions.
Vendor multiplicity: is this capability available at comparable quality from more than one credible provider today?
Price versus value trend: is the price of this capability falling faster than the value it delivers is growing?
Structural replication difficulty: does this capability get genuinely harder for a competitor to copy the longer you operate it, or does it stay exactly as copyable next year as it is today?
Rate each dimension one to five and add them up per layer.
A high vendor multiplicity score and a steep price versus value decline signal a commoditizing layer, worth renting cheaply and never overcommitting to.
A high structural replication difficulty score signals a compounding layer, worth the deliberate long term investment.
Run this scoring exercise as a standing gate on every new AI commitment this quarter, because the commoditization cycle is now running at roughly twelve months and getting shorter with every release.
The posture differs by organization. For enterprise companies, this is a capital allocation and vendor contract question that belongs with the CFO and the board, not just the CIO.
Model your multi-year vendor commitments differently depending on which side of the line they fall: favor short, flexible terms at the commoditizing model layer, and favor longer-term ownership or capacity commitments at the compounding layers, the way the hyperscalers already locked in GPU allocation and power contracts years ahead of need.
Ask your Chief Data and AI Officer a harder question than how many seats are activated. Ask which of your internal builds are being defended as proprietary that actually sit one layer too high, competing on the ground a vendor's next release will commoditize for free.
For mid-market, you will never outcapitalise a hyperscaler on the infrastructure layer, and that is not the fight to pick.
Run a two column audit instead. In one column, what you are buying, model access, generic copilots, off the shelf interfaces.
In the other, what you are building or capturing, proprietary workflow data, customer interaction history, documentation of institutional judgment that nobody else has. Stop overpaying for the first column.
A custom AI platform that is really a commoditizing model wrapper with a markup can do real damage to a mid-market P&L over a multi-year contract.
Redirect what you save into the second column, even in small increments, because a mid market firm that correctly identifies its own compounding layer early can build a defensible position before a larger competitor finishes forming its AI governance committee.
The Takeaway, and What Comes Next
Cheaper intelligence doesn't mean cheaper advantage. It means the advantage moved.
If you don't know where it moved to, you are still paying full price for something that is now free.
Part 1 told you value migrates toward the constraint. Part 2 tells you why the migration happens at different speeds across the stack, and gives you a line to draw through your own organization to find out which side you are standing on.
The line is not fixed as it runs inside layers as much as between them; it moved through the PC industry in the 1990s exactly the way it is moving through the Intelligence Supply Chain now, and it will keep moving as fast as the next model release and the next interconnection queue update allow.
Pick your three largest AI line items. For each one, ask honestly which side of the Commoditization Line it sits on, and whether your organization is treating it accordingly, renting the commoditizing ones cheaply and flexibly, and investing deliberately in the compounding ones before the queue for them gets to seven years.
Next week, in Part 3, we go beneath the Commoditization Line itself and into the Infrastructure Inversion, the mechanics of why physical infrastructure now scales ahead of the intelligence it powers, and what that inversion means for how you plan capital, not just capability.
Talk soon,
Sameer Khan










