Intelligence Economy

Intelligence Economy

The Constraint Multiple: Stop Chasing GPUs or Compute

Every AI infrastructure review I've sat in tracked chip procurement.

Sameer Khan's avatar
Sameer Khan
Jul 25, 2026
∙ Paid

Hey Transformation Leader,

Here is a question I have been thinking about.

If every constraint on your AI project disappeared tomorrow except one, which one would still delay it?

Most people would probably say GPUs.

I would have said the same thing a year ago. But something changed my thinking.

We were in a meeting about expanding our AI infrastructure when someone raised a concern about our significant dependency on one of the major AI providers behind the copilots we use.

It was a valid concern. But it also made me think about dependency differently.

We experience a copilot as software. We open it, type a prompt, and expect an answer.

Underneath that simple experience is a lot of physical infrastructure.

  • Data centers.

  • GPUs.

  • Power.

  • Transmission.

  • Cooling.

  • Transformers.

  • Permits.

And people who know how to build and operate all of it.

Our dependency was not limited to one AI provider.

We also depended on whether that provider could keep expanding the physical infrastructure underneath its software.

That is a very different problem.

Table of Content

  1. The bottleneck has moved below the compute layer

  2. What’s Actually Happening

  3. Why “The Stack” Isn’t Ordered, and Why That Matters

  4. The Constraint Stack Audit (Download the Constraint Audit Calculator)

  5. What This Means Depending on Where You Sit

  6. Where This Leaves You

The bottleneck has moved below the compute layer

One lesson the energy industry has taught me is that physical reality always wins.

Technology teams often assume infrastructure will scale as demand grows because that is how the cloud era worked.

Need more storage?

Buy more.

Need more compute?

Increase your cloud capacity.

The physical complexity stayed mostly hidden from the customer. Energy does not work that way.

Power, transmission, permitting, and supply chains move on multi-year timelines.

No amount of software can compress them.

That is why I spend less time asking what the best technology is and more time asking what the real constraint is.

Because the constraint is what determines how fast anything can scale.

I have been in attendance at these conversations. The conversation always narrowed to GPU availability and compute contracts. Nobody was talking about interconnection queues.

It’s also increasingly wrong, and the market is now telling on itself.

H100 cloud rental prices have fallen 64 to 75 percent from their 2024 peak. Chip supply, in the narrow sense of raw GPU availability, has genuinely eased. If chips were still the binding constraint, that easing should be translating directly into faster AI deployment timelines.

It isn’t.

Source: SemiAnalysis H100 Rental Price Index, IntuitionLabs and Thunder Compute 2026 GPU pricing trackers.

Between 30 and 50 percent of scheduled 2026 U.S. data center capacity is now expected to be delayed or canceled, up sharply from a historical baseline where project cancellations went from 6 in 2024 to 25 in 2025. Something is gating delivery that has nothing to do with chip availability.

That something is the physical and regulatory layer underneath the compute layer, and it is currently the least-modeled part of almost every AI capital plan I’ve seen.

What’s Actually Happening

For the last two years, almost every AI infrastructure conversation has focused on chips.

That focus made sense. GPU supply was tight. Prices were high. Companies were competing for allocations.

But access to some types of compute has started to improve.

The other timelines have not. A data center still needs access to a large and reliable block of power.

It may need a new grid connection. It may need a substation, transformers, switchgear, cooling equipment, permits, and utility upgrades. Those things are not delivered at software speed.

Think about it.

Interconnection queues, the process by which a new facility gets connected to the grid, are now running 4 to 7 years in major hubs like Northern Virginia, Phoenix, and Dallas.

PJM data shows projects spending roughly 3 years just reaching an interconnection service agreement, then another 4 years waiting to actually go live.

Even after approval, you’re not done. High-power transformers now carry 3 to 5 year delivery schedules. Switchgear is sold out through 2028. Substation transformer lead times have climbed from about 140 weeks in 2023 to over 160 weeks in 2026.

Compare that to chip lead times.

Data-center GPU lead times currently run 36 to 52 weeks. Do the math, and you get something close to a five-to-one gap between the constraint everyone is managing and the constraint that actually governs when the facility goes live.

Five to one. Most project plans don’t account for that ratio.

CBRE’s 2026 Data Center Outlook makes clear that demand in 2026 is no longer gated by capital, connectivity, or land. It’s gated by the ability to deliver large blocks of power on an aggressive timeline.

Northern Virginia’s data center vacancy has fallen to 0.3 percent. Atlanta sits at 1 percent. There is functionally no slack left in the markets everyone wants to build in.

Then there’s water, which is the constraint almost nobody is pricing into their models yet.

Hyperscale campuses can consume 1 to 5 million gallons a day. U.S. data centers used an estimated 17 to 20 billion gallons directly in 2024 and 2025, and that figure is projected to roughly quadruple by 2030.

More than 20 states are now weighing data center water bans or moratoriums. In the first quarter of 2026 alone, more than $130 billion in projects were delayed or abandoned over water-related issues, more than all of 2025 combined.

Source: Photo by Erik McGregor/LightRocket via Getty Images

When I think about this from inside a Fortune 100 energy company, the physical constraint problem isn’t new.

We’ve been dealing with multi-year infrastructure lead times in oil and gas for decades. What surprises me is watching the tech industry rediscover what heavy industry already knew if you don’t control the physical inputs, you don’t control the timeline.

But if the site cannot be energized, you do not have an operating AI facility.

You have expensive equipment waiting for electricity.

That is why the difference between the GPU timeline and the power timeline matters.

Your procurement team may be managing to one clock while the project is actually governed by another.

I’ll be straight with you about the counter-signal, because ignoring it would make this argument weaker, not stronger. Sam Altman said it plainly this year: “It goes back and forth. Right now, again, it’s chips.”

He’s not wrong. HBM memory lead times for data-center GPUs still run 36 to 52 weeks, and TSMC’s CoWoS advanced packaging capacity remains sold out into 2026, with Nvidia alone holding over 60 percent of that output.

The honest picture is that chip and packaging constraints govern how fast you can procure compute.

However, power, water, and permitting constraints govern how fast you can physically build the place that houses it.

Those are two different bottlenecks gating two different things, and most boards are still treating them as one undifferentiated “AI scarcity” problem. That conflation is exactly why capex gets approved against timelines that were never achievable.

Why “The Stack” Isn’t Ordered, and Why That Matters

I am not saying chips are no longer a constraint.

They are.

High-bandwidth memory, advanced packaging, and demand for the latest GPUs can still slow a project down. But chips govern how quickly you can acquire compute.

Power, water, construction, and permitting govern how quickly you can operate it at a particular site. Those are not the same question.

And the constraint can change over time. Today it may be GPUs. Tomorrow it may be power.

After that, it may be cooling, labor, or a permit.

So I do not think this is a fixed stack that we solve from top to bottom.

It is a network.

When I started outlining this piece, I wanted to give you a seven-layer diagram: power, interconnection, water, chips, labor, permitting, capital, stacked in a fixed sequence, solve them top to bottom. I used to think that kind of clean sequencing was the most useful thing I could hand an executive.

I’ve revised that view, and I want to tell you why, because the reasoning matters more than the diagram would have.

A fixed sequence implies that once you clear layer one, layer two is next in line, waiting its turn.

The Southern Nevada example alone disproves that. Water and power aren’t queued; they’re coupled. Fixing one changes the other. Altman’s own comment tells you the same thing from a different angle: which constraint binds hardest isn’t fixed over time, it oscillates depending on where the market is in a given cycle.

A rigid stack order is the single easiest thing for a sharp reader, or a McKinsey partner, to dismantle, because it will be visibly wrong the moment they bring their own project to it.

So here’s the more accurate, more useful model.

These constraints aren’t a stack you climb but a network or process you manage.

Power, interconnection, water, chip and packaging capacity, labor, permitting, and capital all interact, and resolving one can tighten another.

Although the visual metaphor of a stack is still useful for communicating the idea of “underneath the layer you can see, there are more layers,” the moment you treat it as a strict sequence, you’ve built something fragile.

Treat it as a system of tradeoffs you manage continuously, and it holds up against scrutiny from people who actually run these projects.

This is also, not coincidentally, a lesson heavy industry already learned.

For example, Aluminum smelters were sited next to hydropower before roads existed to reach them, because power was cheaper to build around than to transport.

Nobody in that era thought of “power access” as one item on a sequential checklist behind capital and permitting. It was the organizing constraint the entire site decision was built around, and everything else, roads, labor, logistics, got engineered to follow it.

What looks like a 2026 AI infrastructure innovation, on-site gas plants, fuel cells, small modular reactors, is really a rediscovery of a century-old industrial siting discipline.

Oracle’s partnership with Volta Grid to deploy a 2.3 gigawatt modular gas fleet.

Crusoe Energy building dedicated on-site gas plants.

Equinix’s agreement with Bloom Energy for over 100 megawatts of fuel cells across 19-plus sites.

Oracle’s own plans to power a 1 gigawatt facility with small modular reactors.

Source: VoltaGrid

These are not ad-hoc decisions. It’s the predictable response of sophisticated operators who’ve concluded that owning generation beats leasing grid access.

Watching who owns power generation versus who’s still waiting in an interconnection queue is now a better predictor of who wins in AI infrastructure than watching who has the most GPUs.

This is also where this week’s argument connects directly back to the framework underneath everything I wrote in the Intelligence Supply Chain post, and its core claim that intelligence is manufactured, and value migrates to the bottom of the stack.

Now you can see precisely why that migration happens. Value moves downward because the bottom of the stack is where the multi-year, non-substitutable lead times live.

Whoever controls that longer, harder-to-substitute layer, or can work through the lead times to secure it, captures a disproportionate share of the value the whole system creates.

I wanted a simple way to make this problem visible to executives.

So I came up with something I call the Constraint Multiple.

The calculation is simple:

Constraint Multiple = longest critical-path lead time ÷ lead time used in the current project plan

Let’s say your team is planning around a one-year GPU procurement timeline. But the realistic power-delivery timeline for the site is five years. Your Constraint Multiple is five.

That number does not tell you everything, but it tells you something important.

You are managing the project against the wrong timeline.

In this market, that number is currently running somewhere around 5.

The Constraint Stack Audit

Before approving the next AI infrastructure investment, look at every major constraint.

  • Power.

  • Interconnection.

  • Water.

  • Chips and packaging.

  • Labor.

  • Permitting.

  • Capital.

Score each one in three areas.

  1. Lead time: How long will it realistically take to resolve?

  2. Severity: Will it stop the project, reduce capacity, or just increase cost?

  3. Mitigability: Can you address it through investment, contracting, redesign, or a different location?

Then rank the constraints by their impact on the go-live date. The output should show you which constraint is really controlling the project. It should also show you which constraint is likely to appear next. That matters because a completed checklist can create false confidence.

Seven green status boxes do not mean the facility will open on time. The real question is whether the slowest critical dependency supports the date in the presentation.

Here is a Constraint Multiple calculator I created.

Below you will find a link to download this calculator.

User's avatar

Continue reading this post for free, courtesy of Sameer Khan.

Or purchase a paid subscription.
© 2026 Sameer Khan · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture