Intelligence Economy

Intelligence Economy

NVIDIA Is Telling Us What the Next AI Bottleneck Will Be

What is the shift from cost per token to a new metric and what it means for enterprise AI.

Sameer Khan's avatar
Sameer Khan
Aug 08, 2026
∙ Paid

Hey Transformation Leader,

Yesterday, I was reflecting on my AI home setup and the output I get. My home setup is unique, as I have a Mac mini that is always on as an AI server. Then, I built a Mac app that uses ChatGPT's voice-to-voice API to talk to Hermes, which runs everything, including all my LLMs (API + local) and my laptop.

I built this setup because I wanted to be hands free like Jarvis with always on voice based AI management so I don’t have to be on my computer. I can just say HeyChatGPT, and it wakes up.

You may ask why I don’t use the ChatGPT voice app on my phone?

The answer is cost. If you keep ChatGPT voice to voice on, you are incurring costs. Additionally, it does automatically switch off if you don’t use it after a few minutes to prevent token usage, so it’s not always on like Alexa or Google Home. That’s why I built a custom setup.

But that’s not the point for my post today. The point I am making is that there is always something new happening with AI, compute, and energy every week, and it’s super intriguing.

Because of this, I find myself constantly absorbing it all and building something new at work or at home. But over the past few days, one idea kept sticking with me. I couldn’t really move past it.

Performance per watt.

Why?

Well, first NVIDIA used to talk in terms of teraflops, then they changed their lingo to GPU clusters. Now it’s token factory revenue and performance per watt.

The first number that really caught my attention was the GB300 NVL72. NVIDIA claims it can generate 25x more tokens per watt compared to Hopper.

Performance chart with tokens per second per megawatt on the y axis and years on the axis showing Kepler at the bottom left beginning with less than 1 tok/sec/MW in 2012 going to the top right with Rubin at 700 K tok/sec/MW in 2026.

Source: https://developer.nvidia.com/

Initially, I thought it may just be another hardware stat that I can ignore. We don’t have to obsess or remember every config in the market.

But the more I looked, I started seeing the same pattern.

Qualcomm is designing its Dragonfly system around tokens per watt instead of just raw performance. Meta is also shifting inference workloads to Google TPUs and reporting 4.7x better performance per dollar and 67% lower power usage for inference.

That’s when I started asking myself why this specific metric is suddenly becoming so important. What I discovered is that these companies are not chasing a smarter model as average companies do.

In fact, they are asking a completely different question.

How much output can we get from the same amount of electricity?

I think that’s where the AI infrastructure conversation is starting to shift.

Right now, most organizations still think about AI in terms of cost.

  • What’s the API bill?

  • Which model is cheaper?

  • Which team is burning the most tokens?

  • How many people are actually using the tools?

Those questions are valid because your invoice reflects it. What’s not obvious is the energy usage/consumption behind all of this.

However, hyperscalers who are building the infrastructure some of them I called about are starting to think differently. They are thinking in terms of tokens per watt and performance per watt.

Simply put: how much intelligence can you produce without scaling more electricity/power at the same pace?

That got me thinking about the impact on the enterprise and mid-markets.

We spent a lot of time digging into the rat hole of AI cost as compute or tokens instead of asking how much useful output we actually get from the electricity we burn.

If power is becoming the limiting factor or a blocker to AI infrastructure, then the denominator matters more than you think.

The race is no longer about compute; it’s about what you can power.

So the real infrastructure question is slowly becoming: how much intelligence can you extract from every watt before electricity impedes what gets built and what doesn’t?

So let’s dive into this today.

Table of Contents

  1. Why Performance per Watt Caught My Attention

  2. The Metric Most Organizations Cannot Answer

  3. Why Power Is Becoming the Constraint

  4. The Megawatt Denominator

  5. Who Actually Owns This?

  6. What Happens When Power Shows Up in the Price

  7. What I Am Watching Next

The Metric Most Organizations Cannot Answer

Your finance team can probably tell you exactly how much you spent on AI last month, including your API bill, cloud spend, software licenses, and probably which teams are using the most tokens.

But I don’t think they can answer the question we raised above:

How much business value did we create from the electricity behind all the AI the organization uses?

I am not saying your finance team is not capable.

What I am saying is we have been managing AI through dollars because that is what we can actually see, and it’s physical/digital. The invoice, model comparison, and the teams that are consuming it, i.e., usage and adoption.

What’s not super clear is the electricity behind it beacuse it’s hidden somewhere inside your cloud provider, data center, or infrastructure cost.

So most companies either don’t worry about it or don’t look at it. But once you notice hyperscalers focusing on performance per watt, it made me wonder what similar metric we should focus on within most organizations.

Because tokens per watt makes sense if you are NVIDIA, Qualcomm, or Google. I don’t think that metric means much beyond these super majors.

The enterprise cares about what intelligence is actually producing. Did it increase revenue, reduce costs, save time, or improve something that matters to the bottom-line margin?

Keep reading with a 7-day free trial

Subscribe to Intelligence Economy to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 Sameer Khan · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture