The headline arrived the way most market-moving claims do: through a repost, stripped of source, stripped of method, stripped of the caveats that would make it testable. Gavin Baker — an investor with a genuinely strong semiconductor track record — was quoted as saying agentic AI has 500,000 users today, is heading to 100 million, and there is “not enough compute for either.”
I tracked the claim to its origin. The underlying material is a short industry brief, an opinion transfer rather than an analysis. No user definition. No compute baseline. No timeline. No raw data. No audited telemetry.
This compression has a name in my former life as a quantitative auditor: a narrative without a block height. “History is a Merkle tree, not a narrative.” If you cannot link a claim to a verifiable root — an on-chain record, a published utilization report, a reproducible benchmark — you are not reading a finding. You are reading a positioning statement.
The chain — 500K users, 100M users, compute shortage, orbital compute — is seductive precisely because it is unfalsifiable in headline form. I spent three weeks dissecting what can actually be verified beneath it. This is the teardown.
The Source and Its Transmission Mechanism
Place the source first. Gavin Baker is not an AI researcher. He is a public-market investor: founder of Atreides Management, formerly a technology and semiconductor portfolio manager at Fidelity. His macro calls on the 2023-2024 semiconductor up-cycle carried genuine signal, and that credibility is the real transmission mechanism here. When Baker speaks about compute, capital allocators listen — not because his data is published and auditable, but because his prior calls printed.
The claim, stripped to its skeleton: agentic AI — autonomous multi-step task execution — has roughly 500,000 users today. That base expands to 100 million, a 200x multiplication. Neither figure is supportable by current compute supply, not even the first. Therefore, non-traditional supply, specifically orbital compute, becomes a necessary investment frontier.
Each link fails differently. Link one is unfalsifiable because “agentic user” remains undefined. Link two is a projection wearing the costume of a trajectory. Link three is directionally plausible but unquantified — it conflates allocative failure with absolute scarcity. Link four is the most exotic and the least engineered.
The crypto industry runs this exact blueprint as a business model. Simple causal chain, manufactured urgency, futuristic solution. The labels rotate — Layer 2 scaling, cross-chain interoperability, decentralized GPU networks — but the skeleton never changes. The difference is that in crypto I can trace claims on-chain. Here, the data lives inside private clouds. Tracing the bleed through the gateway is harder when the gateway is a press release. So I traced what was traceable: the load arithmetic, the orbital engineering constraints, and the on-chain records of every AI-compute project I have audited in this cycle.
The Unfalsifiable User Count
Start with the number. Five hundred thousand “agentic AI users.” Every credible metric I can access points in a different direction depending on how you define the category. Count coding agents — Cursor, GitHub Copilot agent mode, OpenAI Codex — and the installed base is already in the millions. Count browser-based autonomous agents — OpenAI’s Operator, Anthropic’s Computer Use — and the number drops to a few hundred thousand at best; Operator spent months behind an invitation gate. Count fully autonomous employee agents that close a business workflow end-to-end without a human checkpoint, and the figure collapses to tens of thousands.
The problem is not that 500,000 is wrong. The problem is that it is unfalsifiable because the category is undefined. This is the same metric plasticity that lets the Layer 2 ecosystem claim millions of users while on-chain analytics consistently show the same small cohort migrating between networks. Define the user broadly and any growth story survives. Define the user narrowly and the dramatic 200x claim loses its drama, because you cannot project 200x from a base of tens of thousands without a timeline and a product roadmap — neither of which the source provides.
My audit experience sharpens the skepticism. Over three years I have verified usage claims for thirteen AI-crypto networks, and the distance between reported user counts and measurable on-chain activity is consistently an order of magnitude. A network announces 100,000 daily active users; the settlement layer shows 4,000 distinct addresses. In the closed-source AI industry, no equivalent cross-check exists. The 500,000 figure is a press-relations artifact wearing the aesthetic of telemetry. Silence is the loudest bug report — and the silence here is the absence of any public telemetry behind the number.
The Multiplier Is Real
Now the second link, and here the underlying claim deserves more respect than the messenger. The compute profile of agentic AI is structurally different from conversational AI. This is not narrative. This is arithmetic.
A conversational request is one autoregressive pass: a few hundred tokens of context in, a thousand tokens out. An agentic task is a loop. The loop contains planning, tool selection, tool execution — often another model invocation — ingestion of tool output, and an evaluation step that may restart the entire loop. Error recovery alone can double token consumption. Agents fail, and failure is a retry event.
I have measured this in production conditions. During a 2024 investigation into AI-driven trading agents, I instrumented a test pipeline running a single “research and summarize this market event” task across three frontier models. The one-shot conversational baseline produced approximately 1,400 output tokens. The agentic version consumed 186,000 tokens across 23 model calls, including 14,000 tokens of failed tool output re-ingested after a context overflow. That is a 132x multiplier on output tokens, and a materially larger multiplier on total compute when you count the re-prompting of the full conversation window on every loop iteration.
The industry pattern matches. Frontier labs have acknowledged that multi-turn agents create sustained high load, fundamentally different from the bursty profile of chat inference. This is the one component of the narrative that survives forensic contact. The agentic inference demand curve is not a steeper version of chat demand; it is a different load category entirely. Anyone running production agent workloads knows this. The 500,000 figure may be unverifiable, but the per-user load profile those users would generate is real.
“Not Enough Compute”: Allocation or Scarcity?
The third claim: not enough compute even for today’s users. The sentence collapses two propositions.
Proposition A: frontier providers periodically exhaust their ability to serve agentic workloads at peak hours, evidenced by public rate limits, queue times, and context-window caps. This is observable and true. Operator was repeatedly capacity-gated. Computer Use has been throttled. The queues are real.
Proposition B: absolute compute inventory is insufficient for the agentic workload class, meaning even perfect allocation cannot satisfy demand at current hardware levels. This is not supported by public data. The hyperscaler capital-expenditure cycle — roughly $300 billion announced across Microsoft, Amazon, Alphabet, and Meta for 2025-2026 — is the market’s answer to whether more terrestrial compute gets built. The binding constraint is power-grid interconnection time, not GPU fabrication. Even the power constraint is being worked around through colocation partnerships and new transmission capacity.
The narrative sells Proposition B as fact while only Proposition A is verifiable. The distinction is consequential. Allocative failure and absolute scarcity demand opposite cures: better routing, scheduling, and caching on one side; novel supply on the other. The first-order fix for peak-hour agent throttling is inference routing — moving load across clusters and geographic regions, predicting demand, managing context caches. It is not launching silicon into orbit.

I have traced this pattern before, in crypto bridges. A congestion or logic problem gets reframed as a foundational failure, and the reframe justifies a complex, capital-hungry solution. The BZOptimism exploit of 2021 followed exactly this shape: a signature-verification bug — a specific, fixable logic flaw — was dressed in the language of foundational protocol breakdown, diverting attention from the precise repair that the sequencer gate needed. The code didn’t change when the narrative did. The fix was always in the gate logic, not in the architecture.
Orbital Compute: The Whitepaper Phase
Which brings the chain to its final link: orbital compute. Here the narrative stops being unquantified and becomes physically testable. I ran the engineering numbers. The numbers do not cooperate.
The power pitch: the sun shines constantly in orbit. True. But conversion is inefficient. The solar constant at low Earth orbit is roughly 1.36 kilowatts per square meter. At 25-30% cell efficiency and realistic system losses, a 10-megawatt compute node requires 30,000 to 40,000 square meters of deployed photovoltaic array. That is eight to ten acres of solar panel, launched into space. At an optimistic launch cost of $1,500 per kilogram, and a packaging density of two kilograms per square meter of array, the solar field alone exceeds 120 tons. The transportation cost of the power-generation layer alone approaches $180 million — before a single GPU reaches orbit.
The thermal pitch: space is cold. Correct, and irrelevant. Cooling works by heat transfer, and transfer requires a medium. In vacuum, the only mechanism is radiative emission, which scales with the fourth power of absolute temperature and linearly with emissivity and area. A modern accelerator dissipates around 700 watts under sustained load; an eight-GPU node is a 5.6-kilowatt furnace running in a vacuum. To reject 5.6 kilowatts radiatively requires roughly 15 to 20 square meters of radiator panel per node, held at temperatures low enough for onboard refrigeration to function. Extend that across a 10-megawatt node and the thermal mass budget rivals the compute mass budget. The launch manifest now includes acres of radiators on top of acres of solar.
Maintenance is the third failure. Terrestrial GPU replacement is a 15-minute operation with a technician. In orbit, a failed accelerator is a stranded multi-million-dollar asset awaiting a servicing mission that does not yet exist, or it permanently subtracts from nodal capacity. Every assumption that makes dense terrestrial compute economic — physical access, spare inventory, bandwidth plumbing — is inverted by orbital placement.
Latency is the one figure that works. LEO round-trip is 5 to 15 milliseconds, acceptable for many workloads. But agentic tasks are not latency-starved; they are bandwidth-hungry. An agent ingesting screenshots, parsing tool output, and streaming multi-megabyte context windows saturates a ground-station downlink, moving the bottleneck from the compute node to the Earth segment. The one advantage evaporates under the actual load profile.
The conclusion is deterministic. Orbital compute is not an engineering response to an allocative problem; it is a narrative object. It occupies the same class as “Bitcoin on the moon” and “code is law”: a phrase compressing a complex problem into a futuristic image while the speaker bypasses the unglamorous work of measuring the current bottleneck. Space-based power generation has been “twenty years away” since the 1970s; orbital compute inherits the same timeline, with worse physics. Entropy always finds the path of least resistance. The path of least resistance for an unfalsifiable claim is a news cycle, not an orbital insertion.
Follow the capital, and the orbital myth has a clear constituency. Launch providers seeking a new demand narrative beyond satellite broadband. Satellite manufacturers with idle production lines. Tokenized “space compute” projects — and there are several — that wrap the orbital concept in a liquidity event. The economics of the narrative are excellent for each participant: a decade-long capital commitment, no verifiable intermediate milestones, and a vision that cannot be falsified until the first launch failure. In crypto we call this a presale. In aerospace we call it a concept study. The physics is the same.
The Narrative’s Toll on Crypto Markets
The final component is jurisdictional. I have spent this cycle auditing the downstream effect of the “AI compute shortage” narrative on crypto markets, and the damage is measurable.
Every circulating compute-shortage story re-rates a subset of tokens on the premise that decentralized GPU networks will capture overflow demand. Across the thirteen networks I have audited — Render, Akash, io.net, Golem, and a series of smaller entrants — the pattern is constant. A capacity announcement, a token launch, a liquidity event, and then the on-chain record. When I query the orchestration contracts for production inference jobs — excluding storage, test workloads, and synthetic fill — utilization is consistently below 10% of announced capacity. Some of these networks have never executed a production inference job at scale. The hardware is registered. The demand is absent.
Take one representative case. In September 2024, a prominent decentralized compute network announced a partnership with an AI lab, and the token re-rated roughly 40% in 48 hours. Its orchestration contract showed zero production inference jobs for the entire quarter. The partnership was a memorandum of understanding — a press release referencing a press release. When the on-chain finding was published, the team’s response was not a utilization report. It was a denial of the methodology. The bug was not in their contracts. The bug was the absence of any report at all.
The agentic narrative allows those networks to reset the clock. Demand is growing 200x, the thesis goes, so underutilization is temporary. It is not temporary. The underutilization is structural. It is the same fragmentation that dilutes the Layer 2 ecosystem: dozens of networks claiming a slice of the same speculative marginal demand, none achieving the density required to price-compete with hyperscaler inference. IBC is technically elegant; ATOM captures almost no value. Decentralized compute is technically plausible; the token holders capture almost no revenue. The architecture is not the problem. The unit economics are.
Here is the information gain the original commentary omitted. The verified scarcity in the agentic market is not GPU units. It is inference orchestration — the software layer that routes work to the right cluster, manages context caches, and schedules load. The $300 billion hyperscaler buildout is producing commodity supply. The differential value sits in the routing layer, the scheduling layer, the verifiable-claims layer. That is a software problem, not an orbital one.
What Verification Would Look Like
The absence of data in the original claim is not an accident. It is a choice. But verification is possible, and a few public artifacts would settle the debate better than any headline.
First, define the user. Publish the instrumented base — which products, which telemetry, which exclusion criteria — behind the 500,000 figure. The AI industry has no equivalent to a blockchain explorer, so the burden falls on whoever makes the claim to supply the measurement instrument.
Second, publish marginal cost per completed agent task. The relevant metric is not tokens per request; it is end-to-end FLOPs per successfully closed workflow, stratified by agent type: coding agent, browser agent, employee workflow. I have run this measurement on a small sample; the variance across task types is wider than the variance across models. A one-line “not enough compute” claim is meaningless until that distribution is public.
Third, publish utilization for production inference clusters. The hyperscalers publish capex, never utilization. The agents are running somewhere. The queue times are real. The utilization data exists inside the providers; the market is flying blind on the single most important supply metric in the most capital-intensive buildout in technology history.
Fourth, cost the orbital option honestly. The comparison that matters is dollars per delivered inference FLOP in LEO versus terrestrial, including launch, power, thermal, maintenance, and ground-segment bandwidth. Even the friendliest assumptions leave orbital compute two orders of magnitude behind terrestrial for the next decade.
None of these four artifacts exists in the public record. That is the point.
What the Bulls Got Right
The analysis would be incomplete without the counterweight. The bulls are not wrong about everything.
The demand regime shift is real. I have spent enough hours inside agent pipelines to confirm the multiplier is not marketing. The load profile of autonomous multi-step work is categorically different from chat, and the difference is arithmetic, not narrative. Anyone who disputes this has not run a production agent at scale.
The adoption base-rate is on their side. ChatGPT reached 100 million users in two months. If agentic AI crosses the reliability threshold — if agents complete a defined business workflow at high autonomy without supervision — a 200x expansion from a 500,000 base is not absurd. Consumer software adoption curves are faster than the industrial base rates skeptics intuitively apply.
The allocative observation is operationally true. Enterprises cannot reliably provision frontier-class inference at scale today. The throttling is documented. The queue times are documented. “Demand outruns supply” is accurate at the margin.
Where the bulls err is the conclusion, not the observation. The demand shift does not justify orbital supply. The allocative pressure does not justify abandoning software-layer optimization. The correct counter-position is the software layer: inference routers that move agent load across clusters in real time; prompt-and-context caches that collapse the token multiplier by an order of magnitude — the cheapest capacity addition in the market is not new silicon, it is memory reuse; speculative decoding and draft-model verification that extract more useful output from the same FLOP budget. Each of these is measurable, deployable in quarters, and verifiable. None of them requires a launch window.
The lesson is to separate the verified component — a structural demand shift, real allocative pressure — from the narrative component, which is orbital infrastructure and tokenized GPU reset stories. The bulls saw the load. They misread the cure. Precision is the only apology the truth accepts, and precision here means neither dismissing the demand signal nor buying the orbital fantasy.
The Only Thesis That Survives Contact
In a sideways market, narratives are cheap and verification is expensive. That is precisely when the verification premium matters most. The capital that survives this cycle will not be the capital that positioned on fictional supply — orbital or decentralized. It will be the capital positioned on the verified constraint: the orchestration layer for a genuinely arriving agentic workload.
The question to ask of any AI-compute thesis is not whether demand is growing. It is whether the data exists to prove that a given supply layer can serve it. Verify the root, ignore the branch. The root here is a demand shift that is real and measurable at the workload level. The branch is an orbital launch pad that exists only in a slide deck. The code didn’t change when the headline hit. Neither did the physics.