Monitoring Protocol v1.2.1

The physical substrate: hardware, data centers, electricity, and capital

Frozen baselineEvidence cutoff July 24, 2026

The physical substrate: hardware, data centers, electricity, and capital

AI is often described as weightless software, yet its economic frontier is determined by accelerators, host processors, high-bandwidth memory, advanced packaging, interconnect, server assembly, liquid cooling, semiconductor fabs, substations, turbines, transmission lines, and financing contracts.

Electricity is the most important correction to purely digital accounts of artificial intelligence. Models can be copied, software can be distributed globally, and demand can appear instantaneously. Power systems cannot expand at the same speed. A hyperscale data center may be designed and constructed in a few years, while major generation and transmission projects can require longer planning, interconnection, permitting, equipment, and construction cycles. The resulting mismatch makes power a real constraint - but not a simple ceiling on growth.

The scale of the load

The U.S. Department of Energy estimated that data centers consumed approximately 4.4 percent of U.S. electricity in 2023 and could consume 6.7 to 12 percent by 2028, with annual demand rising from about 176 terawatt-hours to 325-580 terawatt-hours (DOE 2024). The U.S. Energy Information Administration's AEO2026 likewise identifies data centers as a major driver of electricity-demand growth (EIA 2026). The International Energy Agency estimated global data-center electricity use at roughly 415 terawatt-hours in 2024 and projected about 945 terawatt-hours by 2030 in its base case (IEA 2025a). The ranges are wide because model efficiency, utilization, hardware, demand, and project completion remain uncertain.

The Federal Energy Regulatory Commission made the constraint institutionally visible in June 2026 by directing all six regional transmission organizations and independent system operators to justify or reform their treatment of large-load interconnection. The orders focused on speed, cost allocation, reliability, speculative projects, and the need for financial commitments that protect existing customers (FERC 2026a; FERC 2026b). The policy problem is therefore not merely "build more power." It is "connect credible loads quickly without shifting unearned costs or reliability risk to everyone else."

Constraint layer What becomes scarce Typical lead-time problem Economic response
Site Land, water, fiber, permits, tax agreements, proximity to power Local opposition and infrastructure coordination Move workloads, pre-negotiate sites, co-locate with generation
Grid Interconnection capacity, substations, transformers, transmission, firm service Queues and equipment delivery exceed data-center build cycles Flexible load, behind-the-meter generation, long-term power contracts
Generation Reliable energy and capacity with acceptable cost and emissions New plants and fuel infrastructure take years Portfolio of gas, nuclear, renewables, storage, imports, and demand response
Compute Advanced accelerators, networking, memory, packaging, cooling Semiconductor and equipment capacity is concentrated Multi-vendor hardware, co-design, model efficiency, inventory commitments
Capital Long-duration financing for uncertain utilization and revenue Assets arrive before demand is proven Project finance, customer prepayments, vendor financing, equity, guarantees

Table 4.1. "Electricity" is a stack of interdependent physical and financial constraints.

The hardware substrate: chips are systems, not interchangeable units

An AI data center is not a warehouse of interchangeable processors. It is a tightly coupled production system. Accelerators must be paired with host CPUs, high-bandwidth memory, advanced packaging, scale-up links inside the rack or pod, scale-out networking across racks, storage, compiler and runtime software, power conversion, cooling, and a facility designed for the resulting density. A shortage or incompatibility in any layer can strand capital in the others. The economically relevant unit is commissioned, software-usable, networked compute delivered within a power, thermal, latency, and reliability envelope - not chips ordered or nameplate FLOPS.

This distinction changes the thesis. Model capability can diffuse while hardware rents remain concentrated. Open weights can improve model choice but still require a proprietary compiler, scarce HBM, a particular interconnect, or a liquid-cooled rack. Conversely, a weaker individual accelerator can remain strategically important if a domestic ecosystem can manufacture, network, operate, and improve it at sufficient scale. The protocol therefore treats hardware as a stack of measurable control points rather than a single chip-performance league table.

Hardware layer Economic function Primary concentration or failure mode Protocol measure
Accelerator and host CPU Execute tensor, reasoning, orchestration, retrieval, and tool workloads. Allocation, architecture cadence, precision support, and dependence on one platform. Installed and productive share by generation; matched goodput per watt and per dollar.
HBM and advanced packaging Keep model state and data close to compute; join large dies and memory at usable bandwidth. HBM supply, packaging yield, CoWoS-class capacity, substrate and thermal limits. Qualified capacity, lead time, yield, memory bandwidth used, and packaging geography.
Scale-up interconnect Makes many accelerators behave as one logical machine inside a rack or pod. Proprietary fabrics, switch availability, collective-communication efficiency, fault domains. Effective bandwidth, communication overhead, failure recovery, and portability.
Scale-out network Connects racks and sites for distributed training and inference. Optics, switches, congestion control, topology, and network power. Application goodput, tail latency, packet loss, optical power, and time to deploy.
Software and toolchain Compilers, kernels, libraries, schedulers, security, observability, and model optimization. CUDA, Neuron, CANN, or other ecosystem lock-in; skills and code portability. Porting time and cost, performance retained after migration, developer availability, incident rate.
Rack, power, and cooling Turns components into an operable high-density system. Liquid-cooling readiness, busbars, transformers, power quality, serviceability, and spare parts. Commissioned rack count, rack-ready MW, PUE, cooling availability, maintenance downtime.
Fabrication and equipment Convert designs into leading-edge logic, memory, packaging, and tested systems. Foundry and tool concentration, process yield, earthquakes, water, electricity, and geopolitics. Production-qualified wafers and packages by geography; recovery time; supplier concentration.
Facility and grid Provide land, power, water, fiber, permits, and continuity. Interconnection queues, local opposition, firm power, construction, and financing mismatch. Status ladder from announced to productive MW; curtailment, utilization, and contract exposure.

Table 4.2. AI hardware is a coupled production stack; the bottleneck can migrate between silicon, software, packaging, networking, cooling, and the site.

Concentration and security across the development stack

A May 2026 Cloud Security Alliance research note reports approximately 92 percent NVIDIA share of the discrete-GPU market; HBM supply of roughly 50 percent SK hynix, 40 percent Samsung and 10 percent Micron, with 2025-2026 capacity committed; a roughly 30 percent HBM price increase in late 2025; and more than 70 percent of TSMC CoWoS-L capacity reportedly secured by NVIDIA. These figures are useful directional estimates, not audited market statistics: the note itself relies partly on secondary sources. The protocol therefore records the source, date, denominator and uncertainty rather than treating each percentage as an immutable structural fact (Cloud Security Alliance 2026).

The same note connects economic concentration to supply-chain security. It cites the PyTorch nightly dependency attack, the May 2024 Hugging Face Spaces breach, more than 352,000 suspicious issues identified across approximately 51,700 models, and a reported 6.5-fold year-over-year rise in malicious model uploads through 2024. Concentration can therefore create a correlated security exposure even when no physical shortage exists. H1B and the dashboard track common software, model-distribution and hardware-platform incidents separately from firm-specific drawdowns (Cloud Security Alliance 2026).

Layer Reported concentration or incident Protocol interpretation Required measure
Discrete GPUs CSA reports NVIDIA at about 92%. Potential platform and security common mode; denominator is broad discrete GPUs, not only frontier accelerators. Productive workload share, substitutability, porting loss, exposure-adjusted incidents.
HBM CSA reports SK hynix ~50%, Samsung ~40%, Micron ~10%; 2025-26 supply committed. A three-supplier complement can bind otherwise available accelerator capacity. Qualified supply, allocation, lead time, price, geography and recovery.
Advanced packaging CSA reports NVIDIA >70% of TSMC CoWoS-L capacity. Reservation power may constrain competitors while concentrating Taiwan exposure. Customer allocation, package yield, alternative processes and time to qualify.
Framework/distribution security PyTorch dependency attack; Hugging Face breach; malicious model-upload growth. Ubiquity can amplify one compromise across many organizations. Common dependency inventory, provenance, patch time, affected workload share.

Table 4.2A. Reported concentration is a starting point for exposure measurement, not a composite risk score.

NVIDIA: a full-stack AI-factory architecture

NVIDIA is best understood as both a system platform and a material AI-system issuer. Its fiscal 2026 filing describes data-center systems co-designed across GPUs, CPUs, NVLink switches, DPUs, NICs, scale-out networking, CUDA, libraries, models, and management software. The GB300 NVL72 illustrates the architecture: a liquid-cooled rack integrates 72 Blackwell Ultra GPUs and 36 Grace CPUs with NVLink and high-bandwidth memory. Vendor performance claims require matched workloads and facility-level power measurement. Separately, NVIDIA's 7.93528 percent SPY weight on July 23, 2026 and its large-customer concentration place it inside the same issuer and transmission tests applied to SpaceX and future public listings (NVIDIA 2026a; NVIDIA 2026b; State Street Global Advisors 2026a).

The platform can shorten the path from delivered components to productive clusters and reduce integration uncertainty. The same integration creates switching cost across CUDA, collective communication, networking, security, operators, facility design and model optimization. That power is not one-way: NVIDIA reported direct-customer concentration of 22 percent and 14 percent, plus estimated concentration among indirect customers. Large hyperscalers can use purchasing scale, custom silicon and enterprise distribution to bargain with or substitute for the platform even while remaining major buyers (NVIDIA 2026a).

NVIDIA is fabless and depends on external foundries, HBM, advanced packaging, contract manufacturers, power and customer capex. Its control point is therefore powerful but contingent. A common TSMC, HBM, packaging or platform-security shock can affect NVIDIA and multiple customers simultaneously; a customer-specific demand shock can flow in the opposite direction. H1A tests the issuer channel, while H1B tests the shared substrate.

A common seven-dimension cross-firm lens

JPMorgan Chase's country framework evaluates policy, hardware, models and software, energy and resources, finance, socioeconomics, and military and security, while warning that the dimensions should not be reduced to one metric. Version 1.2 adapts those same dimensions to the four focal firms. The entries are structured descriptions, not scores; no row is weighted and no total is calculated (JPMorgan Chase Center for Geopolitics 2026).

Table 4.2B. The same seven dimensions are applied to each focal firm. The table has no totals and is not a composite systemic score.

Dimension SpaceX OpenAI Anthropic NVIDIA
Policy Launch, spectrum, defense, export and merger approvals. AI rules, procurement, safety, listing and infrastructure approvals. AI rules, safety, procurement, investment and listing approvals. Export controls, industrial policy, competition and supply-chain rules.
Hardware Owned/hosted clusters plus launch, satellite and network assets; depends heavily on NVIDIA and external fabs. Large NVIDIA and cloud commitments; capacity still partner-mediated and partly announced. Diversified Trainium, TPU and SpaceX-hosted NVIDIA access; orchestration complexity rises. Accelerator, networking, software and rack platform; fabless dependence on TSMC, HBM and assembly.
Models/software Grok plus prospective Cursor application distribution. Frontier models, consumer distribution, agents and enterprise platform. Frontier models, coding/enterprise distribution and safety research. CUDA, libraries, networking, systems software and model tooling.
Energy/resources Large cluster loads, sites, cooling and connectivity; disclosed capacity contracts create utilization obligations. Multi-GW development ambitions and supplier-specific power exposure. Compute portfolio spans several providers and power systems. Demand depends on customers commissioning power-dense sites; supply depends on materials and manufacturing ecosystems.
Finance Public equity, $25B bond, stock acquisition currency and contracted compute revenue. Large private capital, strategic-investor links, prospective listing and multi-GW commitments. Large private financing, cloud partnerships, prospective listing and contracted compute. Large public index weight, high cash generation, strategic investments and concentrated customers.
Socioeconomics Connectivity, launch, defense, regional construction and developer-workflow exposure. Consumer and enterprise adoption, labor substitution, software and service reorganization. Enterprise and developer productivity, safety and labor effects. Supplier employment, customer capex, market wealth, technology diffusion and regional build-outs.
Military/security Launch, satellite, defense, communications, cyber and supply-chain continuity. Model misuse, government deployments, cyber and information risks. Safety, cyber capability, government use and concentration risk. Export-controlled compute, platform vulnerabilities, defense relevance and common-mode supply-chain risk.

Huawei: a sovereign full-stack alternative

Huawei is not simply a lower-cost substitute for one NVIDIA accelerator. It is constructing a parallel stack around Ascend NPUs, Kunpeng CPUs, the CANN and MindIE software environment, UnifiedBus interconnect, CloudMatrix infrastructure, cloud services, and liquid-cooled SuperPoDs. Huawei states that its Atlas 900 A3 can connect up to 384 Ascend NPUs as one logical system and that more than 300 systems had been deployed to more than 20 customers. Those deployment and performance figures are company disclosures, not neutral benchmarks (Huawei 2025a; Huawei 2026).

Huawei has explicitly described cluster scale and interconnect as a response to restricted access to advanced process nodes. This is strategically important. Export controls can raise unit cost and delay leading-edge production, while also stimulating domestic substitution in chips, software, networking, memory, and operating practice. A platform that is weaker per device can still be viable if it is available, supported, and deployable at scale. The offset is physical: using more devices to reach a target can increase floor area, power, cooling, network complexity, failure domains, and maintenance burden. The relevant comparison is verified task goodput per megawatt, per dollar, and per unit of operator effort - not raw chip count (Huawei 2025a; BIS 2024).

For SpaceXAI, OpenAI, and Anthropic, Huawei is not a currently disclosed direct supply path. Its importance is competitive and geopolitical. Ascend can reduce Chinese dependence on U.S.-controlled accelerators, create a second software ecosystem with its own switching costs, pressure global pricing, and alter the effectiveness of export controls. The thesis should therefore track Huawei as a sovereign industrial system, not as a product benchmark inserted into a U.S. procurement table.

Custom silicon and control-point migration

Google TPUs, AWS Trainium, AMD accelerators, Meta MTIA and Huawei Ascend can reduce dependence on one accelerator vendor. They do not necessarily decentralize the stack. A customer that moves from NVIDIA hardware into Trainium or TPU may exchange accelerator dependence for cloud, compiler, managed-service and contract dependence. The relevant question is not whether one vendor's share falls, but where portability, pricing power, security exposure and operating authority move.

The durable proposition is therefore control-point migration rather than a permanent NVIDIA monopoly. Bargaining power can move among accelerators, interconnects, compilers, clouds, foundries, HBM, packaging, power, application distribution and finance. The dashboard tracks layer-specific shares and switching costs so substitution at one layer cannot be misreported as full-stack decentralization.

TSMC and Taiwan: concentrated production under active geographic diversification

TSMC is a different kind of control point. It does not compete with NVIDIA or Huawei as a model platform; it turns designs into high-yield wafers and advanced packages. In 2025, TSMC shipped 15.0 million 12-inch-equivalent wafers, and technologies at 7 nanometers and below accounted for 74 percent of wafer revenue. The company continued preparing multiple phases of 2-nanometer fabs in Hsinchu and Kaohsiung and expanding leading-edge and advanced-packaging capacity in Taiwan (TSMC 2026a).

The Arizona build-out is meaningful diversification rather than immediate duplication of the Taiwan cluster. TSMC reports that its first Arizona fab entered high-volume N4 production in the fourth quarter of 2024, a second fab is moving toward production in the second half of 2027, and a third began construction in 2025. On July 16, 2026, the City of Phoenix announced an additional $100 billion plan, bringing total announced Arizona investment to $265 billion and describing ten fabs, two advanced-packaging facilities and an R&D center. The $265 billion figure is announced capital and facilities, not commissioned capacity; timing, yield, supplier density and production qualification remain separate measures (TSMC 2026a; City of Phoenix 2026).

Taiwan concentration extends beyond wafers, but it is no longer analytically sound to describe the location as a static, singular chokepoint. NVIDIA reported a dense Taiwan systems ecosystem participating in the Vera Rubin ramp; TSMC identifies earthquakes, water and electricity as operational risks; and leading-edge packaging and engineering cadence remain concentrated. At the same time, Arizona and other geographic investments are active diversification. The protocol therefore tracks announced, built, qualified, volume-producing and recoverable capacity separately by geography rather than assuming either complete dependence or completed duplication (NVIDIA 2026c; TSMC 2025b; City of Phoenix 2026).

The classification remains disaggregated. TSMC/Taiwan concentration is operationally and strategically critical because substitution and qualification take time. It becomes macro-financially consequential if delayed hardware cuts capital expenditure, revenue, utility load, supplier cash flow, or regional investment. It is financially systemic only if those effects materially impair credit, funding, payments, clearing, or major financial intermediaries. Geographic importance does not waive the protocol's channel-specific qualification tests.

Dimension NVIDIA Huawei TSMC / Taiwan
Primary role Global full-stack AI-compute platform and ecosystem. Sovereign full-stack alternative centered on Ascend. Leading foundry, advanced-packaging base, and supplier cluster.
Core control points GPU/CPU, HBM integration, NVLink, networking, CUDA, rack designs, management. NPU/CPU, UnifiedBus, CANN/MindIE, CloudMatrix, SuperPoD and cloud deployment. Process technology, yield, CoWoS-class packaging, qualification, volume ramp, supplier density.
Strategic advantage Installed software base, system codesign, time to productive deployment. Domestic availability, stack control, substitution under technology restrictions. Scale, yield, process cadence, customer diversity, and advanced packaging.
Main external dependency Foundries, HBM, packaging, Asian assembly, power and customer capex. Advanced-node equipment and memory access, software maturity, power and cooling at cluster scale. Equipment, materials, water, electricity, natural-hazard resilience, and geopolitical continuity.
Lock-in vector CUDA, interconnect, network operations, model optimization, facility design. CANN, UnifiedBus, Ascend tooling, domestic cloud and procurement ecosystem. Long design and qualification cycles, process rules, packaging and capacity reservations.
Do not compare by Vendor peak FLOPS or projected token claims alone. NPU count or vendor comparisons to NVIDIA alone. Headline fab count or announced investment alone.
Decisive measures Matched goodput/MW, utilization, porting cost, uptime, total system cost. Matched goodput/MW, reliability, developer support, software burden, total system cost. Qualified yield and capacity, package lead time, geographic share, recovery and time to volume.

Table 4.3. NVIDIA, Huawei, and TSMC occupy different hardware control points. Vendor claims require matched-workload and production evidence.

Hardware implications for SpaceXAI, OpenAI, and Anthropic

The three focal firms do not have the same hardware posture. Their dependence should be measured through contracts, productive fleet, software portability, power, and ownership of the facility rather than inferred from model rankings. The table records disclosed evidence available at the baseline; it does not assume that announced systems have shipped or that vendor-reported cluster size equals economically productive capacity.

Firm Disclosed hardware and build-out evidence Strategic implication Critical measure
SpaceXAI A 2024 NVIDIA disclosure described 100,000 Hopper GPUs as a historical milestone. 2026 SEC filings disclose capacity agreements tied to about 325,000 NVIDIA GPUs for Anthropic and about 110,000 for Google. Those contracts do not by themselves establish one non-overlapping productive fleet or internal allocation. Direct infrastructure and long-term customer contracts can support scale, but create delivery, utilization, termination, platform, site, power and capital-allocation risk. Installed, delivered, simultaneously available and utilized fleet; customer allocation; termination exposure; task goodput/MW; site diversity; refresh burden.
OpenAI OpenAI disclosed a 2026 financing that included $30B from NVIDIA and next-generation NVIDIA inference capacity. A 2025 letter of intent contemplated at least 10 GW of NVIDIA systems and up to $100B of progressive NVIDIA investment as capacity is deployed. Capital, compute and supplier relationships can accelerate deployment while creating negotiation, execution, concentration and demand-quality questions. Signed versus deployed capacity; investment conditions; provider mix; commissioned MW; utilization; third-party revenue; margin after depreciation and power.
Anthropic A diversified but interdependent portfolio spans AWS Trainium, Google TPU access, a SpaceX agreement tied to about 325,000 NVIDIA GPUs, and an NVIDIA agreement to invest up to $10B subject to conditions. Multiple accelerator paths can improve bargaining power and continuity, while increasing contractual, financing, orchestration and portability complexity. Workload share by platform; effective cost; porting time; retained quality; availability; contract concentration; independent fallback; related financing.

Table 4.4. Hardware posture of the focal firms at the July 24, 2026 baseline. Sources are company and supplier disclosures; counts and performance are not independently normalized.

Migrating bottlenecks: hardware scarcity and grid stress can coexist

Hardware scarcity and power stress are not opposites. Scarcity is measured against desired demand, not against zero deployment. Hundreds of thousands of accelerators can remain scarce while their absolute load is large enough to strain a local grid. The binding constraint can move by quarter, site, workload and supply-chain stage.

PRODUCTIVE COMPUTE = min(accelerators, HBM and packaging, networking, energized power and cooling, software-ready capacity, operator capacity, economic demand). The minimum is the binding constraint for that site and period. Substitution at one layer may relieve one bottleneck while deepening dependence at another.

Constraint regime What binds Observable symptom What would move the bottleneck
Chip-bound Accelerators, HBM, packaging or interconnect Energized halls or signed demand wait for systems; hardware prices and delivery times rise. New supply, better packaging yield, smaller models, or custom silicon.
Power-bound Grid connection, generation, transformers or cooling Delivered racks wait for power; sites curtail or miss commissioning dates. New generation, transmission, load flexibility, denser cooling, or geographic relocation.
Software-bound Portability, compilers, networking, schedulers or operators Installed systems show low productive utilization or poor reliability. Porting, orchestration, training, tooling, and matched-workload optimization.
Demand-bound Weak customer ROI or overbuilt capacity Available compute has falling utilization, prices or margins. Real workflow value, lower task cost, new demand, or project restructuring.

Table 4.4A. Migrating constraint regimes reconcile hardware scarcity with grid stress without assuming that one bottleneck is permanent.

Protocol implication: H1B tests shared hardware shocks; H2 tests financing and redeployment; H4 tests demand, efficiency, commissioned supply and latency. A site can move from chip-bound to power-bound to software-bound and finally demand-bound. The protocol records the binding constraint at each observation instead of forcing all periods into one static chokepoint story.

Data-center build-outs: from announced gigawatts to productive compute

Data-center build-outs were already central to this chapter through demand forecasts, interconnection, financing, environmental burden, and the distinction between announced and energized capacity. The hardware amendment makes the status ladder more exact. A project can have land and a power agreement while lacking transformers, chillers, HBM, packaged accelerators, switches, optical modules, software qualification, or an operating team. It can be energized while still failing to deliver production workloads.

Announced: A company or government states a target, site, investment, chip count, or gigawatt figure. This is evidence of intent, not capacity.

Contracted and financed: Land, power, construction, equipment, tenant, and financing agreements are executed with enforceable deposits, guarantees, and cancellation terms.

Under construction: Civil works, substations, generation, cooling, networking, and data halls have observable progress; duplicated announcements are removed.

Energized and rack-ready: The site has usable power and cooling at the density required by the selected hardware, not merely a grid connection at the property line.

Commissioned and software-usable: Accelerators, memory, network, storage, security, schedulers, and compilers pass acceptance tests and can execute the intended workloads.

Utilized and economically productive: Capacity delivers successful, latency-compliant work at a measured margin after power, depreciation, networking, maintenance, and human operations.

The named cases occupy different rungs. OpenAI's multi-gigawatt figures remain capacity under development rather than an energized baseline. SpaceX's 100,000-GPU Colossus milestone is historical; the more current public evidence is two customer agreements tied to approximately 325,000 and 110,000 NVIDIA GPUs, subject to ramp, delivery, termination and allocation conditions. Amazon's Project Rainier disclosure describes operational Trainium capacity, while utilization and comparable task economics remain private. TSMC's first Arizona fab had reached volume production, while later fabs and packaging facilities remained construction or announced capacity. Quarterly reviews must move each project through the status ladder instead of summing announcements, contracts and productive systems (OpenAI 2025c; SpaceX 2026e; SpaceX 2026g; Amazon 2026a; TSMC 2026a).

Power is not a fixed revenue ceiling - but it is a physical and latency constraint

A firm earns revenue from useful work produced by an energized system, not from electricity in the abstract. Chip efficiency, architecture, routing, cache reuse, utilization, task success, latency, reliability, price, and the share of customer value captured determine economics within a physical envelope. Better economics can raise revenue per megawatt; they cannot make a fixed local power and cooling system perform unlimited sequential computation.

CAPACITY AND ENERGY PRODUCTIVITY - MEASUREMENT WINDOW T
Economic output rate ($/hour) = verified outcomes delivered within the required service level x average economic value per outcome / hours in T.
Capacity productivity ($/MW-hour of energized capacity) = economic output in T / (average energized MW x hours in T).
Energy productivity ($/MWh consumed) = economic output in T / electricity consumed in T.
"Verified outcome" already includes correctness, reliability and latency compliance; do not multiply those factors again.

Illustrative worked example (hypothetical, not an observation): a system averages 20 MW of energized capacity for 24 hours, consumes 360 MWh, and delivers 120,000 verified outcomes worth $2.50 each. Economic output is $300,000; capacity productivity is $625 per MW-hour; energy productivity is about $833 per MWh. Outcomes missing the service-level deadline are excluded rather than discounted a second time.

Input Hypothetical value Calculation
Energized capacity window 20 MW x 24 hours 480 MW-hours of available capacity
Electricity consumed 360 MWh 75% average electricity-to-capacity ratio in the illustrative window
Verified outcomes 120,000 x $2.50 $300,000 economic output
Capacity productivity $300,000 / 480 MW-hours $625 per MW-hour
Energy productivity $300,000 / 360 MWh $833 per MWh

Table 4.5. Worked example showing consistent units for capacity and energy productivity.

The original claim should therefore be modified in both directions. Power scarcity increases the premium on successful work per unit of compute, but high-value tasks are not infinitely compressible. A provider can improve revenue through routing, utilization, smaller models, higher success, and higher-value outcomes. It can still hit local limits when workloads require many serial reasoning and tool-use steps, strict service levels, geographic residency, or continuous firm capacity.

The rebound effect

Efficiency does not necessarily reduce total electricity use. Lower cost per inference can make AI viable in more applications, increase frequency of use, and enable longer or more autonomous tasks. The result resembles a rebound effect: each task uses less compute, but the number and scope of tasks expand faster. The IEA's wide scenarios reflect this uncertainty. Energy efficiency is essential for affordability and reliability, yet it may accelerate aggregate demand rather than end the power build-out (IEA 2025a; IEA 2025b).

This is one reason the economic metric should be cost per successful task rather than cost per token. A system that generates many cheap tokens but requires repeated attempts, human correction, or downstream remediation can use more energy and labor than a more expensive model that completes the task correctly. Energy productivity and business productivity must be measured together.

Latency, sequential reasoning, and the local physical ceiling

Agentic systems can consume many inference calls, retrievals, tool executions, verifications, and retries for one user intent. Some work can be batched or shifted; real-time settlement, operations, diagnostics, coding feedback, and customer interactions often cannot. P50 averages are therefore insufficient. P95 and P99 completion latency, deadline compliance, queueing, retry depth, and the number of sequential model and tool steps belong in the economic denominator.

The relevant distinction is local physical throughput versus global revenue. A firm may grow by moving work, adding sites, raising price, or selecting higher-value tasks, but a specific grid zone and facility have finite power, cooling, networking, and latency capacity. The thesis is weakened if efficiency and new supply consistently improve cost and latency faster than demand; it is strengthened when demand, sequential depth, and service requirements produce persistent scarcity rents or missed service levels.

Firm load, flexible load, and interruptible intelligence

Not every AI workload requires identical reliability. Online consumer inference, safety systems, and live enterprise agents may need low latency and high availability. Training runs, evaluation batches, synthetic data generation, and some research tasks can be delayed, moved, or interrupted. This creates a potential new class of flexible industrial load. Providers that can shift work across time and regions may receive lower power costs and help integrate variable generation, while providers that demand continuous firm capacity will require more generation and grid investment.

The distinction should be measured rather than assumed. A data center can advertise flexibility while preserving contractual service levels that make actual curtailment rare. Analysts should track the percentage of load that is technically and contractually interruptible, the duration of acceptable curtailment, geographic mobility, backup generation, and the compensation paid for grid services.

Environmental and community constraints

Electricity is only one local input. Data centers can affect water withdrawal and consumption, land use, construction, backup generation, noise, heat, air emissions, transmission corridors, and municipal infrastructure. These effects vary sharply by cooling technology, climate, fuel mix, siting, and whether waste heat or reclaimed water is used. A credible infrastructure thesis must measure the burden rather than treating community acceptance as a permitting delay.

The dashboard should therefore track water intensity, carbon intensity by hour and region, backup-generator use, heat reuse, land and transmission footprint, local rates, tax and employment benefits, complaints, and community-benefit commitments. A project that improves national compute capacity while transferring unpriced environmental or infrastructure costs to one locality has not demonstrated socially efficient scale.

The financing bridge from AI optimism to the utility system

The physical build-out becomes a financial-system issue through contracts. Utilities may construct generation and transmission for a customer whose future demand is uncertain. Infrastructure funds and special-purpose vehicles may finance data centers under long-term leases. Cloud companies may sign take-or-pay agreements; chip suppliers may extend credit; strategic investors may invest in a lab that later spends the proceeds on the investor's infrastructure. Each arrangement reallocates risk.

The central questions are who guarantees the load, what happens if the customer delays or cancels, whether the asset has alternative users, and whether costs can be recovered from ordinary ratepayers. FERC Commissioner David Rosner emphasized cost-recovery agreements and protections against speculative projects in the 2026 large-load proceedings (FERC 2026b). This is the appropriate prudential lens: the danger is not high electricity use alone, but long-lived regulated or leveraged assets built against weakly secured demand.

Financing structure Who receives the upside Who may bear downside Key diligence question
Balance-sheet capex The frontier firm or hyperscaler Equity and unsecured creditors Can operating cash flow support depreciation, maintenance, and refresh cycles?
Project-financed data center Developer, lenders, equity sponsors, tenant Project lenders and sponsors; sometimes guarantor Is the lease investment-grade, transferable, and matched to debt maturity?
Utility rate-base investment Utility shareholders and large-load customer Ratepayers if cost allocation or exit protection is weak Does the customer fund dedicated facilities and termination risk?
Take-or-pay capacity Supplier and customer through secured availability Customer if demand disappoints; supplier if guarantee fails Are commitments enforceable, and is capacity fungible?
Strategic/vendor financing Both parties through ecosystem growth Investor-supplier and firm if demand is circular Is third-party demand independent of financing relationships?

Table 4.6. Infrastructure risk follows the contract, not the press release.

Circularity and demand quality

The frontier ecosystem contains observable capital-compute-revenue interdependence. OpenAI announced a $110 billion financing that included $30 billion from NVIDIA and said it secured next-generation NVIDIA inference compute. A prior OpenAI-NVIDIA letter of intent contemplated at least 10 GW of systems and up to $100 billion of progressive NVIDIA investment. NVIDIA disclosed an agreement, subject to conditions, to invest up to $10 billion in Anthropic. SpaceX disclosed customer agreements tied to approximately 325,000 NVIDIA GPUs for Anthropic at $1.25 billion per month and approximately 110,000 GPUs for Google at $920 million per month after ramp-up. These facts do not prove improper circularity; they show that capital, supplier demand, compute contracts and reported revenue can be mutually dependent (OpenAI 2026g; OpenAI 2025d; NVIDIA 2025a; SpaceX 2026e; SpaceX 2026g).

Amazon provides another named example of the same linkage. It announced a $50 billion OpenAI investment, beginning with $15 billion and followed by $35 billion when stated conditions are met. The partnership also names AWS as the exclusive third-party cloud distribution provider for OpenAI Frontier, commits OpenAI to consume 2 GW of Trainium capacity, and expands the companies' infrastructure agreement by $100 billion over eight years. The investment, distribution right and compute commitment should be recorded as separate flows rather than treated as one proof of demand (Amazon 2026b; OpenAI 2026h).

Headline contract arithmetic also requires denominator discipline. Dividing the disclosed monthly payment by the disclosed GPU count produces about $3,846 per GPU-month for Anthropic and $8,364 for Google. Those ratios are not comparable unit prices. Public disclosures do not normalize GPU generation, delivery timing, guaranteed versus maximum capacity, networking, storage, power, software, support, priority rights, duration, utilization or financing services.

Customer Disclosed monthly payment Disclosed GPU count Headline ratio Interpretation
Anthropic $1.25B About 325,000 About $3,846 per GPU-month Arithmetic ratio only; contract scope and ramp are not normalized.
Google $920M after ramp-up About 110,000 About $8,364 per GPU-month Arithmetic ratio only; services, generation and rights may differ.

Table 4.6A. Headline payment-per-disclosed-GPU ratios are not normalized contract prices.

The relevant demand-quality tests are cash from end users unrelated to strategic financing; contract cancellation and delivery rights; capacity utilization; renewal without new investment; gross margin after depreciation, networking and power; and whether customer commitments are enforceable, transferable and matched to asset life. Analysts should map each flow separately: equity investment, hardware purchase, cloud or compute commitment, capacity-backed financing, guarantee, customer payment and independent end-user revenue. Appearance on both sides of a transaction is a reason for measurement, not a finding of misconduct.

What to monitor

Energized capacity: Megawatts actually connected and available, not merely announced or permitted.

Utilization: Productive accelerator hours divided by available hours, adjusted for maintenance and reserved capacity.

Revenue and gross profit per MW: A measure of whether economic value is scaling faster than the physical footprint.

Successful tasks per kWh and latency: Include retries, routing, tool calls, human remediation, and P50/P95/P99 completion against the required service level.

Interconnection quality: Firm versus interruptible service, queue position, completion milestones, and customer-funded facilities.

Financial exposure: Debt, guarantees, termination payments, take-or-pay obligations, and ratepayer cost allocation.

Refresh burden: The cadence at which accelerators and cooling systems become economically obsolete relative to financing maturity.

Rebound and environmental intensity: Compare growth in total task demand with efficiency gains, and track water, emissions, land, backup generation, and community cost per successful task.

Hardware platform mix and portability: Productive workload share by accelerator and generation; CUDA, Neuron, CANN, and other porting time; performance retained after migration; software and operator availability.

Fabrication, memory, and packaging: Production-qualified leading-edge and advanced-packaging capacity by geography; HBM allocation; yield, lead time, recovery time, and customer concentration.

Commissioning conversion: Ordered chips, delivered racks, rack-ready MW, commissioned systems, productive accelerator hours, and economically productive MW reported as separate stages.

PRIMARY INFERENCE

Power scarcity does not prove that frontier labs must become services firms. It strengthens the incentive to increase the value captured from each successful task. Whether that incentive produces vertical integration depends on model pricing, customer bargaining power, partner economics, liability, and organizational capability.

By Rocky DeStefano, Apeiris AI. Version 1.2.1, evidence cutoff July 24, 2026. This is a monitoring protocol, not a tested theory, and no primary hypothesis has been adjudicated. Apeiris publishes the evidence model and the research artifact; it does not claim to currently monitor these markets or adjudicate these hypotheses. Copyright 2026 Apeiris. All rights reserved. This publication is separately and restrictively licensed and is not covered by the Apeiris corpus CC BY 4.0 license.