Pith. sign in

REVIEW 3 major objections 7 minor 28 references

Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency

T0 review · 3 major / 7 minor · reviewed 2026-07-09 · glm-5.2

Pith's one-line read AI Incumbents' Cost Advantage Never Closes, 2026-2030 Analysis Shows

desk verdict The depreciation-conveyor result is a genuine structural insight, but the solvency corridor rests on layered assumptions whose joint uncertainty isn't propagated. read the letter →

arxiv 2607.07207 v1 pith:NQC3IRGO submitted 2026-07-08 econ.GN cs.AIcs.ARcs.CEcs.PFq-fin.EC

classification econ.GNcs.AIcs.ARcs.CEcs.PFq-fin.EC
keywords memoryindustrypremiumsolvencytokenanalysisbandwidthcost
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that four concurrent forces — historic memory price surges, frontier-capable open-weight models, rapid inference-efficiency gains, and incumbent entry into compute resale — have fundamentally restructured the AI industry's economics around a mechanism the author calls the depreciation conveyor. Because hardware depreciates on a fixed schedule while memory prices remain elevated, whoever bought fleets in the previous cycle continuously receives newly-amortized (cheap) capacity. This creates a cost advantage, measured in dollars per petabyte of memory bandwidth delivered, that runs 3.2x in 2026, narrows to 1.9x in 2027, and re-widens to 3-4x by 2029-30. The advantage rotates among incumbents but never transfers to new entrants. The paper then derives a solvency corridor showing that the announced infrastructure buildout earns its cost of capital only if token demand sustains roughly 2x annual growth for four consecutive years with premium pricing remaining sticky in absolute terms. A measurement critique identifies five systematic upward biases in public token-demand trackers and argues that all optimistic projections predate a Q2 2026 regime break from token maximization to token minimization, making them upper bounds rather than base cases. Training economics bifurcate into a luxury tier ($18-38B per frontier run by 2030) and a mass tier (previous-frontier parity falling toward $5M), producing a luxury-car market structure. A vintage-breakeven analysis shows 2026 and 2028-29 capacity purchases are each fatally exposed to one pricing regime, with only the 2027 vintage robust in all futures. The paper concludes that AI infrastructure solvency now depends on three variables — monetized bandwidth demand, premium-price stickiness, and vintage ownership — rather than on gross token growth alone.

What carries the argument

depreciation conveyor

What would settle it

Two consecutive quarters of platform token growth annualizing below roughly 1.7x while premium API prices begin cutting toward the mass tier — the crash signature the paper itself identifies as the observable that would confirm the Commoditization Crash scenario.

Watch

Extended reading notes

Core claim

The paper's central discovery is the depreciation conveyor: the cost gap between entrants buying at current prices and incumbents holding depreciated fleets is structural and non-closing, because amortization continuously delivers newly cheap capacity to prior buyers faster than hardware prices normalize. This reframes the industry's central question from who has the best technology to who bought last cycle's hardware, and shows that the answer rotates among incumbents rather than opening to new entrants. Combined with a solvency corridor analysis (approximately 2x annual token growth required for four years) and a measurement critique showing public token trackers overstate monetizable by a

Load-bearing premise

The solvency corridor's central threshold of roughly 2x annual token-demand growth depends on a 30% per year bytes-per-token efficiency assumption that the paper itself identifies as its most consequential parameter. The paper notes that KV-cache compression is near its information-theoretic limit, which could mean efficiency decelerates (lowering the threshold to 1.6x and widening the corridor) or that token minimization accelerates efficiency gains beyond 30% (raising the阈值

Editorial extensions

If this is right

  • If the depreciation conveyor result holds, new entrants into AI compute provision face a structural cost disadvantage of 2-4x that no operating skill can overcome, making the Meta/xAI move into compute resale a defining institutional pattern rather than a transitional one.
  • The solvency corridor's ~2x annual token growth threshold means any sustained deceleration below this rate — from enterprise budget rationing, token minimization, or on-premises migration — would trigger impairments concentrated on peak-vintage holders, with the 2028-29 delivery window most exposed.
  • The training-cost divergence (luxury tier at $18-38B vs. mass tier at $5M) implies that closed labs can remain profitable only as premium providers serving a minority segment, and that infrastructure financed as if premium demand were mass-sized carries internally inconsistent assumptions.
  • China's domestic HBM production on a standard instruction set decouples its cost curve from the memory crisis, potentially allowing state-backed capacity to undercut Western providers in non-US markets during any downturn — one side's capacity is a bet, the other's is a plan.
  • The vintage breakeven analysis implies that infrastructure investment timing dominates operating skill: 2027 is the rational entry year in all scenarios, while 2026 and 2028-29 purchases each carry regime-specific fatal exposure.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This paper analyzes how memory scarcity, open-weight models, inference-efficiency gains, and compute-resale market entry restructure the AI industry over 2026–2030. The central analytical contribution is a bandwidth-denominated cost framework ($/PB) for inference economics, from which the author derives three main results: (1) a 'depreciation conveyor' that perpetually advantages incumbents over entrants, (2) a 'solvency corridor' requiring roughly 2× annual token-demand growth for infrastructure to earn its cost of capital, and (3) a U-shaped vintage-breakeven analysis showing 2027 capacity is robust while 2026 and 2028–29 vintages are exposed. The paper also provides a measurement critique of public token-demand trackers, scenario probabilities, and a decision procedure for greenfield custom-silicon entry. The framework is internally coherent and the parameters are explicit, though several load-bearing assumptions require stronger justification.

Significance. The paper makes a genuine methodological contribution by formulating inference economics in a bandwidth-denominated unit ($/PB, Equation 1) that cleanly separates hardware economics from model choice. The depreciation-conveyor result (Section 3) is a useful formalization of a real structural advantage. The vintage-breakeven analysis (Section 6) and the staged go/no-go procedure (Table 4) are falsifiable and decision-relevant. The measurement critique (Section 9.1) identifies concrete upward biases in public token trackers. The paper is unusually transparent about its own assumptions and their sensitivity, including an explicit stress test of the efficiency parameter (Section 7.1). However, the significance of several results is limited by the issues identified below.

major comments (3)
  1. Section 3 (depreciation conveyor) and Section 4 (capacity-bandwidth trade-off) contain a coupling that is not analyzed. The conveyor result—that the entrant gap 'never closes'—is stated as structurally robust, but Section 4 shows legacy fleets are 'capacity-poor': a GLM-5.2-class model requires ~750 GB residency, needing 10 H100s versus 3 GB300s. The paper notes that 'every success of [the compression] stack transfers workload from the premium niche to the incumbent floor' (Section 4), but does not model how the economically relevant workload share servable by legacy fleets evolves over time. At low efficiency gains (15%/yr, Section 7.1), legacy fleets may not compress future models sufficiently, shrinking the workload share where the incumbent floor is economically meaningful. This coupling—where low efficiency simultaneously widens the solvency corridor (helpful for solvency) but weaks
  2. Section 7.1 identifies the 30%/yr bytes-per-token efficiency baseline as 'the most consequential parameter' and provides a stress range of 15–45%/yr, moving the solvency threshold from 1.6× to 2.4×. However, the paper provides no independent derivation of the 30% baseline beyond noting it is 'consistent with' observed cost trends (Section 9) that the paper itself says 'conflate hardware efficiency, model right-sizing, routing, and margin compression.' Since the corridor's utility as a decision tool depends on which efficiency trajectory is realized, and since the quality-adjusted demand growth estimate of 2–3× (Section 9.1) sits within the 1.6–2.4× threshold band, the corridor's practical discriminating power is uncertain. The paper should either provide a bottom-up decomposition of the 30% baseline into its constituent drivers (KV compression, sparsity, routing, quantization) with各自的独立e
  3. The 'Q2 2026 regime break' from token maximization to token minimization (Section 9.2–9.3) is treated as an axiom: it is the basis for treating all pre-break projections as 'optimistic bounds' and for elevating the downside case to co-equal status. The paper itself acknowledges the post-break evidence window 'spans only weeks' (Section 9.3). Since this assumption is load-bearing for the scenario probability revisions (Section 8) and the projection-vintage argument, the paper should either (a) provide more systematic evidence that the regime break is persistent rather than a transient response to budget cycles, or (b) explicitly propagate the uncertainty about whether the break is real into the scenario probabilities, rather than treating it as established.
minor comments (7)
  1. Table 5 lists MBU values (0.50/0.55/0.60–0.62 for Hopper/Blackwell/Rubin) without citing a source. Given that MBU directly enters Equation 1 and affects all cost calculations, a source or justification for these specific values should be provided.
  2. Section 2.1 defines 'incumbent' as a market-position term (holding prior vintages), but the limit-pricing argument in Section 3 ('an incumbent pricing tokens anywhere above its own marginal cost but below the entrant's full cost') implicitly assumes incumbents have market power to set prices. The relationship between sunk-cost advantage and pricing power should be made explicit.
  3. The Goldman Sachs 24× projection (Section 9, [8]) is cited as 'almost exactly the solvency threshold derived independently in Section 7.' This is circular: the 2.2×/yr compound rate equals the threshold because both are derived from similar capacity and efficiency assumptions. The coincidence should not be presented as independent validation.
  4. Section 10.1 references 'the three-author Do We Still Need GPUs? series by Dongarra, Hoefler, and Matsuoka' [17,18]. The author of this paper is Matsuoka. This self-citation is relevant but should be disclosed more prominently given the paper's argument about the LX2 architecture.
  5. Figure 8 (described in text) projects demand under four cases against 'frozen announced capacity.' The paper notes capacity would chase demand in reality, but the frozen-capacity comparison could mislead readers into interpreting the gap as a forecast rather than a diagnostic. The caption or text should clarify this is an analytical device, not a prediction.
  6. The paper uses 'fatally exposed' (Sections 6, 11) to describe vintages that are underwater under one pricing regime. 'Fatally' is strong; 'structurally impaired under one regime' would be more precise.
  7. Reference [14] cites both a Substack aggregation and an official Chinese government figure (140T/day vs 180T/day) for the same metric. The discrepancy should be reconciled or the preferred figure identified to improve reproducibility.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: derivations are computed from stated parameters and external market data, not from self-cited results.

full rationale

The paper's central quantitative results—the depreciation conveyor cost gap (Section 3, Equation 1), the solvency corridor (Section 7, Figure 6), the vintage breakeven analysis (Section 6, Figure 4), and the greenfield custom-silicon outcome distribution (Section 10.2, Table 3)—are all derived from explicitly stated model parameters (Table 5) and external market data (DRAM prices, accelerator specs, token-volume trackers). None of these derivations reduce to their inputs by construction. The author cites two of his own works [17,18] on matrix-enhanced CPUs, but these are used as architectural context for the LineShine discussion (Section 10.1), not as inputs to any quantitative model or as premises for the paper's central claims. The scenario probabilities (Table 1) are explicitly labeled as 'elicited subjective probabilities—judgmental assessments,' not as derived predictions, so they cannot be circular. The solvency corridor threshold (~2x annual token growth) is derived from the interaction of announced capacity buildout, the 30%/yr efficiency assumption, and demand projections—each independently specified. While the skeptic correctly notes that the depreciation conveyor result and the solvency corridor both depend on the efficiency parameter, this is a coupling/correctness concern (the paper acknowledges it in Section 7.1), not circularity: the efficiency parameter is an externally assumed input, not a self-cited result or a fitted quantity renamed as a prediction. The paper is self-contained against external benchmarks.

Assumptions & free parameters 10 free parameters · 5 assumptions · 3 invented entities

The paper is transparent about its subjective probability assessments and flags its most consequential parameter (30%/yr efficiency). The main concern is that the solvency corridor, while derived from a clean cost model, depends on layered assumptions (efficiency rate, pricing regime, premium share, demand growth quality adjustment) whose joint uncertainty is not formally propagated. The regime break concept is the most speculative element, acknowledged by the author as insufficiently confirmed.

free parameters (10)
  • 30%/yr bytes-per-token efficiency = 0.30
    Baseline efficiency gain rate; acknowledged as most consequential parameter. Derived from observed cost trends that conflate multiple effects. Stress range 15-45%.
  • MBU values by generation = 0.50/0.55/0.60-0.62
    Memory-bandwidth utilization for Hopper/Blackwell/Rubin. Stated as assumptions without independent measurement.
  • System multiplier k = 1.5
    System cost as multiple of accelerator price. Standard but unverified for current builds.
  • Premium pricing regimes = coupled: 7x mass; sticky: $0.40/PB
    Two pricing poles used throughout; chosen to bracket the space but specific values are assumptions.
  • Plausible premium share band = 10-20%
    Assumed range of tokens earning premium pricing; used as benchmark in Figure 4 but not independently derived.
  • Scenario probabilities = 25/20/18/25/12%
    Explicitly stated as elicited subjective probabilities, not model outputs.
  • Greenfield outcome probabilities = 25/34/41%
    Elicited subjective probabilities conditioned on scenarios; bands of +/-7 points acknowledged.
  • Capacity buildout trajectory = +45/40/30/15%/yr
    Announced capex trajectory 2027-30; sourced from public announcements but uncertain.
  • Frontier training compute growth = 2.5x/yr
    Ambition scaling rate; stated as observation but future trajectory is an assumption.
  • Good-enough training cost decline = -40%/yr
    RL post-training + distillation cost decline rate; projected trend.
assumptions (5)
  • domain assumption Decode-phase inference is bandwidth-bound: throughput is proportional to delivered HBM bandwidth regardless of model identity.
    Section 2.1. This is the foundational axiom enabling the $/PB unit. It holds for autoregressive decode but not for prefill-heavy or compute-bound regimes. The paper acknowledges scope qualification but the entire framework depends on this being the dominant cost mode.
  • domain assumption Straight-line four-year amortization reflects the economic life of accelerator hardware.
    Section 2.1, Table 5. The depreciation conveyor result depends on this schedule; faster or slower amortization changes the gap trajectory.
  • ad hoc to paper The Q2 2026 regime break from token maximization to token minimization is real and persistent.
    Section 9.3. The paper acknowledges 'the post-Q2 2026 regime break is not yet confirmed by a sufficient time series' but still uses it to treat all pre-break projections as upper bounds. This is a load-bearing assumption for the Crash scenario's elevation to co-modal probability.
  • domain assumption Public token trackers systematically overstate monetizable demand due to five identified biases.
    Section 9.1. The five biases (supply-injected volume, substitution, subsidy, quality composition, base effects) are plausible but their magnitudes are not quantified, making the quality-adjusted 2-3x growth estimate a judgment call.
  • domain assumption AI4SIS demand is rationing-proof and provides an inelastic floor for bandwidth demand.
    Section 7. Mission-funded compute is treated as structurally insulated from enterprise budget cycles. This supports the Jevons Absorption scenario probability but is asserted rather than demonstrated with data.
invented entities (3)
  • Solvency corridor independent evidence
    purpose: Framework for determining whether announced capacity earns its cost of capital as a function of demand growth and efficiency trends.
    The corridor is defined by the intersection of capacity growth and bandwidth demand trajectories, both derived from stated parameters. It makes falsifiable predictions about utilization rates that can be checked against future data.
  • Depreciation conveyor independent evidence
    purpose: Mechanism explaining why the entrant-incumbent cost gap never closes: amortization continuously delivers cheap fleets to incumbents.
    This is a structural property of the cost model (Equation 1) given the amortization schedule, not a separately postulated entity. It produces testable predictions about cost gap trajectories.
  • Regime break (token maximization to minimization)
    purpose: Doctrinal inversion separating projection vintages; used to treat pre-Q2 2026 projections as upper bounds.
    The paper states the break is 'not yet confirmed by a sufficient time series' (Section 8). It is observed through indirect signals (budget rationing, metering rollouts, tokens-per-request collapse) but not independently verified. This is the most speculative invented entity in the framework.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency." pith.science (2026). https://pith.science/paper/NQC3IRGO

@misc{pith2026260707207,
  author       = {Pith},
  title        = {Pith review of: Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NQC3IRGO}},
  note         = {Machine review of arXiv:2607.07207}
}
abstract

We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.2), rapid inference-efficiency gains (near-Shannon-limit KV-cache compression, lightweight local runtimes), and the entry of Meta and xAI into compute resale on fleets bought before the memory repricing. Formulating inference economics in dollars per petabyte of bandwidth delivered (\$/PB) -- model-agnostic for bandwidth-bound decode -- we show the entrant-incumbent cost gap never closes: a depreciation conveyor delivers newly amortized fleets to incumbents faster than hardware prices normalize (3.2x in 2026, 1.9x in 2027, re-widening to 3-4x by 2029-30). Training bifurcates into a luxury tier (\$18-38B per frontier run by 2030) and a mass tier (previous-frontier parity via RL/distillation falling toward \$5M). Solvency of the announced buildout is confined to a corridor requiring roughly 2x annual token-demand growth for four years with sticky premium pricing; a measurement critique shows public token trackers overstate monetizable demand, and all pre-Q2-2026 projections predate the industry's shift from token maximization to token minimization. A vintage-breakeven analysis finds 2026 and 2028-29 capacity each fatally exposed to one pricing regime, with only the 2027 vintage robust. A greenfield custom-silicon entrant removes the merchant margin but not the memory premium (central outcome: 25% success/34% mediocre/41% loss, improvable via staged go/no-go gates). China's LineShine LX2 -- domestic HBM on a standard ISA -- decouples its cost curve from the memory crisis. Scenario probabilities: Rotating Landlord Oligopoly 25%, Commoditization Crash 25%, Jevons Absorption 20%, System-Layer Re-differentiation 18%, Geopolitical Bifurcation 12%. Solvency now depends on monetized bandwidth demand, premium stickiness, and vintage ownership.

Figures

Figures reproduced from arXiv: 2607.07207 by the authors.

Figure 1
Figure 1. New-build full cost vs incumbent sunk-fleet marginal floor, $/PB delivered, under two HBM branches. The gap narrows to 1.9× in 2027 and re-widens to 3–4× by 2029–30 [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Delivered bandwidth per watt. Rubin’s 2.2× advantage over H100 governs the power-slot displacement condition that ends each vintage’s service life in power-limited sites. workload from the premium niche to the incumbent floor. 5 Training-Cost Divergence and the Luxury Segmentation Frontier training ambition continues to scale at roughly 2.5× per year in compute, against hardware cost-per-FLOP improving only 25% per … view at source ↗
Figure 3
Figure 3. Training-cost divergence, log scale. Frontier runs reach $18–38B by 2030 while good-enough replication falls toward $5M — a 3,600–7,400× gap. 6 Vintage Breakeven and Pricing-Regime Sensitivity For each purchase vintage we compute the share of served tokens that must earn premium pricing for the fleet to break even, under two pricing regimes. In the coupled regime, routing arbitrage drags premium prices down with the… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Breakeven premium token share by vintage under coupled vs sticky premium pricing and two HBM branches, against a plausible 10–20% realized premium share (shaded). rescue. The solvency threshold is therefore approximately 2× annual token growth sustained for four consec…
Figure 5
Figure 5. Figure 5: Installed capacity vs bandwidth demand at various token growth rates, net of 30%/yr efficiency. The threshold for system solvency is approximately 2×/yr. deceleration is the strongest counterargument to the breakdown thesis — it widens the corridor by roughly a third —…
Figure 6
Figure 6. Figure 6: The solvency corridor. 2029 fleet utilization across token-demand growth and efficiency-gain rates; the 90% contour divides solvency from impairment. Baseline marked at (2.0×, 30%). month in May 2024 to 480 trillion at I/O 2025, 1.3 quadrillion by October 2025, and 3.2…
Figure 7
Figure 7. Figure 7: Observed token-demand trackers (log scale, native units per series). All primary sources are vendor￾reported or third-party aggregations; levels are not directly comparable across series but growth rates are. when, they are spectacular. Two further corrections follow. …
Figure 8
Figure 8. Figure 8: Delivered-bandwidth demand under four cases vs frozen announced capacity, 2026–2031, log scale. Case-to-scenario mapping shown in legend. (6× tapering to 2× by 2030, 30%/yr efficiency): demand exceeds frozen capacity twenty-fold by 2031, the shortage never ends within …
Figure 9
Figure 9. Figure 9: The circular-financing web (2026 reported figures; magnitudes approximate and moving) and the crash transmission path. Note that the aggregate mixes contract classes; see [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 28 canonical work pages

  1. [1]

    TrendForce and Counterpoint Research, DRAM contract price surveys, Q4 2025–Q2 2026 (Counterpoint: 80–90% QoQ into Q1 2026; TrendForce: HBM at 23% of DRAM wafer output). Via https://sourceability.com/post/tracking-memory-price-increases-across- the-last-several-quarters and https://tech-insider.org/memory-chip-shortage-2026- ai-consumer-electronics/

  2. [2]

    Chey Tae-won (SK Group chairman), remarks at NVIDIA GTC (March 2026) and Com- putex (June 2, 2026): shortage to persist until 2030; capacity doubling in five years. Tom’s Hardware: https://www.tomshardware.com/pc-components/dram/sk-hynix-to-double- memory-wafer-capacity-over-five-years ; TechSpot: https://www.techspot.com/news/ 111751-memory-chip-shortage...

  3. [3]

    VentureBeat: https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on- multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost ; N

    Z.ai, GLM-5.2 release (June 13–16, 2026; MIT weights on Hugging Face). VentureBeat: https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on- multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost ; N. Lambert, Interconnects: https://www.interconnects.ai/p/glm-52-is-the-step-change-for-open

  4. [4]

    TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

    TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate (ICLR 2026). arXiv:2504.19874, https://arxiv.org/abs/2504.19874;https://openreview.net/pdf?id=tO3ASKZlok

  5. [5]

    CNBC, July 1–2, 2026: https://www

    Meta compute-resale plans and 2026 capex guidance ($125–145B). CNBC, July 1–2, 2026: https://www. cnbc.com/2026/07/01/meta-stock-cloud-ai-compute.html ; shareholder-meeting remarks, May 27, 2026: https://www.cnbc.com/2026/05/27/mark-zuckerberg-says-meta-starting-cloud- business-on-the-table.html

  6. [6]

    CNBC, July 1, 2026:https://www.cnbc.com/2026/07/01/meta-stock-cloud-ai-compute.html

    SpaceX/xAI Colossus capacity leases (Anthropic $1.25B/month; Google $920M/month). CNBC, July 1, 2026:https://www.cnbc.com/2026/07/01/meta-stock-cloud-ai-compute.html

  7. [7]

    https://www.softwareseni.com/understanding-the-2025- dram-shortage-and-its-impact-on-cloud-infrastructure-costs/

    OpenAI Stargate DRAM procurement (up to 900k wafers/month, 40% of global output; Samsung and SK hynix agreements, October 2025). https://www.softwareseni.com/understanding-the-2025- dram-shortage-and-its-impact-on-cloud-infrastructure-costs/

  8. [8]

    Decoding the Agentic Economy

    Goldman Sachs Research, “Decoding the Agentic Economy” (May 2026): 24 × token growth to 120 quadrillion/month by 2030. https://www.goldmansachs.com/insights/articles/ai-agents- forecast-to-boost-tech-cash-flow-as-usage-soars

Show all 28 references
  1. [9]

    https://hai

    Stanford Institute for Human-Centered AI, AI Index Report 2026, inference-cost chapter. https://hai. stanford.edu/ai-index

  2. [10]

    https://a16z

    Andreessen Horowitz, LLMflation — LLM inference cost, 2024, with 2025–2026 updates. https://a16z. com/llmflation-llm-inference-cost/

  3. [11]

    https://openrouter

    OpenRouter, public usage rankings and State of AI usage data, April–June 2026. https://openrouter. ai/rankings

  4. [12]

    Pichai, Google I/O 2026 keynote (May 20, 2026): 3.2 quadrillion tokens/month, 7 × YoY; trajectory from 9.7T (2024) and 480T (2025)

    S. Pichai, Google I/O 2026 keynote (May 20, 2026): 3.2 quadrillion tokens/month, 7 × YoY; trajectory from 9.7T (2024) and 480T (2025). https://blog.google/innovation-and-ai/sundar-pichai-io- 2026/

  5. [13]

    Summarized in https://www.uncoveralpha.com/p/why-token-optimization-is- a-gift

    Microsoft Corp., FY2026 Q3 earnings call: 300+ Foundry customers on track for >1T tokens each, acceler- ating 30% QoQ. Summarized in https://www.uncoveralpha.com/p/why-token-optimization-is- a-gift

  6. [14]

    Third-party aggregation of Chinese token consumption ( 180T/day, Feb 2026; V olcano Engine 2T→63T/day): https://robonomics.substack.com/p/token-tracker-and-implications . Official National Data Administration figure: 140T/day, March 2026: https://news.cgtn.com/news/2026-06-18/...

  7. [15]

    LineShine Debuts at No. 1,

    TOP500 Project, 67th TOP500 list and announcement “LineShine Debuts at No. 1,” June 2026. https: //top500.org/lists/top500/2026/06/ ; https://top500.org/news/lineshine-debuts-no-1- top500-enters-new-global-exascale-era/

  8. [16]

    A Deep Dive On China’s LineShine All-CPU, Exaflops-Class Supercomputer,

    T. P. Morgan, “A Deep Dive On China’s LineShine All-CPU, Exaflops-Class Supercomputer,” The Next Platform, June 25, 2026, https://www.nextplatform.com/hpc/2026/06/25/a- deep-dive-on-chinas-lineshine-all-cpu-exaflops-class-supercomputer/5262439 ; Tom’s Hardware, https://www.tom...

  9. [17]

    Do We Still Need GPUs? Rethinking AI and Scientific Com- puting on Future CPUs,

    J. Dongarra, T. Hoefler, and S. Matsuoka, “Do We Still Need GPUs? Rethinking AI and Scientific Com- puting on Future CPUs,” CACM submission (short paper, 9 pp.), revised June 2026. Initial circulated draft (J. Dongarra, June 11, 2026): https://www.dropbox.com/scl/fi/ckafvk9ldx...

  10. [18]

    Do We Still Need GPUs? Rethinking AI and Scientific Computing on Matrix-Enhanced CPUs,

    J. Dongarra, T. Hoefler, and S. Matsuoka, “Do We Still Need GPUs? Rethinking AI and Scientific Computing on Matrix-Enhanced CPUs,” extended (merged) edition, 24 pp., empirical companion led by Matsuoka (A64FX/Fugaku and LX2 evidence); revision of July 6, 2026; arXiv posting in...

  11. [19]

    https://alatirok.com/ ai-circular-financing-explained/ ; https://tech-ish.com/2026/02/03/nvidia-openai- oracle-circular-financing-loop/

    Compiled analyses of circular AI financing (>$800B), H1 2026. https://alatirok.com/ ai-circular-financing-explained/ ; https://tech-ish.com/2026/02/03/nvidia-openai- oracle-circular-financing-loop/

  12. [20]

    Analysis: https:// intuitionlabs.ai/articles/oracle-openai-300b-deal-analysis

    Oracle Corp., FY2026 Q3 disclosures: RPO $523B (from $455B in Q1). Analysis: https:// intuitionlabs.ai/articles/oracle-openai-300b-deal-analysis

  13. [21]

    Reporting on the NVIDIA–OpenAI $100B letter of intent and its stall (early February 2026). Fortune: https://fortune.com/2026/02/02/why-did-oracle-stock-fall-openai-exposure- nvidia-microsoft/ ; https://finance.yahoo.com/news/stalled-nvidia-openai-megadeal- ai-131959187.html

  14. [22]

    Japan’s an AI Laggard. That Could Be Its Edge,

    C. Thorbecke, “Japan’s an AI Laggard. That Could Be Its Edge,” Bloomberg Opinion, May 24,

  15. [23]

    https://www.bloomberg.com/opinion/articles/2026-05-24/japan-s-an-ai-laggard- that-could-be-its-edge

  16. [24]

    https://www.nvidia.com/en-us/products/ workstations/dgx-spark/

    NVIDIA Corp., DGX Spark product documentation. https://www.nvidia.com/en-us/products/ workstations/dgx-spark/

  17. [25]

    PYM- NTS, May 2026: https://www.pymnts.com/news/artificial-intelligence/2026/token-shock- hits-silicon-valleys-biggest-spenders/

    Enterprise AI budget exhaustion and metering (Uber; Microsoft internal license cancellations). PYM- NTS, May 2026: https://www.pymnts.com/news/artificial-intelligence/2026/token-shock- hits-silicon-valleys-biggest-spenders/

  18. [26]

    arXiv:2606.11690,https://arxiv.org/abs/2606.11690

    Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estima- tion (2026): $0.21–$15.25 per million output tokens on identical H100 hardware across offered loads. arXiv:2606.11690,https://arxiv.org/abs/2606.11690

  19. [27]

    Via Business Insider: https://uk.finance.yahoo.com/news/ ubs-says-majority-enterprise-companies-090701077.html

    UBS research note (Keirstead, Arcuri, McGinnis; June 23, 2026): 60% of surveyed enterprises throttling AI spend; routing to open/Chinese models. Via Business Insider: https://uk.finance.yahoo.com/news/ ubs-says-majority-enterprise-companies-090701077.html

  20. [28]

    https://www.anthropic.com/news/claude-fable-5-mythos-5 ; context: https://www.bloomberg.com/professional/insights/markets/ai-boom-built-on- shaky-geopolitical-footing/ 22

    Anthropic, statement on export controls applied to Claude Fable 5 and Claude Mythos 5 (applied June 12, 2026; lifted June 30, 2026). https://www.anthropic.com/news/claude-fable-5-mythos-5 ; context: https://www.bloomberg.com/professional/insights/markets/ai-boom-built-on- shak...

Pith tools

Reviewed July 9, 2026 · model on record in the stance chip above.