REVIEW 4 major objections 5 minor 2 cited by
Demand for LLMs: Descriptive Evidence on Substitution, Market Expansion, and Multihoming
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Three model releases on OpenRouter show that LLM demand is horizontally and vertically differentiated, so non-frontier providers can retain demand and pricing power.
desk verdict The paper is a useful first measurement of LLM demand, but its substitution-versus-expansion claim is muddier than it looks once you account for OpenRouter's own 5x growth during the sample. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The analysis rests on interrupted time-series case studies of three model releases, comparing token usage of the new model against ex-ante comparable incumbents in the days around each launch, under stated assumptions of no concurrent market shocks and no platform switching. For the multihoming fact, the paper uses OpenRouter's weekly top-20-app-per-model token counts, across models, to reconstruct app-level model mixes for four apps. These two devices turn public scraping data into evidence about substitution versus expansion and about within-app model variety.
What would settle it
Track token-level usage around a future major release where no concurrent events are known: if the release causes the dominant incumbent's usage to collapse while the new model grows, substitution rather than expansion would dominate, and the horizontal-differentiation reading would be weakened. Likewise, if the top apps' weekly model mixes converge on one model after successive releases, multihoming would fade and the conclusion that many models keep stable niches would fail.
Extended reading notes
Core claim
The paper's central claim is that demand for LLMs is not winner-take-all: models are differentiated along both quality and taste dimensions. Looking at the releases of Claude Sonnet 3.7, Gemini 2.0 Flash, and Gemini 2.5 Pro on OpenRouter, the paper observes that each release is adopted quickly, but the three releases behave differently. Claude 3.7 draws demand almost entirely from Claude 3.5, while Gemini releases appear to grow the market rather than cannibalize rivals. The same apps use a mix of models, with coding apps splitting usage mainly between Claude Sonnet 3.7 and Gemini 2.5 Pro, and chat and persona apps favoring different cheap models. The author reads these patterns as evidence that models compete partly on objective quality and partly on horizontal attributes such as speed, price, latency, branding, and integration with particular coding tools, so no single model captures a price point.
Load-bearing premise
OpenRouter usage is a valid window into LLM demand, and the case-study comparisons are uncontaminated by concurrent events or by releases causing users to switch onto or off the platform.
Editorial extensions
If this is right
- Providers whose models trail on headline benchmarks can still keep substantial usage and charge markups, as long as they hold a horizontal niche such as coding fit, speed, or integration.
- A new release need not steal share to be valuable: market-expanding releases grow total platform usage, so competition can raise overall LLM adoption rather than merely reshuffle it.
- Substitution can be highly localized, as with Claude 3.7 drawing from Claude 3.5, so a provider's main competitive threat may be its own next model.
- App-level demand is a portfolio of models, not a single choice, so model vendors can coexist inside the same application and usage-based revenues are split.
- If differentiated demand persists, price competition may be muted even as the frontier advances, because customers will not all flock to the single best benchmark model.
Reading between the lines
- Editorial inference: because OpenRouter usage skews toward coding and persona or chat apps, the observed horizontal differentiation may be stronger for those use cases than for general enterprise text generation; the same analysis on native ChatGPT or enterprise logs could show more concentration.
- Editorial inference: the prevalence of multihoming suggests the app or routing layer, not the model alone, may hold significant market power, since apps decide the mix and could steer usage across vendors.
- Editorial inference: a natural extension is to estimate diversion ratios and price elasticities from future releases with clearer price variation, treating each release as a quasi-experiment while controlling for the platform-growth trend.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper uses daily, model-level token data scraped from OpenRouter between January 11 and April 11, 2025, plus weekly top-app information, to document three stylized facts: (i) new models are adopted quickly and stabilize within weeks; (ii) model releases differ in whether they substitute demand from existing models or expand the market; and (iii) apps multi-home across models. Three release events are studied as case studies: Claude 3.7 Sonnet, Gemini 2.0 Flash, and Gemini 2.5 Pro. The author argues that the evidence implies horizontal and vertical differentiation in the LLM market, with implications for pricing power. The analysis is explicitly descriptive and the paper states two interrupted-time-series assumptions (no concurrent events; no platform switching induced by releases). The central contribution is descriptive evidence from a novel marketplace dataset, but the substitution-versus-expansion interpretation rests on visual inspection of raw token levels rather than share-based or counterfactual comparisons.
Significance. If the stylized facts hold, the paper would be a useful early look at LLM demand with relevance for competition policy and IO modeling: persistent demand for non-frontier models and multi-homing would suggest that providers can sustain markups. The paper's strengths include a novel and carefully described dataset, transparency about scraping and cleaning, explicit statements of identifying assumptions, and a clearly written set of case studies with log-scale robustness figures in the appendix. However, the central claim about substitution versus market expansion is not yet supported by a formal or even a share-normalized analysis, and the multihoming fact rests on a top-apps snapshot. The contribution is therefore conditional on the robustness of the visual patterns to platform growth and to the selected case studies.
major comments (4)
- [Section 3.2, Fact 2; Figures 2-4] The substitution-versus-expansion interpretation is confounded by OpenRouter's rapid platform growth. Figure 1 shows total daily token usage rising from roughly 50 billion to over 250 billion tokens during the sample period, and revenue roughly doubling. In this environment, an incumbent model with flat raw token levels is losing share, not holding its own. The statement in Fact 2 that 'we see no obvious movement' in DeepSeek, GPT-4o, Llama, or Gemini 1.5 around releases is therefore not evidence against substitution unless these series are normalized by platform-wide usage or compared with a counterfactual trend. Conversely, a model whose raw tokens rise at the platform growth rate may simply be keeping pace. I suggest redoing the case studies with shares of total OpenRouter tokens (or tokens per active app) and reporting the series in levels and shares side by side. This is load-bearing because Fact 2 is the main evidence for the paper's differentiation conclusion.
- [Section 3.1, Gemini 2.5 Pro case; Figure 4] The Gemini 2.5 Pro case mixes two different products: a free, rate-limited experimental variant and a later paid preview. Rate limits can flatten observed adoption, so the fact that Claude 3.7 Sonnet retains higher demand after Gemini's release is not clean evidence of horizontal differentiation; it may reflect supply constraints. The paper itself notes that the experimental version was rate limited, but the implication for the substitution interpretation is not drawn. I recommend either restricting the case-study window to the paid preview, or explicitly discussing how rate limits affect the comparison with Claude, which was paid and not rate limited in the same way.
- [Section 2 and Section 4] The paper states that it provides 'some of the first evidence of diversion ratios in the market for LLMs,' but no diversion ratio is actually computed or defined beyond a qualitative discussion of whether demand for incumbent models fell. A diversion ratio is a quantitative object (the fraction of a new product's demand that comes from a specific incumbent), and it requires either a formal demand model or at least an event-study-style calculation on shares. As written, the paper provides case-study narrative evidence of substitution, not diversion-ratio evidence. I recommend either computing a simple share-based diversion measure (e.g., the change in an incumbent's share of a relevant app or market segment divided by the new model's share gain) or removing the diversion-ratio claim from the contributions.
- [Section 3.2, Fact 1] Fact 1 ('rapid adoption that stabilizes within weeks') is asserted from visual inspection of logs without any formal measure of the adoption speed or stabilization, and for Gemini 2.5 Pro the sample includes only a few weeks of data. I am not asking for a structural model, but a small set of summary statistics (e.g., days from launch to 80% of the first-peak level, or the slope of the log-token regression in the first two weeks versus the following four weeks) would substantiate the claim and make the three case studies comparable. Without such quantification, the 'stabilization within weeks' claim is hard to evaluate across models.
minor comments (5)
- [Table 1] The decimal alignment in Table 1 is inconsistent (e.g., means are formatted with varying number of digits after the decimal point), and the table would be easier to read with a consistent format and clearly labeled units.
- [Section 3.1, Figure 4 text] The sentence 'Demand for Gemini 2.5 Pro rises quickly, but remains below that of Calude 3.7 Sonnet' contains a typo: 'Calude' should be 'Claude'. Also, the figure caption refers to 'Gemini 2.5 Flash Release,' but the model is Gemini 2.5 Pro, not Flash.
- [Section 3.2, Fact 2] The choice of comparison models is made 'ex-ante' according to the paper, but no criteria are given for which models are plotted in Figures 2-4 versus which are relegated to the full model list. Since the substitution interpretation depends on which incumbents are examined, I recommend listing the selection rule (e.g., all models in the same provider or price tier, or all models with above-median usage) and showing the full set of comparable models in an appendix figure, with the selected subset highlighted.
- [Section 3.2, Fact 3; Figures 5-6] The multihoming measure is based on weekly top-20 public apps per model, which means the app-level usage is censored at the 20th app. The paper acknowledges this, but the magnitude of the censoring is unclear. Reporting the fraction of each model's tokens accounted for by the top-20 apps (or, alternatively, the number of apps that hit the top-20 cap) would help readers judge whether the multihoming shares are representative.
- [Section 4] The discussion of Bertrand-Nash pricing equilibrium is brief and does not connect the stylized facts to a particular model of demand (e.g., logit or nested logit). A short paragraph explaining which demand primitives (cross-price elasticities, outside-good share, within-app variety) the paper's facts would inform would make the implications more concrete for IO readers.
Circularity Check
No circularity: the paper is a descriptive data analysis with no fitted parameters, no predictions derived from fitted inputs, and no load-bearing self-citations.
full rationale
The paper's claims are descriptive summaries of observed OpenRouter usage data. It documents three stylized facts: rapid adoption of new models, differing substitution versus market-expansion patterns across releases, and multihoming among apps. None of these facts is derived from a model fitted to the data; they are visual and tabular characterizations of token-usage series. The central inference about horizontal and vertical differentiation is an interpretation of those facts, not an input to their construction. There is no parameter estimation, no 'prediction' of a quantity that was used to fit anything, and no self-citation chain invoked to justify the core claims. The case-study assumptions in Section 3 are stated explicitly and are threats to causal validity, not circularity: observing raw token levels without share normalization may confound substitution with platform growth, but this is a measurement/identification concern, not a reduction of the conclusion to its premise. The multihoming measure is constructed from OpenRouter's top-app-per-model reports, but the fact that some apps appear next to multiple models is an empirical observation rather than a definitional equivalence. The author also candidly lists market-coverage limitations. Therefore no circular step meets the evidentiary bar of this analysis.
Assumptions & free parameters
free parameters (2)
- Model creation cutoff date =
Jan 1, 2024
- Model class labels (SOTA, Fast & Cheap, Old)
assumptions (3)
- domain assumption OpenRouter usage is a valid window into LLM demand despite excluding native apps like ChatGPT and apps such as Cursor.
- domain assumption For each case study, no concurrent events affected demand to a similar degree, and the model release did not cause users to switch into or out of OpenRouter.
- domain assumption Weekly top-apps data is sufficient to construct app-level model usage.
Cite this review
Pith. "Pith review of Demand for LLMs: Descriptive Evidence on Substitution, Market Expansion, and Multihoming." pith.science (2026). https://pith.science/paper/22YWZJG4
@misc{pith2026250415440,
author = {Pith},
title = {Pith review of: Demand for LLMs: Descriptive Evidence on Substitution, Market Expansion, and Multihoming},
year = {2026},
howpublished = {\url{https://pith.science/paper/22YWZJG4}},
note = {Machine review of arXiv:2504.15440}
}
read the original abstract
This paper documents three stylized facts about the demand for Large Language Models (LLMs) using data from OpenRouter, a prominent LLM marketplace. First, new models experience rapid initial adoption that stabilizes within weeks. Second, model releases differ substantially in whether they primarily attract new users or substitute demand from competing models. Third, multihoming, using multiple models simultaneously, is common among apps. These findings suggest significant horizontal and vertical differentiation in the LLM market, implying opportunities for providers to maintain demand and pricing power despite rapid technological advances.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 2 Pith papers
-
Freemium Is All You Need
Under a stylized uniform-value model, an optimal freemium policy can be expressed by two value thresholds, but the paper's case analysis and dynamic optimality claim are not correct.
-
LLM Performance for Code Generation on Noisy Tasks
LLMs solve heavily obfuscated benchmark tasks, and performance decay under obfuscation differs sharply between old and new datasets, which the authors interpret as a signature of training-data contamination.
Reference graph
Works this paper leans on
-
[1]
The Simple Macroeconomics of AI
Acemoglu, Daron. 2024. “The Simple Macroeconomics of AI.” w32487, National Bureau of Eco- nomic Research
work page 2024
-
[2]
Brynjolfsson, Erik, Danielle Li, and Lindsey Raymond. 2025. “Generative AI at work.” The Quar- terly Journal of Economics : qjae044
work page 2025
-
[3]
Empirical properties of diversion ratios
Conlon, Christopher, and Julie Holland Mortimer. 2021. “Empirical properties of diversion ratios.” The RAND Journal of Economics 52 (4): 693–726. Dell’Acqua, Fabrizio, Raj Agarwal, Marco Iansiti, and Shi Zheng. 2023. “Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality...
work page 2021
-
[4]
Microeconometric models of consumer demand
Dube, Jean-Pierre. 2019. “Microeconometric models of consumer demand.” In Handbook of the Economics of Marketing , edited by Jean-Pierre Dube and Peter E Rossi, vol. 1, chap. 1, 1–68
work page 2019
-
[5]
GPTs are GPTs: Labor market impact potential of LLMs
Eloundou, Tyna, Sam Manning, Pamela Mishkin, and Daniel Rock. 2024. “GPTs are GPTs: Labor market impact potential of LLMs.” Science 384 (6702): 1306–1308
work page 2024
-
[6]
Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conver- sations
Handa, Kunal, Tyna Eloundou, Anand Thakkar, Richard Mansfield, and Daniel Rock. 2025. “Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conver- sations.” arXiv preprint arXiv:2503.04761 . A Addition Figures A.1 Claude 3.7 Sonnet Release Figure A1: Claude 3.7 Sonnet Release Comparison (Log Scale) 12 Figure A2: Gemini 2.5 Pro Rel...
arXiv 2025
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.