{"id":"40e12294-1915-4b49-85a3-f75fd2e2d105","arxiv_id":"2504.15440","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using OpenRouter usage data, the paper documents that LLM adoption is fast, that model launches differ in whether they substitute for or expand demand, and that multihoming across models is common.","lead":"This paper uses data from OpenRouter, an LLM marketplace, to describe how demand changes when new models launch. It finds that new models gain users quickly, that some launches pull users from rivals while others expand the market, and that apps typically use several models at once.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Raw-token case studies are confounded by OpenRouter's rapid platform growth; substitution versus market-expansion conclusions need share-based normalization.","rationale":"The reader's weakest assumption identifies OpenRouter usage validity and the interrupted time-series assumptions as load-bearing. My concern is a sharper version of the same issue: even taking OpenRouter as the relevant market, the raw series are non-stationary because total platform usage grew by a factor of five during the sample. The paper's own Section 3 states assumptions of no concurrent events and no platform switching, but the platform growth documented in Figure 1 is a concurrent event affecting all models and all three case studies. Without normalization or a counterfactual, the visual distinction between substitution and market expansion is not identified. This does not undermine the descriptive facts about adoption speed or multihoming, and the author is appropriately cautious in several places, so the conditional verdict remains appropriate. I would not move the verdict to reject or accept because the concern is testable and the underlying data, if released, could resolve it with share-based or detrended analysis. The reader also flags the diversion-ratio language and lack of formal inference; I agree these are secondary, but the normalization issue is the most central threat to the headline conclusion about differentiation and pricing power.","tokens_in":10417,"tokens_out":3583,"duration_ms":38734,"concrete_test":"Recompute Figures 2, 3, and 4 using each model's daily share of total OpenRouter tokens (and, as a robustness check, share of total revenue), plus detrended residuals from a pre-release log-linear trend. If, after the Gemini 2.0 Flash release, incumbent shares (DeepSeek, Gemini 1.5, Llama variants) decline significantly, then the market-expansion conclusion of Fact 2 fails. For the Gemini 2.5 case, split the free experimental and paid preview series and rerun the comparison using only the paid release; if the free rate-limited variant drives the adoption curve, the persistence of Claude demand is not clean evidence of differentiation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference that LLM demand is horizontally and vertically differentiated rests on Fact 2 (Section 3.2): new model releases differ in whether they substitute demand from existing models or expand the market. The evidence for this fact is visual inspection of raw token levels in Figures 2, 3, and 4. However, Figure 1 shows that total OpenRouter tokens grew from roughly 50 billion to over 250 billion tokens per day during the sample period. In a rapidly growing platform, an incumbent model with flat or slowly rising raw tokens is losing share, not holding its own; conversely, a small raw drop in one series may be masked by the platform-wide upward trend. Thus the statement in Fact 2 that 'we see no obvious movement' in DeepSeek, GPT-4o, Llama, or Gemini 1.5 around releases is not evidence against substitution unless these series are compared with a counterfactual trend or normalized by total platform usage. The paper's Section 3 assumptions rule out concurrent events and users switching onto/off OpenRouter, but the platform's own fivefold growth is an internal confound affecting every case study, not an external caveat. A second, related issue is that the Gemini 2.5 case mixes a free, rate-limited experimental variant with a later paid preview; supply constraints can flatten the observed adoption curve, so the persistence of Claude demand relative to Gemini is not clean evidence of horizontal differentiation. The conclusion about pricing power would still be plausible, but the descriptive facts as presented do not yet separate differentiation from trend, supply constraints, or platform selection effects.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses daily, model-level token data scraped from OpenRouter between January 11 and April 11, 2025, plus weekly top-app information, to document three stylized facts: (i) new models are adopted quickly and stabilize within weeks; (ii) model releases differ in whether they substitute demand from existing models or expand the market; and (iii) apps multi-home across models. Three release events are studied as case studies: Claude 3.7 Sonnet, Gemini 2.0 Flash, and Gemini 2.5 Pro. The author argues that the evidence implies horizontal and vertical differentiation in the LLM market, with implications for pricing power. The analysis is explicitly descriptive and the paper states two interrupted-time-series assumptions (no concurrent events; no platform switching induced by releases). The central contribution is descriptive evidence from a novel marketplace dataset, but the substitution-versus-expansion interpretation rests on visual inspection of raw token levels rather than share-based or counterfactual comparisons.","tokens_in":10678,"tokens_out":2537,"duration_ms":25811,"significance":"If the stylized facts hold, the paper would be a useful early look at LLM demand with relevance for competition policy and IO modeling: persistent demand for non-frontier models and multi-homing would suggest that providers can sustain markups. The paper's strengths include a novel and carefully described dataset, transparency about scraping and cleaning, explicit statements of identifying assumptions, and a clearly written set of case studies with log-scale robustness figures in the appendix. However, the central claim about substitution versus market expansion is not yet supported by a formal or even a share-normalized analysis, and the multihoming fact rests on a top-apps snapshot. The contribution is therefore conditional on the robustness of the visual patterns to platform growth and to the selected case studies.","major_comments":[{"comment":"The substitution-versus-expansion interpretation is confounded by OpenRouter's rapid platform growth. Figure 1 shows total daily token usage rising from roughly 50 billion to over 250 billion tokens during the sample period, and revenue roughly doubling. In this environment, an incumbent model with flat raw token levels is losing share, not holding its own. The statement in Fact 2 that 'we see no obvious movement' in DeepSeek, GPT-4o, Llama, or Gemini 1.5 around releases is therefore not evidence against substitution unless these series are normalized by platform-wide usage or compared with a counterfactual trend. Conversely, a model whose raw tokens rise at the platform growth rate may simply be keeping pace. I suggest redoing the case studies with shares of total OpenRouter tokens (or tokens per active app) and reporting the series in levels and shares side by side. This is load-bearing because Fact 2 is the main evidence for the paper's differentiation conclusion.","section":"Section 3.2, Fact 2; Figures 2-4"},{"comment":"The Gemini 2.5 Pro case mixes two different products: a free, rate-limited experimental variant and a later paid preview. Rate limits can flatten observed adoption, so the fact that Claude 3.7 Sonnet retains higher demand after Gemini's release is not clean evidence of horizontal differentiation; it may reflect supply constraints. The paper itself notes that the experimental version was rate limited, but the implication for the substitution interpretation is not drawn. I recommend either restricting the case-study window to the paid preview, or explicitly discussing how rate limits affect the comparison with Claude, which was paid and not rate limited in the same way.","section":"Section 3.1, Gemini 2.5 Pro case; Figure 4"},{"comment":"The paper states that it provides 'some of the first evidence of diversion ratios in the market for LLMs,' but no diversion ratio is actually computed or defined beyond a qualitative discussion of whether demand for incumbent models fell. A diversion ratio is a quantitative object (the fraction of a new product's demand that comes from a specific incumbent), and it requires either a formal demand model or at least an event-study-style calculation on shares. As written, the paper provides case-study narrative evidence of substitution, not diversion-ratio evidence. I recommend either computing a simple share-based diversion measure (e.g., the change in an incumbent's share of a relevant app or market segment divided by the new model's share gain) or removing the diversion-ratio claim from the contributions.","section":"Section 2 and Section 4"},{"comment":"Fact 1 ('rapid adoption that stabilizes within weeks') is asserted from visual inspection of logs without any formal measure of the adoption speed or stabilization, and for Gemini 2.5 Pro the sample includes only a few weeks of data. I am not asking for a structural model, but a small set of summary statistics (e.g., days from launch to 80% of the first-peak level, or the slope of the log-token regression in the first two weeks versus the following four weeks) would substantiate the claim and make the three case studies comparable. Without such quantification, the 'stabilization within weeks' claim is hard to evaluate across models.","section":"Section 3.2, Fact 1"}],"minor_comments":[{"comment":"The decimal alignment in Table 1 is inconsistent (e.g., means are formatted with varying number of digits after the decimal point), and the table would be easier to read with a consistent format and clearly labeled units.","section":"Table 1"},{"comment":"The sentence 'Demand for Gemini 2.5 Pro rises quickly, but remains below that of Calude 3.7 Sonnet' contains a typo: 'Calude' should be 'Claude'. Also, the figure caption refers to 'Gemini 2.5 Flash Release,' but the model is Gemini 2.5 Pro, not Flash.","section":"Section 3.1, Figure 4 text"},{"comment":"The choice of comparison models is made 'ex-ante' according to the paper, but no criteria are given for which models are plotted in Figures 2-4 versus which are relegated to the full model list. Since the substitution interpretation depends on which incumbents are examined, I recommend listing the selection rule (e.g., all models in the same provider or price tier, or all models with above-median usage) and showing the full set of comparable models in an appendix figure, with the selected subset highlighted.","section":"Section 3.2, Fact 2"},{"comment":"The multihoming measure is based on weekly top-20 public apps per model, which means the app-level usage is censored at the 20th app. The paper acknowledges this, but the magnitude of the censoring is unclear. Reporting the fraction of each model's tokens accounted for by the top-20 apps (or, alternatively, the number of apps that hit the top-20 cap) would help readers judge whether the multihoming shares are representative.","section":"Section 3.2, Fact 3; Figures 5-6"},{"comment":"The discussion of Bertrand-Nash pricing equilibrium is brief and does not connect the stylized facts to a particular model of demand (e.g., logit or nested logit). A short paragraph explaining which demand primitives (cross-price elasticities, outside-good share, within-app variety) the paper's facts would inform would make the implications more concrete for IO readers.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a descriptive, data-descriptive contribution in the spirit of a 'first look' at a new market. The main risk is that the central interpretation (substitution vs. expansion) is not yet robust to the obvious platform-growth confound, and the presentation currently oversells the diversion-ratio contribution. With share-based normalization and a more careful treatment of the free-vs-paid Gemini 2.5 case, the paper could become a solid descriptive piece. My major_revision recommendation reflects that the central claim is defensible but needs load-bearing changes, not that the data or the research question are weak."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is worth reading and worth sending to referees. The OpenRouter scrape is a genuinely new window into LLM demand, and the three stylized facts — fast adoption, heterogeneous substitution patterns, and multihoming — are concrete and mostly new to the literature. The writing is honest about the descriptive nature of the exercise. But the central contrast between 'substitution' and 'market expansion' is not as clean as the paper suggests, for the reason the stress-test note points out. The case studies plot raw token levels while total platform tokens grow from roughly 50B to 250B per day. In that environment, an incumbent holding flat raw usage is losing share, and a new model's rise could reflect platform growth more than true market expansion. The paper's own Section 3 assumptions rule out concurrent events and people switching onto/off OpenRouter, but the platform's own trend is an internal confound, not an external caveat. The 'no obvious movement' in DeepSeek, GPT-4o, and Llama around Claude 3.7 is exactly the kind of visual claim that needs share-based normalization or a counterfactual trend. The Gemini 2.5 case is also muddied by the free, rate-limited experimental launch; the supply constraint can flatten adoption, so comparing it to Claude's persistence is not clean evidence of horizontal differentiation. This doesn't kill the paper's broader conclusion — horizontal and vertical differentiation is plausible and probably true — but the descriptive facts as presented don't separate differentiation from trend or supply constraints.\n\nThe multihoming fact is on firmer ground: the app-level top-20 data is coarse, but the concentration differences between coding and persona apps are themselves evidence of heterogeneous use cases. The paper's self-limitations are well stated, and the author doesn't oversell the platform's representativeness. The main problems are the raw-token analysis and the overclaim about 'first evidence of diversion ratios' — no actual diversion ratios are computed. Also, no data or code is released; for a descriptive paper built on a scrape, that's a bigger omission than it would be elsewhere.\n\nWho it's for: anyone working on the economics of AI, platform competition, or demand estimation. It's a solid first measurement that should go through peer review, with the expectation that the author either normalizes by platform trend or adds formal tests, and releases the data. I'd cite it once it's revised.","headline":"The paper is a useful first measurement of LLM demand, but its substitution-versus-expansion claim is muddier than it looks once you account for OpenRouter's own 5x growth during the sample.","tokens_in":11174,"tokens_out":4054,"would_cite":true,"duration_ms":32754,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Three model releases on OpenRouter show that LLM demand is horizontally and vertically differentiated, so non-frontier providers can retain demand and pricing power.","keywords":["LLM demand","horizontal differentiation","vertical differentiation","market expansion","substitution","multihoming","OpenRouter","diversion ratio"],"falsifier":"Track token-level usage around a future major release where no concurrent events are known: if the release causes the dominant incumbent's usage to collapse while the new model grows, substitution rather than expansion would dominate, and the horizontal-differentiation reading would be weakened. Likewise, if the top apps' weekly model mixes converge on one model after successive releases, multihoming would fade and the conclusion that many models keep stable niches would fail.","tokens_in":10221,"feed_emoji":"🤖","tokens_out":4606,"duration_ms":42144,"temperature":0.7,"pith_summary":"This paper uses daily token usage from the OpenRouter marketplace to ask whether LLMs are interchangeable commodities or differentiated products. It documents three facts: new models gain demand within weeks; different releases either steal usage from incumbents or expand the market without visible cannibalization; and individual apps routinely use several models side by side. The paper argues these patterns imply the LLM market has both vertical and horizontal differentiation. If correct, a provider that is not at the top of every benchmark can still retain demand and charge a markup. This matters because it determines whether competition in AI is a race to the bottom or a market where quality differences, brand, and fit support ongoing profits.","feed_headline":"LLM demand is divided, not winner-take-all","feed_subtitle":"OpenRouter data: new models can grow the market, and apps use several models at once.","key_machinery":"The analysis rests on interrupted time-series case studies of three model releases, comparing token usage of the new model against ex-ante comparable incumbents in the days around each launch, under stated assumptions of no concurrent market shocks and no platform switching. For the multihoming fact, the paper uses OpenRouter's weekly top-20-app-per-model token counts, across models, to reconstruct app-level model mixes for four apps. These two devices turn public scraping data into evidence about substitution versus expansion and about within-app model variety.","core_discovery":"The paper's central claim is that demand for LLMs is not winner-take-all: models are differentiated along both quality and taste dimensions. Looking at the releases of Claude Sonnet 3.7, Gemini 2.0 Flash, and Gemini 2.5 Pro on OpenRouter, the paper observes that each release is adopted quickly, but the three releases behave differently. Claude 3.7 draws demand almost entirely from Claude 3.5, while Gemini releases appear to grow the market rather than cannibalize rivals. The same apps use a mix of models, with coding apps splitting usage mainly between Claude Sonnet 3.7 and Gemini 2.5 Pro, and chat and persona apps favoring different cheap models. The author reads these patterns as evidence that models compete partly on objective quality and partly on horizontal attributes such as speed, price, latency, branding, and integration with particular coding tools, so no single model captures a price point.","pith_inferences":["Editorial inference: because OpenRouter usage skews toward coding and persona or chat apps, the observed horizontal differentiation may be stronger for those use cases than for general enterprise text generation; the same analysis on native ChatGPT or enterprise logs could show more concentration.","Editorial inference: the prevalence of multihoming suggests the app or routing layer, not the model alone, may hold significant market power, since apps decide the mix and could steer usage across vendors.","Editorial inference: a natural extension is to estimate diversion ratios and price elasticities from future releases with clearer price variation, treating each release as a quasi-experiment while controlling for the platform-growth trend."],"forward_implications":["Providers whose models trail on headline benchmarks can still keep substantial usage and charge markups, as long as they hold a horizontal niche such as coding fit, speed, or integration.","A new release need not steal share to be valuable: market-expanding releases grow total platform usage, so competition can raise overall LLM adoption rather than merely reshuffle it.","Substitution can be highly localized, as with Claude 3.7 drawing from Claude 3.5, so a provider's main competitive threat may be its own next model.","App-level demand is a portfolio of models, not a single choice, so model vendors can coexist inside the same application and usage-based revenues are split.","If differentiated demand persists, price competition may be muted even as the frontier advances, because customers will not all flock to the single best benchmark model."],"supporting_citations":[{"why":"Supplies the diversion-ratio concept the paper uses to interpret whether a new model's demand comes from competitors or from new users.","marker":"Conlon and Mortimer (2021)"},{"why":"Frames the demand-modeling perspective on product differentiation and substitution in which the case-study evidence is interpreted.","marker":"Dube (2019)"},{"why":"Provides evidence on which economic tasks are performed with AI, motivating the demand-side question of which models get used.","marker":"Handa et al. (2025)"},{"why":"Establishes the labor-task exposure context that makes the market demand for better AI models an economically meaningful question.","marker":"Eloundou et al. (2024)"}],"fun_headline_variants":["LLM demand splits: some models expand, some substitute","Apps multihome: LLM demand isn't winner-take-all","New LLMs can expand the market, not just steal users","Model releases differ: some substitute, some grow demand"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"OpenRouter usage is a valid window into LLM demand, and the case-study comparisons are uncontaminated by concurrent events or by releases causing users to switch onto or off the platform.","fun_headline_variants_meta":{"raw":{"variants":["LLM demand splits: some models expand, some substitute","Apps multihome: LLM demand isn't winner-take-all","New LLMs can expand the market, not just steal users","Model releases differ: some substitute, some grow demand"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000374,"raw_usage":{"total_tokens":1933,"prompt_tokens":815,"completion_tokens":1118,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":431,"completion_tokens_details":{"reasoning_tokens":1048}},"tokens_in":431,"tokens_out":1118,"duration_ms":7868,"temperature":1.0,"reasoning_tokens":1048,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:25:28.913226+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Track token-level usage around a future major release where no concurrent events are known: if the release causes the dominant incumbent's usage to collapse while the new model grows, substitution rather than expansion would dominate, and the horizontal-differentiation reading would be weakened. Likewise, if the top apps' weekly model mixes converge on one model after successive releases, multihoming would fade and the conclusion that many models keep stable niches would fail.","supporting_citations":[{"cited_title":"Empirical properties of diversion ratios","cited_arxiv_id":null,"evidence_quote":"Supplies the diversion-ratio concept the paper uses to interpret whether a new model's demand comes from competitors or from new users."},{"cited_title":"Microeconometric models of consumer demand","cited_arxiv_id":null,"evidence_quote":"Frames the demand-modeling perspective on product differentiation and substitution in which the case-study evidence is interpreted."},{"cited_title":"GPTs are GPTs: Labor market impact potential of LLMs","cited_arxiv_id":null,"evidence_quote":"Establishes the labor-task exposure context that makes the market demand for better AI models an economically meaningful question."}],"review_version":1}