{"id":"06347975-0314-4ac9-860e-30e560934fae","arxiv_id":"2607.07652","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":7,"one_line_summary":"ChatGPT retains 94.8% of information-seeking occasions without outbound referrals and wider access displaces 9.4% of traditional search queries, with losses concentrated on informational and ad-supported destinations.","lead":"This paper uses U.S. desktop clickstream data to show that ChatGPT Search sends outbound clicks in only 5.2% of sessions versus 31.1% for Google, and that wider ChatGPT Search access reduces traditional search queries by 9.4%. A smart generalist should read it because it quantifies how AI search may be breaking the web's traffic-for-content bargain that publishers, advertisers, and platforms have relied on for two decades.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Referral-ratio comparison may conflate different task mixes across intermediaries; the 5.2% vs 31.1% gap is descriptive, not causal, and the paper's context-control check is too coarse to rule out task selection as the primary driver.","rationale":"The reader correctly identified parallel trends as the weakest assumption for the causal displacement claim, and that concern is real but well-addressed (pre-trend tests, L-vs-A comparison, Callaway-Sant'Anna robustness). My concern targets the descriptive referral-ratio comparison instead, which the reader treated as robust. The paper itself acknowledges the task-selection limitation (Section 3.1) and explicitly states the comparison 'does not separate differences in the tasks brought to each intermediary from differences created by the intermediaries.' The surrounding-context check (Table OA.2.4) is the paper's main defense, but it operates at the level of broad content categories (12 buckets), which is too coarse to hold task difficulty or complexity fixed. A 'reference/knowledge' context could include anything from a quick dictionary lookup to a multi-paragraph research question. The +23.1pp 'solo' share for ChatGPT is particularly telling: ChatGPT sessions are much more likely to occur with no other browsing context, suggesting users bring different, more self-contained tasks. That said, several factors mitigate the concern's severity: (1) the paper is transparent about this limitation and frames the finding as descriptive; (2) industry benchmarks (Peec AI 3-5%, Similarweb ~7%) independently corroborate the low ChatGPT referral rate; (3) the message-level ratio (1.2%, Table OA.2.2) is even lower, consistent with absorption rather than task effects; (4) the causal displacement estimate is independent of this comparison and has its own (adequate) identification support. The concern does not invalidate the paper's contribution but suggests the 5.2% vs 31.1% gap should be read as an upper bound on intermediary-driven retention, with task-mix differences potentially inflating it. The verdict remains ACCEPT because the paper is appropriately cautious in its framing, the causal estimate is independently identified, and the descriptive finding is corroborated by external benchmarks even if its exact magnitude is uncertain.","tokens_in":53093,"tokens_out":897,"duration_ms":917047,"concrete_test":"Construct a within-household comparison restricted to sessions where the surrounding context contains a search-engine query within ±5 minutes of the ChatGPT session or Google query (i.e., both intermediaries are invoked in the same narrow task episode). If the referral-ratio gap persists at similar magnitude in this matched-task-episode sample, task-selection differences are unlikely to explain the gap. If the gap shrinks substantially, task mix is a primary driver.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central descriptive claim—ChatGPT routes only 5.2% of sessions vs Google's 31.1%—compares conversation sessions to search queries as equivalent 'information-seeking occasions.' But these units may capture systematically different task distributions. A ChatGPT conversation session (one UUID, 1-hour gap) may bundle multiple information needs or represent longer, more complex tasks than a single Google query. The paper acknowledges this: 'this comparison does not separate differences in the tasks brought to each intermediary from differences created by the intermediaries' (Section 3.1). The surrounding-context fixed effect (Table OA.2.4) is offered as a check, but it controls only for the dominant content category of nearby browsing—a coarse proxy. It cannot distinguish, say, a quick factual lookup from a multi-step research task within the same 'reference/knowledge' context bucket. If ChatGPT users bring harder, more self-contained tasks (the +23.1pp 'solo' share supports this), the routing gap partly reflects task selection rather than intermediary absorption. The paper frames this correctly as descriptive, but the headline claim and economic-bargain framing lean on it as if it reflects structural intermediary behavior. The causal displacement estimate (9.4%) is independent of this concern and rests on separate identification; the concern is specifically about whether the 5.2% vs 31.1% gap measures intermediary-driven retention vs task-mix differences.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This paper uses URL-level Comscore U.S. desktop clickstream data (October 2024–July 2025) to study how AI search changes web traffic allocation. It makes three contributions: (1) a descriptive comparison showing ChatGPT produces clean outbound referrals in only 5.2% of conversation sessions versus 31.1% of Google queries, with the gap persisting within household-week; (2) an analysis of the composition of ChatGPT's residual referral traffic, showing it skews toward specialized, non-ad-supported destinations; and (3) a stacked difference-in-differences design exploiting three ChatGPT Search access expansions (October 31, 2024; December 16, 2024; February 5, 2025) estimating that wider access reduces traditional search queries by 9.4% on average and 17.0% after twenty weeks, with losses concentrated in informational categories. The paper is transparent about the descriptive nature of the referral-ratio comparison and about the gap between its preferred displacement estimate and a cleaner within-adopter comparison.","tokens_in":53223,"tokens_out":2448,"duration_ms":192827,"significance":"The paper addresses a timely and economically important question: whether AI search structurally reallocates web attention away from the routed-visit economy. Its strengths include a within-household-week comparison design that controls for persistent household differences and common calendar shocks; a stacked DiD with externally timed treatment (access expansions) rather than endogenous adoption; formal parallel-trends tests that pass (joint Wald p=0.161 and p=0.743); ACS reweighting for population representativeness; a blind human validation of the domain classifier (Cohen's κ = 0.76 content type, 0.82 monetization); extensive robustness sweeps across session gaps, window lengths, contamination rules, dominance thresholds, classifier confidence, and control definitions; reconciliation with industry benchmarks; and an estimand-ladder framework (Lundberg et al. 2021) linking theoretical target to regression. The L-vs-A comparison that differences out selection into ChatGPT use is a valuable falsifiability check. The robots.txt analysis showing that opt-out does not bind the referral margin is a useful supplementary finding.","major_comments":[{"comment":"Section 3.3, Table 4: The gap between the preferred three-event estimate (−17.0% at w≥20) and the cleaner December-only L-vs-A comparison (−8.2% at w≥20) is substantial—roughly a factor of two. The paper attributes this to residual selection in the Nitt control and reports both estimates, which is commendable. However, the abstract and headline claims feature only the preferred (larger) estimate. Given that the L-vs-A design is explicitly described as the 'cleanest single comparison' (Section 3.3) and differences out the selection concern the paper itself raises, the authors should clarify in the main text why the preferred estimate is the primary headline figure rather than the more conservative L-vs-A estimate, or at minimum present both estimates with equal prominence in the abstract. This is load-bearing because the magnitude of displacement is central to the paper's economic-bargain","section":null},{"comment":"Section 3.1 and the economic-bargain framing: The paper acknowledges that the 5.2% vs 31.1% referral-ratio comparison 'does not separate differences in the tasks brought to each intermediary from differences created by the intermediaries' (Section 3.1). The surrounding-context fixed effect check (Table OA.2.4) moves the coefficient from −0.302 to −0.300, which the paper interprets as evidence that broad task context does not explain the gap. However, as the paper itself notes, this control captures only the dominant content category of nearby browsing—a coarse proxy that cannot distinguish, e.g., a quick factual lookup from a multi-step research task within the same 'reference/knowledge' bucket. The +23.1pp 'solo' share for ChatGPT is consistent with users bringing harder, more self-contained tasks to ChatGPT. The paper frames the result correctly as descriptive, but the abstract and the","section":null},{"comment":"Section 3.3, Figure 8: The category-level displacement results (e.g., −32.8% for academic research, −26.5% for reference/knowledge) are described as 'descriptive heterogeneity within the preferred design, not a separately identified mechanism.' This is appropriate, but the Discussion (Section 4) leans on these category-level patterns to connect retention inside ChatGPT to downstream losses in routed traffic. The connection between the descriptive referral-ratio findings and the causal displacement estimates is suggestive but not formally tested. The authors should clarify that the category-level displacement results are correlational with the intent patterns, not evidence that retention causes displacement within those categories.","section":null}],"minor_comments":[{"comment":"The abstract states 'ChatGPT produces outbound clicks in only 5.2% of conversation sessions' without specifying the denominator unit (conversation session vs. query). Adding 'of conversation sessions' would improve precision.","section":null},{"comment":"Figure 6: The repeated domain labels (e.g., 'python.org' appearing 17 times) make the figure difficult to read. Consider using point markers without text labels, or labeling only the most extreme domains.","section":null},{"comment":"Section OA.1.3, footnote 5: The f/conversation endpoint is described as 'not separately documented in public reverse-engineering sources.' The paper should note the risk that this endpoint's behavior may change, potentially affecting the July 2025 message-level referral ratio (Table OA.2.2).","section":null},{"comment":"Table OA.6.4: The $150–200K income cell has a post-rake SMD of 0.016, which the paper notes is the binding constraint on ESS. It would help to state the effective sample size (NESS/N=0.81) in the main text rather than only in the table notes, as it bears on the precision of the displacement estimates.","section":null},{"comment":"The paper uses 'Nitt' and 'Nitt' interchangeably (e.g., Section 2 vs. Appendix OA.6.1). Standardize the notation.","section":null},{"comment":"Section 3.2: The statement 'ChatGPT's referral ratio is highest in developer/technical contexts (13.4%)' could note that this is still well below Google's per-query ratio in the same context (68.1% from Table OA.4.1), to give readers a sense of the within-context gap.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a careful empirical study that is transparent about its limitations. The skeptic's concern about task-mix confounding the referral-ratio comparison is valid but the paper acknowledges it explicitly and frames the result as descriptive; the concern does not undermine the causal displacement estimate, which rests on separate identification. The gap between the preferred and L-vs-A displacement estimates is the most substantive issue, but the paper reports both and discusses the difference. I recommend minor revision with attention to framing balance in the abstract and main text. The paper fits the journal's scope well."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. All three major comments are well-taken and will be addressed in the revised manuscript. Two require revisions to the abstract and framing (Comments 1 and 2); one requires a clarifying statement in the Discussion (Comment 3). No standing objections remain.","responses":[{"response":"The referee is correct that the abstract features only the preferred (larger) estimate and that this choice is load-bearing for the paper's economic-bargain framing. We will revise the abstract to present both estimates with equal prominence. Specifically, the abstract will report that wider access cuts search use by 9.4% on average (17.0% after twenty weeks) under the preferred three-event design, and note that the cleaner December-only L-vs-A comparison—which differences out selection into ChatGPT use—yields a smaller but directionally consistent 8.2% decline after twenty weeks. We will also add a brief sentence in Section 3.3 explaining why the preferred design serves as the primary specification: it pools three externally timed access expansions, providing substantially greater statistical power and a longer common-support window, while the L-vs-A comparison trades power for cleaner identification. The paper already reports both estimates transparently in Table 4 and discusses the selection concern; the revision ensures the abstract and headline framing reflect that transparency.","revision_made":"yes","referee_comment":"Section 3.3, Table 4: The gap between the preferred three-event estimate (−17.0% at w≥20) and the cleaner December-only L-vs-A comparison (−8.2% at w≥20) is substantial—roughly a factor of two. The paper attributes this to residual selection in the Nitt control and reports both estimates, which is commendable. However, the abstract and headline claims feature only the preferred (larger) estimate. Given that the L-vs-A design is explicitly described as the 'cleanest single comparison' (Section 3.3) and differences out the selection concern the paper itself raises, the authors should clarify in the main text why the preferred estimate is the primary headline figure rather than the more conservative L-vs-A estimate, or at minimum present both estimates with equal prominence in the abstract."},{"response":"The referee's point is well taken. The surrounding-context fixed effect is indeed a coarse proxy: it controls for the dominant content category of nearby browsing but cannot distinguish task complexity within a category. The +23.1pp solo share is consistent with the interpretation that users bring harder, more self-contained tasks to ChatGPT, and the current framing does not adequately flag this as a specific, uncontrolled form of task selection. We will make two changes. First, we will add a sentence to Section 3.1 explicitly acknowledging that the solo-share pattern is consistent with task-complexity selection that the surrounding-context control cannot address, and that this is a specific limitation of the descriptive comparison. Second, we will adjust the abstract to qualify the referral-ratio comparison as descriptive and note that task-selection differences—including the possibility that users bring more self-contained tasks to ChatGPT—cannot be fully ruled out. The paper already states that the result 'does not separate differences in the tasks brought to each intermediary from differences created by the intermediaries'; the revision makes this caveat more specific and ensures the abstract carries it.","revision_made":"yes","referee_comment":"Section 3.1 and the economic-bargain framing: The paper acknowledges that the 5.2% vs 31.1% referral-ratio comparison 'does not separate differences in the tasks brought to each intermediary from differences created by the intermediaries' (Section 3.1). The surrounding-context fixed effect check (Table OA.2.4) moves the coefficient from −0.302 to −0.300, which the paper interprets as evidence that broad task context does not explain the gap. However, as the paper itself notes, this control captures only the dominant content category of nearby browsing—a coarse proxy that cannot distinguish, e.g., a quick factual lookup from a multi-step research task within the same 'reference/knowledge' bucket. The +23.1pp 'solo' share for ChatGPT is consistent with users bringing harder, more self-contained tasks to ChatGPT. The paper frames the result correctly as descriptive, but the abstract and the"},{"response":"The referee is correct. The Discussion (Section 4) draws a suggestive connection between the descriptive finding that informational tasks are prevalent in ChatGPT use and often retained, and the causal finding that informational categories show the largest search-referral losses. This connection is not formally tested: we do not estimate a mediation or mechanism model linking retention to displacement within categories. The category-level displacement results are heterogeneity within the preferred DiD design, and their alignment with the intent patterns is correlational. We will add an explicit clarifying statement to Section 4 stating that the category-level patterns are consistent with—but do not formally identify—a mechanism by which retention inside ChatGPT causes downstream referral losses in the same categories. We will also adjust the relevant sentence in Section 3.3 to make clear that the heterogeneity is correlational with the intent patterns rather than evidence of a causal link.","revision_made":"yes","referee_comment":"Section 3.3, Figure 8: The category-level displacement results (e.g., −32.8% for academic research, −26.5% for reference/knowledge) are described as 'descriptive heterogeneity within the preferred design, not a separately identified mechanism.' This is appropriate, but the Discussion (Section 4) leans on these category-level patterns to connect retention inside ChatGPT to downstream losses in routed traffic. The connection between the descriptive referral-ratio findings and the causal displacement estimates is suggestive but not formally tested. The authors should clarify that the category-level displacement results are correlational with the intent patterns, not evidence that retention causes displacement within those categories."}],"tokens_in":53105,"tokens_out":1281,"duration_ms":170891,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Two things to know up front: this paper provides the first user-side measurement of how often ChatGPT absorbs information-seeking sessions without sending users onward (5.2% referral rate vs. 31.1% for Google), and it uses three externally-timed ChatGPT Search access expansions in a stacked DiD to estimate 9.4% displacement of traditional search queries (17.0% after twenty weeks). Both findings are new and both matter for anyone thinking about web economics, platform regulation, or content production incentives. The paper is empirically careful. The within-household-week comparison with household and week fixed effects is the right design for the descriptive referral-ratio gap, and it barely moves from the raw difference, which tells you the gap isn't a composition artifact of who uses each platform. The stacked DiD uses externally announced treatment dates, passes parallel-trends tests (joint Wald p=0.161 for the three-shock design, p=0.743 for the December comparison), and is corroborated by a Callaway–Sant'Anna estimator and an alternative within-adopter comparison. The destination-composition analysis—showing ChatGPT referrals tilt toward reference, academic, and developer sites and away from ad-supported destinations—is well-executed, with a blind human-validated classifier (κ=0.76 content type, 0.82 monetization) and robustness across confidence thresholds. The concentration analysis distinguishing aggregate vs. within-household patterns is a nice piece of work. The stress-test concern about task-mix confounding the 5.2% vs. 31.1% gap is legitimate but partially addressed. The surrounding-context fixed effect check (Table OA.2.4) moves the coefficient from −0.302 to −0.300, which is reassuring but coarse—it controls for the dominant browsing category, not the actual task. The paper acknowledges this limitation directly, and the headline framing does lean on the gap as structural intermediary behavior more than the evidence strictly supports. That said, the gap is large enough (6x difference) that even substantial task-mix confounding would leave a meaningful difference. The causal displacement estimate is independent of this concern and rests on separate identification. The main soft spot is the gap between the preferred estimate (17.0% at twenty weeks) and the cleaner L-vs-A comparison (8.2%), which the authors attribute to selection on AI adoption in the Nitt control. This is probably right, but it means the headline 9.4% figure sits at the upper end of what the evidence supports—the true effect is likely between 8% and 17%, depending on how much selection remains. Desktop-only scope is a real limitation but doesn't invalidate the findings. This paper is for empirical researchers in digital economics, platform regulation, and web measurement. It deserves a serious referee who can evaluate the DiD design and the clickstream measurement choices. Recommend accept into peer review.","headline":"Solid empirical paper documenting AI search's effect on web traffic routing; deserves serious review","tokens_in":53810,"tokens_out":660,"would_cite":true,"duration_ms":118428,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"AI search sends outbound clicks in only 5.2% of sessions vs 31.1% for Google","keywords":["AI search","web traffic","digital intermediation","attention economy","search advertising","clickstream","referral ratio","ChatGPT"],"falsifier":"If traditional search queries did not decline after ChatGPT Search access expanded—i.e., if the event-study coefficients were zero or positive post-treatment—the displacement claim would fail. Alternatively, if the referral-ratio gap between ChatGPT and Google disappeared under a different session definition or attribution window, the absorption claim would weaken substantially.","tokens_in":53282,"feed_emoji":"🔗","tokens_out":1389,"duration_ms":111338,"temperature":0.7,"pith_summary":"This paper argues that AI search is structurally changing how attention flows on the web. For two decades, search engines routed users from queries to destination websites, and those visits were the economic lifeblood of online content—creating ad impressions, subscription conversions, and audience relationships. The authors use URL-level U.S. desktop clickstream data to show that ChatGPT resolves most information needs inside its own interface: only 5.2% of ChatGPT conversation sessions produce a clean outbound click, compared with 31.1% of Google queries. The residual traffic that does exit ChatGPT is not a scaled-down version of Google's referral stream—it skews toward specialized, non-profit, and developer destinations and away from ad-supported sites. Using three natural experiments in which ChatGPT Search access expanded to new user groups, the authors estimate that wider access reduces traditional search queries by 9.4% on average and 17.0% after twenty weeks, with the largest losses concentrated in informational content categories like academic research, reference, and news. The paper frames this as a shift from routed visits to residual referrals: the intermediary now answers the question itself rather than sending the user onward, which weakens the implicit bargain that has linked search, traffic, and content production on the open web.","feed_headline":"AI search sends outbound clicks in only 5.2% of sessions vs 31.1% for Google","feed_subtitle":"Wider ChatGPT Search access cuts traditional search queries by 9.4%, suggesting some residual traffic may bypass the open web entirely.","key_machinery":"The paper uses three empirical components. First, a within-household, within-week comparison of ChatGPT and Google referral ratios using Comscore U.S. desktop clickstream data (October 2024–July 2025), with household and week fixed effects to absorb persistent differences. Second, a stacked difference-in-differences design exploiting three ChatGPT Search access expansions (paid subscribers October 31 2024, free logged-in users December 16 2024, anonymous browsers February 5 2025), with each treated cohort matched to a reweighted control of households with no pre-expansion ChatGPT or Claude activity. Third, a domain classification system (4,266 destinations labeled by content type and monetiz","core_discovery":"The central object the paper identifies is the referral ratio—the share of information-seeking occasions at an intermediary that produce at least one clean outbound click to a third-party website. By measuring this symmetrically for ChatGPT conversation sessions (5.2%) and Google queries (31.1%), the authors establish that AI search absorbs roughly six times more information needs internally than traditional search does. This is not a welfare claim about whether users are better served; it is a traffic-allocation claim about where observable attention ends. The paper then connects this absorption to downstream substitution: when households gain ChatGPT Search access, their traditional search","pith_inferences":["If AI search absorption rates hold or increase as models improve, the web's traffic-based attribution system becomes increasingly incomplete: a growing share of information needs are satisfied without any observable visit that a website can count, monetize, or convert. This creates a measurement problem for the entire digital advertising and publishing ecosystem, not just for search engines.","The gap between the preferred estimate (17.0% displacement at 20 weeks) and the cleaner within-adopter comparison (8.2%) suggests the true causal effect may lie between these bounds. If so, the paper's headline displacement figure may overstate the effect for policy purposes while still confirming the direction.","The finding that different households reach different specialty destinations through ChatGPT (aggregate dispersion) while individual households concentrate on few destinations implies that AI search may fragment web audiences in ways that undermine the network effects large destination sites have relied on, without necessarily benefiting smaller sites in aggregate.","If content producers respond to declining routed traffic by reducing investment in informational content—the categories most affected—the quality of information available to AI search systems themselves may degrade, creating a feedback loop the paper identifies as its central open question."],"forward_implications":["If the referral-ratio gap persists, ad-supported informational websites that depend on routed search traffic face a structural decline in the visits they can monetize, even if AI search continues to use their content as source material.","The finding that residual ChatGPT referrals avoid ad-supported sites by 27.6 percentage points suggests that the websites most dependent on search-driven attention are precisely those most bypassed by AI search's smaller click-out stream.","Search advertising inventory contracts as traditional queries fall by 9–17%, which would affect pricing and channel allocation in search-ad markets before any equilibrium adjustment in bids or budgets.","The category-specific pattern—academic research referrals down 32.8%, reference/knowledge down 26.5%—identifies which content producers face the most acute exposure in licensing and attribution negotiations with AI intermediaries.","The robots.txt finding that 79% of classified domains block at least one AI training crawler, yet blocking does not reduce runtime referrals, means that existing opt-out mechanisms do not protect content producers from the traffic reallocation the paper documents."],"fun_headline_variants":["ChatGPT absorbs six times more queries internally than Google search","AI search breaks the referral economy: only 5.2% of sessions click out","Wider AI search access cuts traditional search queries by 9.4%","AI search traffic bypasses ad-supported sites for specialized destinations","Referral ratio gap: 5.2% for ChatGPT vs 31.1% for Google search"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The parallel-trends assumption: that households gaining ChatGPT Search access and households without it would have followed similar search-query trajectories absent the expansion. The preferred control group consists of households with no pre-expansion ChatGPT or Claude activity, who may differ from treated households in unobserved ways related to search trends. The gap between the preferred estimate (17.0% long-run displacement) and a cleaner comparison between two groups of","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT absorbs six times more queries internally than Google search","AI search breaks the referral economy: only 5.2% of sessions click out","Wider AI search access cuts traditional search queries by 9.4%","AI search traffic bypasses ad-supported sites for specialized destinations","Referral ratio gap: 5.2% for ChatGPT vs 31.1% for Google search"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":618,"prompt_tokens":516,"completion_tokens":102,"prompt_tokens_details":null},"tokens_in":516,"tokens_out":102,"duration_ms":18075,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T03:33:15.323259+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If traditional search queries did not decline after ChatGPT Search access expanded—i.e., if the event-study coefficients were zero or positive post-treatment—the displacement claim would fail. Alternatively, if the referral-ratio gap between ChatGPT and Google disappeared under a different session definition or attribution window, the absorption claim would weaken substantially.","supporting_citations":[],"review_version":1}