{"id":"fa5ecf31-ef27-4a61-9282-c65dbfce8ec2","arxiv_id":"2605.14021","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Google AI Overviews activate on 13.7% of queries overall and 64.7% of questions, cite more credible sources than standard results but omit key information in 11% of claims, and suppress clicks on over half of cited pages that carry ads.","lead":"This paper measures Google AI Overviews by running 55,393 trending queries over 40 days and tracking activation rates, cited sources, claim accuracy, and ad revenue effects. The results show how AI summaries are reshaping what users see and how publishers are affected.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Atomic claim decomposition into 98k units lacks reported validation, risking bias in the 11% unsupported fidelity rate","rationale":"The reader's weakest assumption correctly flags both query representativeness and decomposition reliability as central risks. The decomposition step is more directly load-bearing for the fidelity finding than query selection, because the 11% unsupported rate and its independence from source quality rest entirely on that measurement. Full methods would be needed to confirm or refute; until then the claim remains provisional.","tokens_in":1824,"tokens_out":320,"duration_ms":18815,"concrete_test":"Select a random 200-response subsample; have two independent human annotators (blind to original decomposition) re-extract atomic claims and label support status; compute Cohen's kappa on claim boundaries and support labels. If kappa < 0.75 or unsupported rate shifts >3 percentage points, the 11% figure and independence claim are unreliable.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The third finding decomposes AIO responses into 98,020 atomic claims and reports 11.0% unsupported (omission dominant), with source quality and fidelity largely independent. This decomposition step is load-bearing: if claim extraction or support judgments are performed via unvalidated LLM prompts or single-annotator rules, systematic over- or under-counting of omissions could inflate the unsupported rate and create spurious independence from source credibility. The abstract provides no inter-rater reliability, prompt details, or human-validation subsample, leaving the metric vulnerable to annotator or model bias.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports results from a 40-day longitudinal measurement study issuing 55,393 trending queries across 19 topical categories to Google. It finds AIO activation at 13.7% overall (64.7% for question-form queries, lower for politically sensitive topics), AIO-cited domains more credible than co-displayed first-page results yet ~30% absent from those results, 11.0% of 98,020 atomic claims unsupported by cited pages (omission dominant) with source quality and fidelity largely independent, and >50% of AIO-cited pages carrying display advertising.","tokens_in":1931,"tokens_out":383,"duration_ms":26120,"significance":"If the measurements are reliable, the study supplies large-scale empirical evidence on activation rates, source selection distinct from ranking, claim fidelity, and publisher revenue displacement in generative search summaries. These observations bear directly on epistemic security, information ecosystem shifts, and advertising economics, providing a useful baseline for future work.","major_comments":[{"comment":"Abstract (third finding) and associated methods: the decomposition of responses into 98,020 atomic claims and the 11.0% unsupported rate are load-bearing for the fidelity and independence claims, yet the manuscript provides no description of atomic-claim extraction criteria, support-judgment rules, inter-rater reliability, or human-validation subsample. Without these details the reported percentage cannot be verified and may embed systematic bias.","section":"Abstract / claim-fidelity section"}],"minor_comments":[{"comment":"The date window March 13–April 21, 2026 appears to lie in the future relative to the arXiv posting; confirm the correct interval or clarify the study timeline.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their careful review and constructive feedback on our manuscript. We address the single major comment below and will revise the manuscript to incorporate the requested methodological details.","responses":[{"response":"We agree that the current version of the manuscript does not include a sufficiently detailed description of the atomic-claim extraction process, the rules used to judge support or omission, inter-rater reliability statistics, or the human-validation subsample. This information is necessary for reproducibility and to allow readers to evaluate potential bias. In the revised manuscript we will add a dedicated subsection to the Methods section that specifies: (1) the annotation guidelines and decision rules for decomposing AIO responses into atomic claims, (2) the criteria for classifying a claim as supported (direct quotation, close paraphrase, or logical entailment from the cited page), (3) inter-rater agreement results (including Cohen’s kappa) obtained from a pilot study on a random subsample of claims, and (4) the size, selection procedure, and validation protocol for the human-reviewed subsample. These additions will be placed immediately before the presentation of the 11.0 % unsupported rate so that the fidelity and independence findings can be properly assessed. No changes to the reported percentages themselves are required.","revision_made":"yes","referee_comment":"[Abstract / claim-fidelity section] Abstract (third finding) and associated methods: the decomposition of responses into 98,020 atomic claims and the 11.0% unsupported rate are load-bearing for the fidelity and independence claims, yet the manuscript provides no description of atomic-claim extraction criteria, support-judgment rules, inter-rater reliability, or human-validation subsample. Without these details the reported percentage cannot be verified and may embed systematic bias."}],"tokens_in":1413,"tokens_out":376,"duration_ms":32708,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper gives the first big empirical snapshot of Google AI Overviews. They ran 55,393 trending queries over 40 days and tracked when the feature activates, which sources it picks, how well its claims hold up, and what that means for publishers. Activation lands at 13.7 percent overall and jumps to 64.7 percent on question queries, while political topics see lower rates. Cited domains are more credible than standard results, yet nearly 30 percent of them do not appear in the first page at all. They split responses into 98,020 atomic claims and report 11 percent unsupported, mostly through omission, with fidelity largely independent of source quality. Over half the cited pages carry display ads, so publishers lose clicks when the summary takes over. The work stays observational and covers several angles at once without extra modeling. The scale and the mix of questions make the numbers useful for anyone following how generative search changes information flow. The soft spot is the claim decomposition. Breaking text into 98k atomic claims and judging support against the cited pages is the load-bearing step for the 11 percent figure, yet the text gives no details on extraction rules, inter-rater checks, or a human validation sample. Without those, the unsupported rate could move with different choices. The trending-query sample is reasonable for this purpose but does not stand in for all user searches. This is for researchers on search engines, AI information systems, and media economics. It deserves peer review because the empirical reach is substantial and the questions are current, even if the methods section needs more transparency in revision.","headline":"First large-scale numbers on AI Overviews activation and claim support, but the atomic claim breakdown lacks reported validation.","tokens_in":2409,"tokens_out":387,"would_cite":true,"duration_ms":31281,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Empirical audit of generative search fidelity and economics; no structural overlap with RS forcing chain","alignment":"orthogonal","rationale":"The paper's core machinery is large-scale query crawling, atomic-claim decomposition (98k claims), LLM-based verification against cited pages, PC1 credibility scoring, and ad-prevalence counting. These are standard measurement and NLP techniques with no connection to J-cost, φ-ladder, 8-tick periodicity, or the distinction-to-spacetime forcing theorems. No RS module (e.g., Cost/FunctionalEquation, Foundation/RealityFromDistinction, AlexanderDuality) is invoked or paralleled.","tokens_in":59767,"confidence":"high","tokens_out":150,"duration_ms":7578,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Google AI Overviews activate on 13.7% of searches, with 11% of their claims unsupported by cited pages and most of those pages carrying ads that suppress publisher revenue.","keywords":["Google AI Overviews","search measurement","generative AI","claim fidelity","source quality","publisher revenue","query activation","information ecosystem"],"falsifier":"Re-running the same queries with a different query sample or having human raters verify the support status of a random subset of the 98,020 claims would show whether the reported activation and unsupported rates hold.","tokens_in":2714,"feed_emoji":"📊","tokens_out":568,"duration_ms":29363,"temperature":0.7,"pith_summary":"The paper measures Google AI Overviews across 55,393 trending queries over 40 days to track how often they appear and how well they perform. Activation reaches 13.7% overall but jumps to 64.7% for question queries while staying low on political topics. Cited sources prove more credible than standard results yet nearly 30% are absent from those results, showing a separate selection process. Breaking answers into 98,020 atomic claims reveals that 11% lack support from the cited pages, mostly through omissions, and source quality does not predict fidelity. Over half the cited pages include display advertising, so publishers lose clicks and revenue when the AI summary replaces the original link.","feed_headline":"AI Overviews appear in 14% of searches with 11% unsupported claims","feed_subtitle":"Measurement of 55,000 queries shows AIOs select credible sources yet omit facts and reduce ad revenue for cited publishers.","key_machinery":"Large-scale longitudinal query measurement combined with atomic claim decomposition to quantify activation rates, source credibility differences, unsupported claims, and advertising presence on cited pages.","core_discovery":"Issuing 55,393 trending queries across 19 categories shows AIO activation at 13.7% overall and 64.7% for question-form queries, with lower rates on politically sensitive topics. AIO-cited domains are more credible than co-displayed results but nearly 30% do not appear in those results. Of 98,020 decomposed atomic claims, 11.0% are unsupported by the cited pages, with omission as the main failure mode, and fidelity is independent of source quality. Well over half of AIO-cited pages carry display advertising, so publishers lose revenue when AIOs suppress clicks while Google's sponsored ads remain.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["AI Overviews trigger for 14% of searches","11% of AI Overview claims lack source support","AI Overviews select credible sources missed by ranking","Over half cited AI pages display ads hurting publishers","Question queries activate AI Overviews 65% of time"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The 55,393 trending queries represent typical user searches and that breaking responses into atomic claims can be done reliably without systematic bias in measuring unsupported content.","fun_headline_variants_meta":{"raw":{"variants":["AI Overviews trigger for 14% of searches","11% of AI Overview claims lack source support","AI Overviews select credible sources missed by ranking","Over half cited AI pages display ads hurting publishers","Question queries activate AI Overviews 65% of time"]},"model":"grok-4.3","cost_usd":0.004412,"raw_usage":{"total_tokens":2189,"prompt_tokens":795,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":44115500,"prompt_tokens_details":{"text_tokens":795,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1322,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":795,"tokens_out":72,"duration_ms":18580,"temperature":1.0,"reasoning_tokens":1322,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-15T02:09:49.978034+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Re-running the same queries with a different query sample or having human raters verify the support status of a random subset of the 98,020 claims would show whether the reported activation and unsupported rates hold.","supporting_citations":[],"review_version":1}