{"id":"9d34dd87-9d2c-406a-bb51-f18d52e11fc7","arxiv_id":"2607.14035","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A critical review of GEO research concludes that already-retrieved content can improve citation and use, but no tested technique reliably raises organic discoverability or downstream traffic across engines.","lead":"This survey of 45 studies on Generative Engine Optimization finds that headline gains only appear once a source is already inside an engine's context, and that no technique has shown a stable cross-platform boost to organic discovery or traffic. It offers a staged pipeline model and a visibility vector to separate being cited from being found.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The absence claim's external force depends on corpus completeness, which §2.4 concedes is unauditable; a missed positive longitudinal multi-engine study would flip the central conclusion.","rationale":"The paper's main contribution is a bounded negative: within 45 reviewed studies, no technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior. I looked for internal inconsistencies between the formal model (§3), the evidence hierarchy (§6), and the synthesis (§10). The paper is consistent: it separates conditional from total effects, distinguishes metrics across a visibility vector, and explicitly limits its negatives to the corpus. The one place where the argument is exposed is the transition from 'within this corpus' to the field-level conclusion that GEO's scientifically warranted promise is conditional on retrieval. That transition requires the corpus to represent the relevant literature. §2.4 and §14 both concede that the search did not retain database-specific hit counts or a complete exclusion ledger, so completeness cannot be audited. If a qualifying study — randomized or strong quasi-experimental, multi-engine, longitudinal — exists outside Table 8, the headline negative would flip. The preprint status of SAGEO Arena and the traffic study affects confidence, but it is not the load-bearing point: even if those preprints were withdrawn, the absence of a positive demonstration would remain. The concrete test is to independently reproduce the search with a full ledger and specifically hunt for counterexamples. This is exactly the condition the reader attached; therefore I keep the verdict unchanged.","tokens_in":20224,"tokens_out":5931,"duration_ms":62595,"concrete_test":"Commission an independent team to reproduce the §2.2 search across the six named databases (arXiv, ACM DL, ACL Anthology, NeurIPS, PMLR, OpenReview) for November 16, 2023–July 14, 2026, using the listed query families and a preregistered inclusion form, while retaining per-database hit counts and exclusion reasons. Have them specifically flag any candidate reporting a causal estimate of an intervention on organic retrieval or downstream traffic measured on more than one commercial engine at more than one time point. If any such study is absent from Table 8, the central absence claim must be re-scoped; if none surfaces, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline negative — 'no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior' — is an existential claim over the surveyed literature. Its external force depends on the 45-study corpus being the relevant population. §2.4 concedes that 'the original search did not retain database-specific hit counts or a complete exclusion ledger,' and §14 marks those files unavailable. A missed or excluded study that does report a multi-engine longitudinal causal gain on retrieval or traffic would directly flip the central conclusion. The survey honestly bounds the claim to 'within this corpus,' which protects it against internal falsification, but the paper's broader intellectual contribution — that GEO's scientifically warranted promise is limited to conditioning on retrieval — implicitly generalizes beyond the corpus. The two strongest negative building blocks (SAGEO Arena's end-to-end degradation, §7.4, and the traffic placebo result, §8.5) are preprints, so they are secondary risks; the primary load-bearing risk is selection completeness, not any single finding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a critical scoping review of 45 studies on Generative Engine Optimization (GEO) published between November 2023 and July 2026. It formalizes GEO as a multistage, partially observable pipeline (activation, retrieval, reranking, generation, citation, absorption, fidelity, downstream behavior), introduces a visibility vector and an evidence hierarchy, and proposes a reproducibility protocol. The paper's central claims are that the widely cited 'up to 40%' result of Aggarwal et al. is a within-context relative gain in position-adjusted word count, that the surveyed literature supports causal effects only for already-retrieved content, and that no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior. It also synthesizes commercial audits, manipulation/defense work, and governance considerations.","tokens_in":20333,"tokens_out":8323,"duration_ms":86368,"significance":"The paper's main strength is its careful epistemic calibration. It correctly separates peer-reviewed studies from preprints (Section 2.3), distinguishes conditional from total effects (Section 3.3), and explicitly bounds its absence claims to the reviewed corpus (Abstract; Section 2.4; Section 15). The reading of the foundational paper's 41% figure as a within-context relative pawc gain is accurate and well documented (Table 1, Section 4.1). The proposed visibility vector, evidence hierarchy (Appendix A), and factorial measurement protocol are useful contributions that can discipline future work in this rapidly growing area. If the synthesis is accepted, it materially corrects the overreading of GEO's promise and redirects attention to retrieval-stage and downstream-outcome measurement. The principal limitation is the unaudited completeness of the corpus: because the negative claim is explicitly corpus-relative, the issue is a constraint on external generalization rather than an internal inconsistency.","major_comments":[],"minor_comments":[{"comment":"The manuscript is careful to bound the negative claim with 'Within this corpus' in the Abstract and Section 15, but Section 10's opening synthesis ('The most important conclusion concerns scope...') and the confidence ratings in Table 5 do not repeat that qualifier. Since Section 2.4 concedes that database-specific hit counts and the complete exclusion ledger were not retained, a reader could over-generalize the absence claim. Please add 'within the reviewed corpus' to Table 5's header or a footnote, and to the first sentence of Section 10, so the epistemic scope is uniform throughout.","section":"Abstract, Section 10, Table 5"},{"comment":"The proposed hierarchical model has an indexing inconsistency: the outcome is indexed by i, q, e, t, r, but the treatment indicator is T_i and the random effect is b_s. Since the source is the treatment unit, the equation should use T_s (or clarify what i denotes). Please also state whether b_q, b_e, b_t, and b_s are crossed or nested, and define the cluster structure explicitly.","section":"Section 11.3, Equation (8)"},{"comment":"The recommendation of 'seven to eight repetitions as a starting point' is appropriately hedged, but 'the appropriate practice is sequential precision analysis' needs an operational stopping rule. Specify a concrete criterion (e.g., continue until the half-width of the confidence interval is below a prespecified threshold) or give an example so that readers can implement it.","section":"Section 6.2"},{"comment":"The entry for Nimase et al. 2026 ('GEO-Bench') shares a name with the original benchmark used by Aggarwal et al. 2024. This is a source of potential confusion. Please add a disambiguating note, for example 'not the original GEO-Bench benchmark,' in the table or in the text.","section":"Appendix B, Table 8"}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid, carefully bounded scoping review. The main residual risk is corpus completeness for the negative claim, but the authors already handle this by framing the claim as 'within the reviewed corpus.' I do not see a need for further experiments or a new search; the requested changes are local and presentational. No concerns about novelty or fit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. This is a survey of GEO that actually does the hard work of reconciling the field's founding 40% figure with the later nulls and mixed results. The key move is formal: the multistage pipeline (Eqs. 1–4) and the visibility vector (Eq. 5) let the author separate retrieval (D_s), citation (C_s), absorption (H_s), and traffic (B_s), so the headline claim — 'already-retrieved content can causally alter its citation, but no technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior' — is stated precisely and defended carefully. That is a real contribution, not just a taxonomy. The reading of the foundational paper is accurate, including the critical point that the 41% gain is a within-context, position-weighted share, not a discovery or traffic effect. The evidence hierarchy in Table 7 is sensible and the protocol in §11 is genuinely useful for anyone designing GEO experiments. It is also honest: §2.4 and §14 flag that the corpus search did not retain database-specific hit counts or the full exclusion ledger. That is the real weakness, and it is disclosed. The central absence claim is carefully bounded to the reviewed corpus, but the broader intellectual contribution — that GEO's warranted promise is limited to conditioning on retrieval — implicitly generalizes. A missed positive longitudinal multi-engine study would crack that. Also, several of the strongest negative building blocks (SAGEO Arena §7.4, the traffic placebo §8.5) come from preprints, so those findings carry preprint-level weight. That said, the author handles this appropriately by distinguishing publication status throughout and not over-claiming. The stress-test note is fair but slightly overstated; the paper's conclusion is explicitly scoped to the corpus, so the internal logic holds. The selection completeness is a real limitation, but it is the kind of limitation that can be addressed by releasing whatever audit trail exists and re-verifying preprint-dependent claims. I disagree with any implication that this is sloppy or circular. It is a disciplined synthesis with a clearly stated negative claim. Soft spots are proportionate: the missing audit trail is moderate, not fatal. Who is this for? Anyone working on GEO, AI search, or content optimization, and skeptical practitioners who need an evidence-grounded map. It deserves a serious referee. I would accept it for peer review and let the review process push on the selection audit. I would also cite it. Bring it to the reading group.","headline":"A genuinely useful critical survey: it separates the conditional citation effect from the discovery/traffic claims in GEO, and the main soft spot is the unauditable corpus selection, not the synthesis itself.","tokens_in":20914,"tokens_out":622,"would_cite":true,"duration_ms":8276,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that GEO is not a single ranking task but a stochastic, partially observable pipeline, and that the field's most cited result—'up to 40% visibility gains'—is conditional on a document already being retrieved, not a general","keywords":["generative engine optimization","GEO","AI search visibility","citation analysis","algorithmic auditing","retrieval-augmented generation","causal measurement","literature survey"],"falsifier":"A single preregistered randomized field experiment would falsify the paper's central absence claim if it showed: a specific, well-defined content intervention (e.g., adding verifiable structured data) that, across at least two major generative engines and over a period of months, durably increases organic retrieval probability (not just citation among already-retrieved sources) in a treated group compared to a randomized control, with a pre-specified primary metric and adequate statistical power.","tokens_in":19964,"feed_emoji":"🔍","tokens_out":3245,"duration_ms":31860,"temperature":0.7,"pith_summary":"This critical survey of 45 studies on Generative Engine Optimization (GEO) tries to establish what the field has actually proven, rather than what its promotional vocabulary suggests. The paper's central claim is that GEO operates across a multistage pipeline—from search activation, crawling, and retrieval, through reranking, generation, and citation, down to user behavior—and that treating visibility as one number obscures where interventions truly take effect. The widely cited 'up to 40%' gain from the foundational 2024 paper is real but narrowly scoped: it measures how much a document already placed in a five-source context gains in position-weighted citation share, not whether a page gets organically retrieved or whether it drives clicks. Across the reviewed corpus, the paper finds strong evidence that already-retrieved content can causally alter its citation or use, but no technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior. A sympathetic reader cares because this reframes GEO's promise: instead of a recipe for ranking highly in ChatGPT, the defendable science is about conditioning on retrieval, and the paper offers a visibility vector, an evidence hierarchy, and a measurement protocol to make that distinction operational.","feed_headline":"Survey: no technique yet lifts organic visibility in AI search","feed_subtitle":"A 45-study review finds the field's 40% gains are conditional on retrieval; discovery and traffic remain unproven.","key_machinery":"The central machinery is the visibility vector V_s = (D_s, K_s, C_s, P_s, H_s, F_s, B_s), which separates discoverability (retrieval probability), context exposure (rank, token allocation), citation probability, prominence (position, repetition), absorption (effective contribution to the answer's facts or language), fidelity (whether attributed claims are supported), and behavioral/economic outcome (click, referral, conversion). This vector is paired with a multistage formal model of the generative engine pipeline—activation, crawling/indexing, retrieval, reranking/context allocation, generation/citation, absorption/fidelity, and user behavior—and a causal estimand τ_T(m) that distinguishes","core_discovery":"The core discovery is a scope restriction on what GEO can currently claim. The paper shows that the foundational 'up to 40%' figure derives from a simulator in which five documents are already placed in context, where one source's position-weighted word share (pawc) rises from 19.3 to 27.2 under a quotation-addition intervention—a relative gain of about 41%. It does not establish that a page will be retrieved organically, nor that it will generate traffic or conversions. The survey's synthesis of 45 studies concludes that within the reviewed corpus the evidence is narrow: already-retrieved content can causally influence an answer, including its rank, citation, or use, but no technique shows","pith_inferences":["If the survey's absence claim holds, a practical consequence is that GEO budgets should be redirected toward retrievability and content quality rather than citation-optimization tricks, because a page that is never retrieved cannot benefit from any downstream effect.","The recognition–discovery gap documented for named products (99.4% recognition vs 3.32% organic discovery for ChatGPT) implies that brand authority, third-party coverage, and entity-level representation may matter more than page-level rewrites, suggesting a network-level view of GEO rather than a page-level one.","A testable extension the paper leaves implicit: a randomized field trial across multiple engines with controlled pages, measuring organic retrieval probability (not just citation given retrieval) over several months, could directly falsify the central absence claim if a positive, stable effect emerges.","The paper's proposed evidence hierarchy and multi-stage measurement protocol could be adopted more broadly by researchers auditing algorithmic surfaces beyond GEO, offering a template for separating conditional from total effects in any black-box optimization context."],"forward_implications":["The 'up to 40%' GEO result should be read as a within-context effect, not a general promise of ranking highly in ChatGPT; it measures position-weighted attribution share for a document already placed in a fixed, five-source context.","Generic GEO heuristics—such as keyword stuffing, fluency rewrites, or formatting tricks—do not generalize; the most reproducible levers are query–document relevance and context position, which shift attention upstream toward retrieval.","Optimizing for citation can backfire on retrieval: the SAGEO Arena experiment shows that body-only rewrites reduce average top-20 presence by ~9%, top-10 presence after reranking by ~16%, and final citation by ~6%.","Commercial engines are heterogeneous and unstable: audits find low source overlap across engines, substantial run-to-run variability (daily Jaccard scores ~0.34–0.42), and persistent fidelity gaps, so visibility must be measured as a distribution over engines, dates, and paraphrases, not a point estimate.","Evidence for traffic or conversion effects is the weakest link: only one suggestive quasi-experiment (an estimated multiplier of 1.82 with a placebo p = 0.16) and one under-specified industry report claim production-level traffic lifts, falling short of causal standards."],"fun_headline_variants":["GEO survey: 40% gains don't mean organic visibility","No GEO technique proves long-term AI search lift","Critical survey: GEO evidence stops at retrieval","AI SEO review: citation gains, no discovery proof","Survey: GEO can alter answers, not organic reach"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central absence claim—that no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior—depends on the 45-study corpus being representative of the field; the paper itself notes that the original search did not retain database-specific hit counts or a complete exclusion ledger, and several key negative findings rely on preprints rather than peer-reviewed work.","fun_headline_variants_meta":{"raw":{"variants":["GEO survey: 40% gains don't mean organic visibility","No GEO technique proves long-term AI search lift","Critical survey: GEO evidence stops at retrieval","AI SEO review: citation gains, no discovery proof","Survey: GEO can alter answers, not organic reach"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0002,"raw_usage":{"total_tokens":1262,"prompt_tokens":845,"completion_tokens":417,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":355}},"tokens_in":589,"tokens_out":417,"duration_ms":4400,"temperature":1.0,"reasoning_tokens":355,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:56:22.603668+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single preregistered randomized field experiment would falsify the paper's central absence claim if it showed: a specific, well-defined content intervention (e.g., adding verifiable structured data) that, across at least two major generative engines and over a period of months, durably increases organic retrieval probability (not just citation among already-retrieved sources) in a treated group compared to a randomized control, with a pre-specified primary metric and adequate statistical power.","supporting_citations":[],"review_version":1}