Pith. sign in

REVIEW 3 major objections 4 minor 5 references

Chinese-language generative search engines exert a second, strong selection stage after citation: only 8.3% of cited-pool brands and 12.4% of contact-information items from cited sources appear in generated answers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-01 22:23 UTC pith:PG55ECVV

load-bearing objection Careful, large-scale empirical baseline for Chinese-language GEO; the 71% unmatched-contact claim is real but needs the crawl-caveat attached. the 3 major comments →

arxiv 2607.15771 v1 pith:PG55ECVV submitted 2026-07-17 cs.IR

What Do Chinese-Language Generative Search Engines Cite and Surface? A Large-Scale Empirical Study

classification cs.IR
keywords generative searchcitation behaviorentity exposurecitation absorptiongenerative engine optimizationsource attributioncross-interface consistencyunmatched exposure
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that, in Chinese-language generative search engines, being cited and being exposed in the generated answer are separate, strongly selective stages. Using 614 controlled queries across eight platform interfaces and a cleaned corpus of 160,860 citation-level records, it finds that only 8.3% of brands present in the citation pool are written into answers and only 12.4% of contact-information items detected in cited sources reach the answer text. Roughly 13% of brand exposures and 71% of contact-information exposures cannot be matched to the contemporaneous citation pool or the crawled body text, and App and Web interfaces of the same platform return substantially different source sets. A sympathetic reader would care because these regularities define where visibility is actually gained or lost in generative search, and they imply that citation counts are a poor proxy for real exposure.

Core claim

The paper's central claim is that Chinese generative search engines filter twice: they select sources for the citation list and then select again which entities from those sources appear in the final answer. The second filter is severe—91.7% of citation-pool brands are eliminated before the answer is written, and 87.6% of contact-information items present in cited sources never appear in answers. The answer text also contains entities that cannot be traced to the visible citation pool: about 13% of brand exposures and, under the paper's body-text extraction definition, 71% of contact-information exposures are unmatched. The paper further claims that interface type is a first-order stratifier

What carries the argument

The central machinery is a four-layer measurement framework built on a unified citation-level corpus: citation selection (a URL entering the citation list), citation absorption (categorized as nominal, general, or deep via a fixed LLM rubric), entity exposure (brands or contact details appearing in answer text), and interface consistency (App versus Web source sets compared by Jaccard overlap). The key objects are the sets P (citation-pool brands), A (answer-exposed brands), and A−P (exposures unmatched to the contemporaneous citation pool). Predictive models, mainly gradient-boosted trees with SHAP importance, are used to rank which external, non-circular factors—content fit, cross-source o

Load-bearing premise

The load-bearing measurement premise is that the crawled body text of cited pages—which excludes buttons, pop-ups, footer contact cards, and JSON-LD—is an adequate basis for deciding that contact information is unmatched to a cited source; if phone numbers, WeChat accounts, and official URLs live in those excluded regions, the 71% unmatched-exposure finding becomes a crawl artifact rather than evidence that answers use information absent from the cited sources.

What would settle it

Re-crawl a sample of cited pages without excluding structured regions—footer contact cards, JSON-LD, buttons, pop-ups—and re-run the contact-information matching. If the traceability rate rises from 28.9% toward 90% or higher, the claim that 71% of contact exposures are unmatched to cited sources collapses; if it stays near 30%, the paper's interpretation survives.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X LinkedIn Reddit HN

If this is right

  • If correct, content providers cannot treat a citation as the goal: being listed yields only an 8.3% chance that a cited brand is written into the answer, so visibility strategies must target answer-body inclusion, not just citation appearance.
  • The monotonic cross-source redundancy gradient—2.1% selection when a brand appears in one citation versus 59.9% when it appears in eleven or more—implies that repeated mention across independent sources is a practical lever for exposure.
  • The fitted half-life of cited pages for high-timeliness queries is about 39 days, giving a concrete freshness window: for time-sensitive questions, cited content clusters within roughly a month of publication.
  • The App–Web divergence, with domain-level Jaccard coefficients as low as 0.19, means evaluations and optimizations must be reported per interface; results from one access route do not transfer to the other.
  • The weak showing of the composite SEO quality score suggests traditional authority metrics are not the main drivers of citation absorption or entity exposure in these systems, pointing toward content fit and redundancy as more actionable signals.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if cited pages were re-crawled including footer contact cards, buttons, pop-ups, and JSON-LD, the 71% unmatched contact-information rate could fall substantially; the paper itself flags this as an open measurement question in its limitations and Section 3.4.1.
  • Editorial inference: the roughly 13% unmatched brand exposures could reflect parametric memory or retrieval sources omitted from the visible citation list; a targeted probe using deliberately obscure brand aliases could separate these explanations.
  • Editorial inference: the interface gap suggests App and Web may draw on separate retrieval indices or ranking policies rather than merely different presentation formats; a controlled test using identical queries with varied device and cookie signatures could test this.
  • Editorial inference: the 39-day half-life for high-timeliness queries offers a testable content-strategy benchmark—for breaking or rapidly changing topics, pages older than roughly a month are unlikely to be cited, so regular updating could be a practical lever.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper reports a large-scale observational audit of Chinese-language generative search engines. It collects 614 controlled queries across the Web and App interfaces of Doubao, DeepSeek, Tencent Yuanbao, and Qwen, with three replications per query–platform–interface combination, producing a cleaned citation-level dataset of 160,860 records. The paper develops a unified measurement framework distinguishing citation selection, citation absorption, entity exposure, and interface consistency, and reports five headline findings: (1) brands in citation pools are selectively surfaced in answers, with an overall selection rate of 8.3%, and 12.4% of cited sources containing contact information have that information exposed in answers; (2) content fit, cross-source occurrence count, and semantic role dominate predictive models, whereas the 5118-Baidu Composite Quality Score is not the leading predictor for any examined outcome; (3) among cited pages with publication dates, fitted half-lives are about 39 days for high-timeliness and 68 days for low-timeliness queries; (4) about 13% of brand exposures and 71% of contact-information exposures cannot be matched to the contemporaneous citation pool or crawled body text, respectively; and (5) source sets differ systematically between App and Web interfaces of the same platform. The paper claims these regularities characterize second-stage selection after citation, provenance gaps in generated answers, and interface-level stratification.

Significance. If the empirical regularities hold, the paper makes a useful contribution: it is one of the first multi-platform, citation-level audits of Chinese generative search engines, and it provides quantitative evidence that visible citations do not fully account for entities exposed in answers and that access interface is a first-order stratifier. The study design has real strengths: a controlled query space, three fixed replications, explicit primary-key construction, a source-authenticity check, a stated feature-exclusion principle (§1.3.2) that prevents citation-depth and semantic-similarity variables from entering their own outcome models, and robustness checks including seed stability (0.991) and set-size-controlled cross-interface comparisons. The paper is also unusually candid in its limitations. However, the strength of the headline conclusions is uneven. The brand-selection and cross-interface findings are anchored in more direct measurements, while the contact-information unmatched-exposure finding depends on a crawl that explicitly excludes the structured regions where contact details are likely to live. The absence of any human gold-standard validation for the LLM-generated an

major comments (3)
  1. [§1.1.3, §3.4.1, Table 3-4, Research Limitations] The 71% unmatched-contact-information finding is not supported by the current crawl definition. §1.1.3 states that the original-text crawl excludes 'interactive and structured components' such as page buttons, pop-up windows, footer contact cards, and JSON-LD. §3.4.1 itself concedes that URLs, WeChat Official Accounts, and hotlines 'may also occur in structured regions.' These are exactly the categories with the highest answer-exposure-to-source-detection ratios in Table 3-4 (3.17, 3.62, 4.20). The abstract's gloss that these exposures are 'not covered by the citation text visible at the time' is therefore stronger than the operational measurement warrants. A full-page re-crawl test, or at minimum a structured-region extraction pass on a sample, is needed before the unmatched-exposure claim can be read as evidence of a provenance gap. Until then, the conclusion should be restricted to 'n
  2. [§1.3.3, §2.2, §3.2] The three model-assisted variables—semantic similarity score, degree of citation absorption, and semantic role—were all generated by a single LLM (Doubao-Seed-2.1-Pro) under a fixed rubric, with no human gold standard and no inter-annotator or model-variation reliability check. These variables are central to the predictive-importance findings: Table 2-2's 'content fit' variable is the LLM similarity score, and semantic role drives the R²=0.622 regression in §2.3. The paper cites Ziems et al. (2024) to justify LLM annotation, but does not report any agreement statistic, error analysis, or sensitivity analysis against human coding on a subsample. Because the headline finding that 'content fit, cross-source occurrence count, and semantic role are relatively important' depends on annotation quality, a validation sample with independent human labels, or at least a second LLM with measured agr
  3. [§1.5.1, §3.4, Table 3-3] The contact-information exposure model reports 'brand-related queries had an odds ratio of approximately 89' with p<0.001 but no confidence interval. Although §1.4.1 says some odds ratios are point estimates and conclusions are limited to direction, the text in §3.4 and the Discussion interprets the 89× odds ratio as evidence that 'query type is the primary stratification factor.' Given the sparsity of positive cases acknowledged in §1.4.1, a coefficient of this magnitude requires at least a confidence interval or a warning about separation/instability. Reporting the interval would strengthen the claim and is a minor addition.
minor comments (4)
  1. [Data and Research-Materials Availability] The full crawled text and analysis scripts are not released, and the public repository contains only raw records and field documentation. Since the unmatched-exposure findings depend on crawl coverage, releasing the extraction code or a detailed crawl-coverage audit would materially aid verification.
  2. [§4.4, Table 4-3] The claim that between-platform differences exceed between-industry differences is based on comparing the range of four platform-level Jaccard means (0.32) with the range of six industry-level means (0.059). Ranges over different sets of units are not a formal comparison; a variance-component or bootstrap interval would be more appropriate. The descriptive nature of the claim should be stated.
  3. [§2.5, Table 2-5] The half-life fits report no uncertainty intervals, although H1's Cliff's δ has a bootstrap CI. Since §1.4.1 states that some half-lives are point estimates, adding approximate CIs (e.g., from bootstrap or profile likelihood) would help readers judge the 39- vs 68-day contrast.
  4. [Abstract and §3.4.1] Minor wording: the abstract says contact-information exposures 'could not be matched to the crawled body text,' which is accurate, but the following clause 'suggesting that some entities presented in answers were not covered by the citation text visible at the time' overstates the implication; the manuscript's own limitation text should be reflected in the abstract.

Circularity Check

1 steps flagged

Headline rates are direct measurements; one acknowledged role/depth label overlap partially inflates the 'semantic role' feature-importance result.

specific steps
  1. self definitional [Section 2.3 (Semantic Role and Degree of Citation Absorption), Table 2-3; cf. §1.3.2]
    "Background context was lower, while the nominal and 'other' categories were close to the nominal baseline by definition."

    The semantic-role analysis of citation depth includes a role category 'Nominal (No Substantive Use)' whose mean depth is 1.002 and whose nominal share is 99.5%. Because nominal citation depth is assigned depth 1, this feature category is a near-restatement of the outcome's baseline. The paper itself states in §1.3.2 that 'the nominal semantic role is nearly identical to a nominal citation at citation-depth level 1' and warns that including such definition-derived variables can inflate R2. Thus the reported R2 = 0.622 and the claim that 'semantic-role variables made a relatively large predictive contribution' are partly true by construction, even though the paper discloses the overlap and excludes semantic similarity from this model.

full rationale

The paper's headline quantities—8.3% brand selection, 12.4% contact-information conversion, 13% unmatched brand exposure, 71% unmatched contact-information exposure, the ~39/68-day half-lives, and the App/Web Jaccard coefficients—are direct measurements from the collected corpus, not outputs of a derivation that feeds back into its own inputs. The modeling sections implement a feature-exclusion principle (§1.3.2, §1.5.3) to avoid target leakage and explicitly discard a shopping-card model whose features overlapped the target definition (§3.5), so the main attribution claims are not forced by construction. The one genuine circular element I can exhibit is localized to §2.3: the semantic-role predictor contains a 'Nominal (No Substantive Use)' category that is nearly identical to citation-depth level 1, making part of the 'semantic role is important' result a restatement of the outcome definition. This is acknowledged in the paper and does not undermine the independent 5118-Baidu-score conclusion or the direct exposure/interface findings. The 71% unmatched-contact-information statistic is sensitive to the body-text crawl definition that excludes buttons, footer cards, pop-ups, and JSON-LD, but that is a measurement-validity limitation, not circular reasoning. Overall, the central empirical claims are self-contained measurements with one disclosed, non-central definitional overlap.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

The empirical claims rest on a small number of measurement assumptions rather than mathematical axioms. Free parameters are minimal: the recency-decay rate is explicitly fitted, and citation-depth thresholds are hand-chosen. The critical assumptions are the fidelity of the proprietary capture tool, the reliability of LLM annotations, and the completeness of the body-text crawl; the last directly affects the headline 71% unmatched-contact figure. No invented entities are introduced.

free parameters (2)
  • Recency decay rate λ (per timeliness group) = Half-lives ≈ 39 days (high timeliness), 68 days (low timeliness), 60 days (full dataset)
    Fitted by the log-linear model N(d) ∝ e^(−λ·d) to observed publication-age distances; reported as a headline result, though it is a descriptive within-sample fit with no uncertainty interval.
  • Citation-depth thresholds on semantic similarity score = 0.30 and 0.65
    Prespecified cutoffs dividing semantic-similarity scores into nominal/general/deep citation absorption; all depth-based results depend on these hand-chosen thresholds.
axioms (5)
  • standard math Validity of standard statistical machinery (LightGBM, SHAP, logistic/ordinal regression, Jaccard, Cliff's δ, log-linear half-life fit) for the data.
    All quantitative claims rely on these standard methods; no proofs are supplied, which is normal for an empirical audit.
  • domain assumption LLM-generated labels from Doubao-Seed-2.1-Pro for semantic similarity, citation absorption, and semantic role are reliable enough operational measurements.
    §1.3.3; no human gold standard; multiple dependent variables and feature-importance results depend on these labels.
  • domain assumption The Aidso monitoring tool captures platform answers, citations, and metadata faithfully across all eight interfaces.
    §1.1.3; proprietary capture with no independent verification; underlies the entire dataset.
  • domain assumption Body-text crawl, excluding structured regions (buttons, pop-ups, footer contact cards, JSON-LD), adequately represents cited-page content for brand/contact matching.
    §1.1.3 and §3.4.1; central to unmatched-exposure rates; paper itself notes potential underestimation.
  • domain assumption The 614 controlled queries and June–July 2026 window are representative enough to describe Chinese-language generative search behavior.
    §1.1 and Research Limitations; non-random sample of query space and no complete candidate pool from the open Web.

reviewed 2026-08-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of What Do Chinese-Language Generative Search Engines Cite and Surface? A Large-Scale Empirical Study." pith.science (2026). https://pith.science/paper/PG55ECVV

@misc{pith2026260715771,
  author       = {Pith},
  title        = {Pith review of: What Do Chinese-Language Generative Search Engines Cite and Surface? A Large-Scale Empirical Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PG55ECVV}},
  note         = {Machine review of arXiv:2607.15771}
}
Share X LinkedIn Reddit HN
read the original abstract

Generative AI question-answering systems increasingly mediate information access, shifting content visibility from ranked search results to retrieval, citation, and presentation in generated answers. We conduct a large-scale empirical study of Chinese-language generative search across the Web and App interfaces of four mainstream platforms. The controlled design covers eight platform interfaces, 614 queries, and three replications per query-platform-interface combination. From 214,119 raw records, we construct a cleaned citation-level dataset of 160,860 records and analyze citation behavior, source attribution, entity exposure, and cross-interface consistency. Five findings emerge. First, brands in the citation pool were selectively surfaced in answers: the overall brand-selection rate was 8.3%, and 12.4% of retrieved sources containing contact information contributed contact information to answers. Second, content fit, cross-source occurrence count, and semantic role were relatively important in predictive models, whereas the 5118-Baidu Composite Quality Score was not the leading predictor for any examined outcome. Third, among cited pages with publication dates, fitted half-lives were approximately 39 days for high-timeliness queries and 68 days for low-timeliness queries. Fourth, approximately 13% of brand exposures could not be matched to the contemporaneous citation pool, and approximately 71% of contact-information exposures could not be matched to the crawled body text. Fifth, source sets differed systematically between the App and Web interfaces of the same platform. These results characterize how Chinese-language generative search systems select, attribute, and surface information and show that interface type is an important dimension of analysis.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

5 extracted references · 1 canonical work pages

  1. [1]

    Vu, T., Iyyer, M., Wang, X., Constant, N., Wei, J., Wei, J., Tar, C., Sung, Y.-H., Zhou, D., Le, Q., & Luong, T. (2024). FreshLLMs: Refreshing large language models with search engine augmentation. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 13697–13720). Association for Computational Linguistics. https://doi.org/10.18653/v...

  2. [2]

    Yu, G., Jin, L., & Su, F. (2026). The characteristics, effect disputes, and potential value of large-model hallucination [in Chinese]. Journalism and Communication Review, 79(1), 14–24. https://doi.org/10.14086/j.cnki.xwycbpl.2026.01.002

  3. [3]

    Yuan, X. (2022). A customized and separated method for building websites (Patent No. CN 114329270A) [in Chinese]. China National Intellectual Property Administration

  4. [4]

    Zhang, K., He, X., & Yao, J. (2026). From citation selection to citation absorption: A measurement framework for generative engine optimization across AI search platforms. arXiv:2604.25707. https://arxiv.org/abs/2604.25707

  5. [5]

    Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z., & Yang, D. (2024). Can large language models transform computational social science? Computational Linguistics, 50(1), 237–291. https://doi.org/10.1162/coli_a_00502

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.