Pith. sign in

REVIEW 3 major objections 6 minor 19 references

Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Across 2,208 queries to four AI assistants, at least 85.6% of the 4,776 restaurants, cafes, and bars in two Balinese markets were never recommended; documentation decides entry, star rating only decides rank.

desk verdict A genuinely new census-denominated audit with a solid invisibility floor, but the two-margin factor story rests on a case-control design that Prentice-Pyke does not justify; the entry-margin estimates need a full-cohort refit before the dissociation is secure. read the letter →

arxiv 2608.07069 v1 pith:PGLTDOFG submitted 2026-08-07 cs.IR cs.CY

classification cs.IRcs.CY
keywords AIrecommendationauditcensus-denominatedmeasurementvenueinvisibilityratetwo-marginvisibilitylocalfood-and-drinkdiscoverygenerativeengineoptimizationentitymatchingvalidationsearch-groundedassistants
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to measure what AI assistants actually surface when a traveler asks where to eat or drink, using a complete enumeration of an entire market as the denominator. It claims that invisibility is the norm — at least 85.6% of the 4,776 cafes, restaurants, and bars across two Balinese submarkets were never recommended by any of four production systems in 2,208 runs — and that visibility is governed by two margins with different drivers. Entry into an answer tracks documentation: review volume, an own website, listed price information, and third-party web mentions all raise the odds, while star rating has no detectable effect at this margin. Once a venue is in an answer, the pattern reverses: rating significantly predicts being named first. A sympathetic reader would care because the factor venues invest most in signaling — their star rating — appears to do nothing for AI discovery, while infrastructure like having a website governs who is surfaced at all.

What carries the argument

The load-bearing object is the census itself: a complete enumeration of 4,776 Google-listed food-and-drink venues across two bounded geographic polygons, built by adaptive grid subdivision of the Places Nearby Search API (recursively splitting saturated 20-result cells down to a 130-meter floor) plus a resolution pass that folded 70 mention-probed venues of atypical listing type into the frame. Because every mention is resolved against this full registry, 'never recommended' becomes a measurable population rate rather than a relative prominence judgment. The second piece of machinery is the two-margin outcome decomposition: the entry margin models whether a venue is recommended at all (a binomial regression model over venue x persona x engine exposure opportunities in a case-control frame, cluster-robust by venue, with Benjamini-Hochberg multiple-comparison correction), and the rank margin models which recommended venue is named first (a conditional logit over within-run choice sets). The dissociation between the two margins — documentation admits, rating ranks — is the identity that carries the argument, and the paper reads it through the retrieval-then-generation architecture of the audited systems.

What would settle it

Re-run the identical 96-query instrument through the four vendors' consumer-facing configurations, or through server settings matching production consumer apps, and compare the invisibility rate and factor estimates against the same 4,776-venue census; if the never-recommended share falls materially below 85.6%, or if star rating becomes a significant positive predictor of entry, the two-margin structure (documentation admits, rating ranks) is an artifact of the API access route rather than a property of the assistants users interact with.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is a census-denominated measurement of AI venue visibility: across 2,208 search-grounded responses from ChatGPT, Claude, Gemini, and Perplexity to 96 persona-conditioned queries about two Bali submarkets, 4,087 of the 4,776 enumerated venues — 85.6%, a floor — were never recommended by any system; even among venues with fifty or more ratings the never-recommended share is 72.6%, and visibility among the surfaced is long-tailed rather than winner-take-all (the leading venue holds 1.9% of recommendations). The paper's explanatory claim is a two-margin dissociation: admission into an answer is associated with documentation — review volume (1.64x), an own website (1.92x), listed price (1.54x), and web mentions (1.44x) as odds multipliers — while star rating is null at entry (0.89x), but within-answer ordering reverses the pattern, with rating significantly predicting first position (1.17x). Presence in the Foursquare open POI dataset shows no positive effect at either margin. The systems almost never fabricate venue names (0.08% of mentions) yet recommended permanently closed venues 93 times, making staleness rather than hallucination the practical failure mode, and a two-week test-retest shows answer churn is sampling stochasticity, not temporal drift.

Load-bearing premise

The whole measurement rests on two proxy choices: search-grounded production APIs standing in for the consumer assistants, and the Google Places listing set standing in for the market, so if provider-side serving configurations differ from the consumer apps, or the English tourist-and-nomad query mix under-represents real demand, the measured invisibility rates and factor associations may not match what users actually see.

Editorial extensions

If this is right

  • A 4.9-star venue with thin documentation is, to these systems, indistinguishable from an absent one: improving a rating without growing the documentation trail should not be expected to change whether a venue is surfaced at all.
  • Single-shot, single-engine visibility checks measure sampling noise; meaningful measurement requires repetition across engines and phrasings.
  • The clearest correctable harm is staleness: 93 recommendations of permanently closed venues show that closure signals propagate more slowly than reputation signals in the sources assistants retrieve.
  • Because top-20 agreement between engines is only 0.33-0.54 (just eight venues appear in all four engines' top-20 lists), optimizing visibility for one assistant is not optimizing for AI in general.
  • For independent venues the action hierarchy is to be documented before being excellent: an own website, complete platform profiles, review volume, and third-party mentions precede any payoff from rating.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the two-margin structure is a property of retrieval-then-generation architecture rather than of these four systems, the same dissociation — documentation admits, rating ranks — should appear in other verticals such as hotels and services, where conjoint experiments already show rating dominance within constructed choice sets; a census-denominated replication in a structurally different market w
  • The Foursquare null is the first direct test of a widely assumed visibility lever, but the observational design cannot settle whether the dataset simply goes unread by retrieval or is redundant with web presence; a clean follow-up would register randomly chosen thin-documentation venues in open POI datasets and measure whether visibility moves.
  • Because the invisibility rate only rises when the frame broadens (above 92% under the broadest defensible estimate), markets with thinner Google coverage than Bali plausibly show even higher invisibility, making the documented associations a lower bound on the discoverability penalty faced by small independent venues.
  • The paper's demonstration that a matching-error remediation reversed one factor's apparent effect from positive to null raises a testable suspicion that earlier brand-level audits without validated matching have unstable factor conclusions; re-running those designs with census denominators and audited matchers would be the direct check.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper presents a census-denominated audit of AI venue recommendation. The authors enumerate 4,776 Google-Places-listed cafés, restaurants, and bars in two Bali submarkets (Canggu and Ubud) and query four production AI systems (ChatGPT, Claude, Gemini, Perplexity) through their search-grounded APIs with 96 persona-conditioned queries over seven days, yielding 2,208 runs and 12,439 valid mentions, under a pre-registered protocol with a frozen analysis snapshot. Five findings are reported: (1) at least 85.6% of venues were never recommended by any system, and 72.6% among venues with fifty or more ratings; (2) entry into answers is associated with documentation signals (own website OR 1.92, review volume OR 1.64, price listed OR 1.54, web mentions OR 1.44) while star rating is null at entry (OR 0.89) yet predicts first position within answers (OR 1.17); (3) Foursquare presence is null at both margins; (4) fabrication is rare (0.08% of mentions) while 93 recommendations went to permanently closed venues; (5) cross-system top-20 overlap is low (Jaccard 0.33–0.54) and a two-week test–retest shows run-to-run churn without temporal drift. The paper also reports two methodological lessons: a matching remediation reversed a factor's apparent effect, and the hours-listed covariate encoded collection provenance. The protocol, instrument, code, and derived data are released, and all tables and figures are claimed to be byte-reproducible.

Significance. If the central claims survive, this is an important and unusually rigorous contribution to the algorithm-audit literature. The census denominator converts 'never recommended' from a relative-prominence statement into a population rate, a structural advance over catalogue-based audits. The two-margin dissociation (documentation admits, rating ranks) is specific, falsifiable, and offers a concrete reconciliation of the conjoint and observational audit traditions. The instrument-validation discipline is exemplary: pre-registration, a frozen analysis snapshot, double-annotated extraction with adjudication, an audited entity matcher whose revision is disclosed to have reversed a coefficient, and a covariate–provenance artifact reported in full. The byte-reproducible pipeline and the publication of null results (Foursquare; rating at entry) that run against the funder's commercial interest strengthen credibility. The principal caveats — English tourist/nomad queries, two Balinese submarkets, search-grounded APIs as proxies for consumer apps, and an operationally defined Google-Places frame — are acknowledged in Section 7.

major comments (3)
  1. [§4.4–4.5, Table 1] The case-control frame in §4.4 samples at the venue level (all 749 ever-mentioned venues plus 620 never-mentioned controls), but M1 in §4.5 is a trial-level binomial GLM over venue × persona × engine cells. The cited Prentice–Pyke (1979) consistency result covers logistic regression under outcome-based sampling of individual units; it does not cover sampling on a cluster-level aggregate of the outcome. Because a venue's sampling probability (ever mentioned) is a nonlinear function of its trial-level success probabilities and hence of the venue-level predictors, the unweighted trial-level likelihood is misspecified with respect to the sampling design, and the slope estimates in Table 1 (website 1.92, rating 0.89, Foursquare 0.84) can be biased in either direction. The robustness analyses (M2/M3/M5) address correlation and subsets, not selection. Since trial outcomes for all census venues are recoverable from the collected corpus (never-mentioned venues have all-zero outcomes by construction), a full-cohort fit over all 4,776 venues is feasible and should be reported as the primary analysis; alternatively, weight venues by the inverse of their selection probability and show that the coefficients are insensitive. This is load-bearing because the paper's headline 'documentation admits, rating ranks' dissociation rests on the M1 odds ratios.
  2. [§4.1] The pre-registered extraction gate was ≥95% accuracy, with the metric level left unspecified. After remediation, strict run-level agreement is 91.5%, below that nominal gate, whereas mention-level precision (97.6%) and recall (99.1%) clear it. The interpretation under which the gate is passed was not fixed in advance, so it needs an argument rather than an assertion: either show that the run-level disagreements involve extractions irrelevant to the outcome definitions, or treat the gate as not met and provide a sensitivity analysis restricted to runs with perfect extraction agreement. The disclosure of all three numbers is commendable, but the confirmatory status of the study depends on the gate being met under a pre-committed interpretation.
  3. [§5.4, Table 3] The selection-artifact caveat is applied selectively in the interpretation of M7. Section 5.4 declines to interpret the negative Foursquare coefficient because 'within-set coefficients of entry-relevant variables can carry selection artifacts,' yet presents the positive rating coefficient (OR 1.17) as evidence that 'rating ranks.' The same conditioning logic applies to every M7 coefficient, since the choice sets are generated by an entry process that appears to depend on the documentation variables (and possibly on rating, if the entry-margin null is revised). Please apply the caveat consistently and demonstrate that the rating-rank finding is not an artifact of differential selection into choice sets, for example by checking the stability of M7 coefficients across choice-set sizes or by estimating a joint entry/rank specification.
minor comments (6)
  1. [§5.4 vs Table 3] The text reports the star-rating p-value in M7 as p = .0002, while Table 3 reports p < .001; align the two statements.
  2. [§4.4] Please clarify the exclusion of 94 venues from the case-control set (1,369 → 1,275): how many were excluded for operational status versus missing rating, and whether exclusion is associated with case/control status or with the outcome.
  3. [§5.3, Table 2] The Foursquare ladder (M4) is fitted on only 480 venues with three non-significant terms and uncorrected p-values; report the minimum detectable effect or a power statement so the reader can gauge how informative the null actually is.
  4. [§5.2] Make explicit that the invisibility rates are conditional on the 96-query instrument and on the search-grounded API access route; the current 'at least 85.6%' phrasing is easily misread as a claim about all user queries or about what consumer apps display.
  5. [§5.7, Figure 9] The Gemini citation domains are recovered from citation titles, with unrecoverable citations excluded; state whether the reported top-domain ordering is robust to alternative recovery assumptions.
  6. [Title/Abstract] Consider qualifying 'census' as 'Google-Places-listed' in the title or abstract; the operational frame in Section 3.2 is clear, but the unqualified phrase overstates the frame at a glance.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the invisibility rate and two-margin associations are direct measurements, not derivations from fitted inputs or self-citations.

full rationale

The paper's central claims are measured outcomes rather than reductions. The invisibility rate (85.6%) is computed directly from 2,208 runs and 4,776 enumerated venues; no parameter is fitted to define it. The entry-margin odds ratios (M1) and rank-margin odds ratios (M7) are estimated from observed trial successes and venue covariates; no predictor is constructed from the outcome, and the two margins are operationally distinct (admission vs. first position), so 'documentation admits, rating ranks' is an empirical dissociation, not a definitional identity. The paper contains no self-citations: all references are to external prior work, and no load-bearing claim rests on a uniqueness theorem or ansatz imported from the same authors. The two disclosed self-referential elements, the funder's commercial interest and the hours-listed covariate provenance contamination, are disclosed and do not make any result true by construction; the hours variable is explicitly excluded from interpretation. A statistical concern about the unweighted cluster-level case-control design (Prentice-Pyke applicability) is a validity issue, not circularity: it does not render the estimates equal to their inputs, and the 85.6% floor is independent of the factor model.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

The analysis is empirical; no invented entities. The main inputs from outside are the Google Places frame, the API behaviors, and the English query mix. Statistical model assumptions are standard. The hand-set matching thresholds and recency window are the clearest free parameters, and they are sensitivity-tested.

free parameters (3)
  • Entity matching score thresholds = full-name >=87; core veto >=80; core >=90; sub-95 stratum for sensitivity
    Hand-set thresholds in the deterministic matcher (Section 4.2). They affect which mentions resolve to census venues; the matching audit estimates 98-99% accuracy, so the impact is likely small but not zero.
  • Review-recency window = 90 days
    Defines the review-recency predictor (Section 4.4). A different window could move the suggestive OR 1.41 estimate across the significance threshold.
  • Extraction filter blocklists = 259 wave mentions removed (2.0% of raw extractions)
    Deterministic post-extraction rules with blocklists (Section 4.1) define valid mentions; the rules were pre-committed but are a model choice that affects the numerator.
assumptions (6)
  • standard math Case-control logistic regression yields consistent slope estimates without weighting (Prentice and Pyke, 1979).
    Relied on in Section 4.4 for the M1 case-control estimation frame.
  • standard math Benjamini-Hochberg correction and cluster-robust standard errors control multiple testing and within-venue correlation.
    Used in M1 and described in Section 4.5.
  • domain assumption The Google Places food-service listing set is the population frame AI systems can plausibly surface.
    Section 3.2 defines the census operationally; venues outside this frame are underrepresented by design, though the authors argue this biases invisibility rates upward as floors.
  • domain assumption The search-grounded APIs audited are faithful proxies for consumer-facing assistants.
    Section 3.3 and Section 7 note that provider-side serving configurations are opaque and that consumer apps may differ.
  • domain assumption The 96 English persona-conditioned queries approximate the demand distribution for local discovery in these markets.
    Section 7 limits generalization; prior work shows linguistic and geographic variation in LLM recommendations.
  • ad hoc to paper Claude's search budget cap of two searches per run does not materially change conclusions.
    Section 3.3 discloses the cap and Section 4.5 (M5) reports a sensitivity refit dropping the Claude arm with conclusions unchanged.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census." pith.science (2026). https://pith.science/paper/PGLTDOFG

@misc{pith2026260807069,
  author       = {Pith},
  title        = {Pith review of: Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PGLTDOFG}},
  note         = {Machine review of arXiv:2608.07069}
}
read the original abstract

AI assistants are becoming a primary interface for local discovery, yet almost nothing is known about which venues they surface -- especially in food and drink, where recommendations carry direct revenue consequences. We present the first census-denominated audit of AI venue recommendation: a complete enumeration of 4,776 cafes, restaurants, and bars across two bounded markets (Canggu and Ubud, Bali), against which we evaluate 2,208 search-grounded responses from four production AI systems (ChatGPT, Claude, Gemini, Perplexity) to 96 persona-conditioned queries, collected over seven days under a pre-registered protocol. Because we observe the full market, we can measure what sampled audits cannot: 85.6% of venues were never recommended by any system -- 72.6% even among established venues with fifty or more ratings. Visibility follows a two-margin structure. Entry into answers is associated with documentation: review volume (OR 1.64), an own website (OR 1.92), listed price information (OR 1.54), and third-party web mentions (OR 1.44) -- while star rating is null at this margin (OR 0.89). Rank within answers reverses the pattern: among recommended venues, rating significantly predicts first position (OR 1.17). Presence in an open POI dataset (Foursquare), a folk-theorized visibility factor, shows no positive effect at either margin. Outright fabrication is rare (0.08% of mentions), but systems recommended permanently closed venues 93 times -- staleness, not hallucination, is the practical failure mode. Cross-system agreement is low (top-20 Jaccard 0.33-0.54). A two-week test-retest shows cross-period answer similarity comparable to same-day rerun similarity: the churn is sampling stochasticity, not temporal drift. We release our protocol, registry construction method, and derived data.

Figures

Figures reproduced from arXiv: 2608.07069 by the authors.

Figure 1
Figure 1. Invisibility by denominator. Share of census venues never mentioned and never [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Concentration among the visible. Cumulative share of all 9,791 pooled recommenda [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Entry-margin factor estimates (M1). Odds ratios with 95% confidence intervals, log [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: The two-margin structure. Entry-margin (M1, BH-corrected) and rank-1-margin (M7, [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]
Figure 5
Figure 5. Figure 5: Repetition vs. paraphrase stability. Mean top-20 Jaccard overlap between venue sets [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Cross-engine agreement. Pairwise Jaccard overlap between the four engines’ top-20 [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: Engine fingerprints. Distinct venues recommended, mention volume per run, repetition [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: The consensus core. Top-20 membership for every venue appearing in at least two [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Citation diets. Top cited domains by share of each engine’s grounding citations. [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: Decomposition of the confirmatory wave’s pre-recovery unmatched pool (2,452 [PITH_FULL_IMAGE:figures/full_fig_p021_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

19 extracted references · 9 canonical work pages

  1. [1]

    Dining and Drinking

    The Jaccard-versus-baseline comparison above is therefore the cleaner stability statement. Together they support the paper’s treatment of wave-period visibility as a persistent venue property measured through a stochastic channel (Section 6.3). B Census completeness The census is a complete enumeration of an operationally defined frame — the Google-listed...

  2. [6]

    M. Chen, X. Wang, K. Chen, and N. Koudas. Generative engine optimization: How to dominate AI search.arXiv preprint arXiv:2509.08919,

  3. [7]

    arXiv:2604.27790. Y. Hou, J. Zhang, Z. Lin, H. Lu, R. Xie, J. McAuley, and W. X. Zhao. Large language models are zero-shot rankers for recommender systems. InAdvances in Information Retrieval (ECIR),

  4. [8]

    arXiv:2305.08845. M. Iannelli and A. Ai. From prompt to purchase: How AI brand recommendations move consumers on the open web.arXiv preprint arXiv:2606.10907,

  5. [10]

    arXiv:2406.13997. A. Kumar and H. Lakkaraju. Manipulating large language models to increase product visibility. arXiv preprint arXiv:2404.07981,

  6. [11]

    P. Kumar. Generative engine optimization at scale: Measuring brand visibility across AI search engines.arXiv preprint arXiv:2606.20065,

  7. [14]

    Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)

    O. Martinez. Optimizing visibility in generative engines: A critical survey of generative engine optimization (2023–2026).arXiv preprint arXiv:2607.14035,

  8. [16]

    An Investigation of Linguistic Biases in LLM-Based Recommendations

    N. Venkateswaran, J. Ang, D. Adhikari, and T. K. Dasari. An investigation of linguistic biases in LLM-based recommendations.arXiv preprint arXiv:2604.25456,

Show all 19 references
  1. [17]

    ˙Zatuchin

    D. ˙Zatuchin. Who owns the AI recommendation? a multi-industry empirical map of brand category ownership across large language models.arXiv preprint arXiv:2606.23057,

  2. [2000]

    Global is good, local is bad?

    W. Jack, N. Lehman, K. Maloney, and S. Xu. Paraphrase brittleness in production retrieval- augmented commercial recommendation: Reproducibility below the rerun-stability baseline. arXiv preprint arXiv:2605.27440, 2026a. W. Jack, N. Lehman, K. Maloney, and S. Xu. Prominence-str...

  3. [2012]

    Andre, G

    A. Andre, G. Roy, E. Dyer, and K. Wang. Revealing potential biases in LLM-based recommender systems in the cold start setting.arXiv preprint arXiv:2508.20401,

  4. [2014]

    Y. Seo, W. Jeong, E. Kim, H. Jang, and D. Lee. Verified misguidance: Measuring structural citation failures in search-augmented LLMs.arXiv preprint arXiv:2605.28565,

  5. [2016]

    Ma, C.-Y

    S.-D. Ma, C.-Y. Chen, B.-A. Li, P.-Y. Chen, S.-Y. Hsu, and Y.-N. Chen. Rethinking fairness in LLM-based recommender systems: A survey.arXiv preprint arXiv:2606.28340,

  6. [2020]

    Bhagat, K

    K. Bhagat, K. Vasisht, and D. Pruthi. Richer output for richer countries: Uncovering ge- ographical disparities in generated stories and travel recommendations.arXiv preprint arXiv:2411.07320,

  7. [2021]

    J. M. Lichtenberg, A. Buchholz, and P. Schw¨ obel. Large language models as recommender systems: A study of popularity bias.arXiv preprint arXiv:2406.01285,

  8. [2023]

    arXiv:2305.07609. K. Zhang, X. He, and J. Yao. From citation selection to citation absorption: A measurement framework for generative engine optimization across AI search platforms.arXiv preprint arXiv:2604.25707, 2026a. M. Zhang, F. Yi, and D. Gursoy. The effects of generativ...

  9. [2024]

    arXiv:2311.09735. M. Anderson and J. Magruder. Learning from the crowd: Regression discontinuity estimates of the effects of an online review database.The Economic Journal, 122(563):957–989,

  10. [2025]

    M. S. A. Baig, S. A. Gillani, and A. Ali. Whose hotel does the AI recommend? an algorithm audit of reputation signals in LLM-assisted hotel selection.arXiv preprint arXiv:2606.16344,

  11. [2026]

    Banerjee, G

    A. Banerjee, G. K. Patro, L. W. Dietz, and A. Chakraborty. Analyzing “Near Me” services: Potential for exposure bias in location-based retrieval.arXiv preprint arXiv:2011.07359,

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.