Pith. sign in

REVIEW 7 cited by

Evaluating Verifiability in Generative Search Engines

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.09848 v2 pith:JCG5HHMN submitted 2023-04-19 cs.CL cs.IR

classification cs.CLcs.IR
keywords generativesearchcitationsenginesqueriessystemsassociatedcitation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative search engines directly generate responses to user queries, along with in-line citations. A prerequisite trait of a trustworthy generative search engine is verifiability, i.e., systems should cite comprehensively (high citation recall; all statements are fully supported by citations) and accurately (high citation precision; every cite supports its associated statement). We conduct human evaluation to audit four popular generative search engines -- Bing Chat, NeevaAI, perplexity.ai, and YouChat -- across a diverse set of queries from a variety of sources (e.g., historical Google user queries, dynamically-collected open-ended questions on Reddit, etc.). We find that responses from existing generative search engines are fluent and appear informative, but frequently contain unsupported statements and inaccurate citations: on average, a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence. We believe that these results are concerningly low for systems that may serve as a primary tool for information-seeking users, especially given their facade of trustworthiness. We hope that our results further motivate the development of trustworthy generative search engines and help researchers and users better understand the shortcomings of existing commercial systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. Equal Accuracy, Unequal Evidence: Search APIs as Decision Surfaces for Tool-Using Agents

    cs.CL 2026-07 conditional novelty 6.5 of 10

    Equal semantic accuracy across Brave, Tavily, and Firecrawl (25–26/100) hides sharply different pre-fetch support, rank-1 concentration, contradiction ratios, and agent exploration regimes under one frozen agent.

  2. Characterizing Web Search in The Age of Generative AI

    cs.IR 2025-10 conditional novelty 6.0 of 10

    AI search engines vary greatly in how much they rely on web pages versus internal model knowledge, and these differences shift which sources and concepts users see.

  3. HIRAG: Hierarchical-Thought Instruction-Tuning Retrieval-Augmented Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A hierarchical chain-of-thought instruction-tuning curriculum for filtering, combination, and reasoning improves zero-shot retrieval-augmented QA.

  4. MORPHEUS: A Multidimensional Framework for Modeling, Measuring, and Mitigating Human Factors in Cybersecurity

    cs.CR 2025-12 conditional novelty 5.0 of 10

    A unified framework that maps 50 human factors, 295 pairwise interactions, and 99 psychometric tools to phishing, malware, passwords, and misconfiguration risks.

  5. Paladin-mini: A Compact and Efficient Grounding Model Excelling in Real-World Scenarios

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A compact fine-tuned grounding model and a purpose-built benchmark aim to expose and close gaps in numerical, temporal, and logical fact-checking.

  6. Hallucination Detection with Small Language Models

    cs.CL 2025-06 reject novelty 5.0 of 10

    A multi-small-model ensemble with sentence splitting, z-score normalization, and harmonic mean detects hallucinations in RAG answers with a reported 10% F1 gain over single-model baselines.

  7. MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation

    cs.AI 2026-07 conditional novelty 4.0 of 10

    MPR-CiteG combines four hand-designed query portfolios with reranking and sentence-level citation grounding; it ranked second in the ScienceON AI Challenge.

Pith tools