REVIEW 3 major objections 5 minor 47 references
Curated retrieval versus open web search in public AI information services: a coverage-trust trade-off
T0 review · 3 major / 5 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read Open web search answers more civic questions than a curated corpus, but experts flag sources in over a third of those answers.
desk verdict Solid pre-deployment expert study that makes source trustworthiness measurable for public AI services; the 35% web-flag rate is real but rests on thin IRR and non-blinded single-reviewer judgments. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A dual review instrument that scores whole-answer quality on a seven-criterion yes/no rubric while separately flagging individual cited sources with mode-specific reasons (outdated for curated; untrustworthy or irrelevant for web), applied by domain experts to paired answers on the same 287 questions.
What would settle it
A replication with balanced, planned multi-reviewer overlap on real user queries that finds the share of web answers carrying an untrustworthy or irrelevant source far below one third, or that finds ordinary surface quality scores reliably predict expert source flags.
Extended reading notes
Core claim
In a controlled pre-deployment comparison of RAG over a vetted local corpus versus open web search, web search answered more questions at the cost of source quality: experts flagged at least one cited source in 35 percent of reviewed web answers (65 of 187), nearly always as untrustworthy or irrelevant, while curated sources were flagged far less often and only for being outdated. Answer fluency and topical fit carried no signal of whether the sources underneath were sound.
Load-bearing premise
That five experts' non-blinded, thinly overlapping flags of individual web sources as untrustworthy or irrelevant are a reliable measure of the source-trust risk that would face real citizens.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a pre-launch expert evaluation of Evrópuvefur, an Icelandic public AI service answering EU-related questions ahead of the 29 August 2026 referendum. Five domain experts produced 551 evaluations of 449 answers under two retrieval modes (curated RAG vs open web search), scoring a seven-criterion answer rubric and flagging individual cited sources. Headline results: at least one source was flagged in 35% (65/187) of reviewed web answers, almost always as untrustworthy or irrelevant, versus 6% for RAG (outdated only); web search answered far more questions while the curated corpus declined when coverage failed; fluency and topical fit did not predict source flags within the web path; a trusted-domain list in the prompt raised on-list citations only from 12% to 21%; and RÚV was never cited across 287 web answers. The authors frame source trustworthiness as a measurable information-quality dimension and discuss transparency-oriented responses.
Significance. If the results hold, the paper makes a timely and practically useful contribution to public-sector AI evaluation. It supplies rare expert-evaluated evidence on source trustworthiness in a low-resource language and high-stakes civic setting, introduces a reusable two-part review instrument that separates answer quality from per-source flags, and documents a coverage–trust trade-off with multiple corroborating strands (matched-question tests, disclaimer analysis, free-text comment audit fully reproduced in Appendix C, Fjölmiðlanefnd audience profiling, and a controlled prompt ablation). The finding that prompt-level domain lists only weakly steer citations, and that surface fluency does not signal source quality, is directly actionable for procurement and governance of public AI information services. Strengths include transparent limitation reporting, Wilson CIs, McNemar/Wilcoxon on the matched subset, and full reproduction of the 77 web flag comments for audit.
major comments (3)
- [§5.8, Abstract, Conclusion] Section 5.8 and Appendix A: the headline 35% (65/187) web flag rate rests on thin dual-review support—only 40 of 128 flags fall on doubly-reviewed answers, and only six answers received independent dual flags—so flag reliability is unestimated. The paper already treats flags as individual judgments rather than adjudicated rulings, but the Abstract and Conclusion still present “more than a third” as a prevalence claim. Please qualify the headline proportion more explicitly in the Abstract and Conclusion (e.g., “a reviewer flagged…” / single-reviewer observational rate) and state that dual-flag agreement could not be estimated, so the figure is not an adjudicated prevalence.
- [§4.2, §4.5, Fig. 1B, Abstract] Sections 4.2, 4.5, and 7: mode was not blinded (reviewers saw local article links vs external URLs), and assignment was a shared queue rather than balanced randomisation (262 RAG vs 187 web evaluations). Cross-mode flag comparisons are already labelled descriptive, but the coverage–trust trade-off narrative still juxtaposes the 35% and 6% rates in Figure 1B and the Abstract. Please either (a) report the matched-subset flag rates only as secondary descriptive statistics and lead with within-web and within-RAG analyses, or (b) add a sensitivity discussion of how mode visibility and unequal review volume could inflate the web flag share, and keep the primary claim as the within-web trust problem plus the coverage gap on ‘answers the question’ (McNemar on 179 matched questions).
- [§5.4] Section 5.4: the claim that fluency and topical fit do not predict source trustworthiness is load-bearing and rests on Fisher exact tests within the web path (flagged vs unflagged on language quality, scope, hallucinations, answers-question; all p>0.05). With only 65 flagged web answers and multiple criteria, power is limited and non-significance is not strong evidence of decoupling. Please report effect sizes (e.g., risk differences or odds ratios with CIs) alongside p-values, note the exploratory multiple-comparison setting already flagged in §7, and soften “did not predict” to “showed no statistically detectable association on surface criteria in this sample.”
minor comments (5)
- [Fig. 1B] Figure 1B caption correctly notes that available flag reasons differ by mode, but the y-axis label “Reviewed answers with a flagged source” still invites a like-for-like reading. Consider adding “(descriptive; reasons not comparable)” in the panel title.
- [§5.6–5.7] Section 5.6 / Figure 6: the production system never cites RÚV, yet the ablation cites RÚV 50 times in each arm. The configuration-sensitivity point is important; a short explicit sentence in the Abstract or Discussion that citation mix is configuration-dependent (structured output vs free-text parse; timing) would help readers avoid over-generalising the RÚV absence.
- [§4.1] Section 4.1: question generation used esbvaktin.is both as seed material and as the trusted-domain classifier. The residual feedback-loop risk is acknowledged in §7; a one-sentence note in Methods that no esbvaktin content entered answer context would make the mitigation easier to find.
- [Appendix A, Table 1] Table 1 (Appendix A): report n of pairwise overlaps per criterion or note that AC1 is computed on the 82 doubly-reviewed answers only, so readers can judge precision of the coefficients.
- [Abstract / body] Minor typography: several run-together words appear in the compiled text (e.g., “reportapre-launchexpertevaluation”, “Wecomparedtwo”). Please re-export with correct spacing before camera-ready.
Circularity Check
No circularity: purely empirical expert evaluation with independent measurements; no derivation reduces to fitted inputs or self-citation.
full rationale
This paper reports a pre-deployment expert evaluation of two retrieval modes (curated RAG vs open web search) for a public AI information service. The central claims are observational proportions (e.g., 35% of reviewed web answers had at least one flagged source; web answered more questions; fluency did not predict source flags) and a prompt ablation measuring list compliance (12% vs 21%). There is no mathematical derivation, no fitted parameter renamed as a prediction, and no uniqueness theorem or ansatz imported from the authors' prior work. The trusted-domain list is taken from an external fact-checking project (esbvaktin.is) and is not tuned to produce the flag rates. Expert flags and Fjölmiðlanefnd survey audience skews are independent measurements. The service's own corpus was deliberately kept off the web-search guidance so that the evaluation would measure external sources. Self-citations (e.g., Einarsson 2026a,b on Icelandic LLM performance) are background only and not load-bearing for the results. The work is self-contained against its own evaluation export; any concerns about thin inter-rater overlap or non-blinding are reliability/correctness issues, not circularity. Score 0 is the correct honest finding.
Assumptions & free parameters
assumptions (4)
- domain assumption Domain-expert binary judgments of individual cited sources as untrustworthy/irrelevant/outdated constitute a valid operationalisation of source trustworthiness for public AI services.
- domain assumption The 287 LLM-generated questions, seeded from esbvaktin.is clusters, adequately span the public debate that real citizens would query.
- domain assumption Wang & Strong (1996) multi-dimensional information-quality framework correctly places source trustworthiness in the intrinsic-quality family alongside believability and reputation.
- domain assumption Mainstream vs. alternative-media distinction (editorial standards vs. viewpoint expression) is a useful lens for interpreting flagged domains in a small media market.
invented entities (2)
-
Two-part expert review instrument (7-criterion answer rubric + per-source flagging with mode-asymmetric reasons)
-
Coverage–trust trade-off framing for public AI retrieval paths
Cite this review
Pith. "Pith review of Curated retrieval versus open web search in public AI information services: a coverage-trust trade-off." pith.science (2026). https://pith.science/paper/IBTMKQW3
@misc{pith2026260705217,
author = {Pith},
title = {Pith review of: Curated retrieval versus open web search in public AI information services: a coverage-trust trade-off},
year = {2026},
howpublished = {\url{https://pith.science/paper/IBTMKQW3}},
note = {Machine review of arXiv:2607.05217}
}
read the original abstract
Public institutions increasingly use large language models (LLMs) to answer citizens' questions, often pairing a curated knowledge base with live web search, yet whether the sources behind these answers can be trusted has received little empirical scrutiny. We report a pre-launch expert evaluation of Evr\'opuvefur, an independent, government-funded service run by the University of Iceland that answers questions about the European Union, conducted as Iceland prepared for its referendum of 29 August 2026 on whether to resume EU accession talks. Five domain experts produced 551 evaluations of 449 AI-generated answers, scoring each against a seven-criterion quality rubric and, separately, flagging individual cited sources. We compared two retrieval paths: a curated local corpus (RAG) and open web search. In more than a third of the reviewed web-search answers (35%, 65 of 187), at least one cited source was flagged, almost always as untrustworthy or irrelevant; curated sources were flagged far less often and only for being out of date. Web search answered more questions, but at the cost of source quality; the curated corpus was trustworthy yet limited in coverage, and the model declined to respond when it fell short. The citation mix also passed over strong sources: across all 287 web-search answers, the system never cited R\'UV, the public broadcaster and the country's most widely used news source. A companion prompt ablation shows how weak prompt-level steering is: a trusted-domain list in the system prompt raised the share of citations to listed domains only from 12% to 21%. Fluency and topical fit did not predict source trustworthiness. We argue that source trustworthiness is a measurable yet largely invisible dimension of information quality in public AI services, and we discuss transparency-oriented responses and their trade-offs.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Government Information Quarterly , year =
Hemesath, Sebastian and Tepe, Markus , title =. Government Information Quarterly , year =
-
[2]
and Strong, Diane M
Wang, Richard Y. and Strong, Diane M. , title =. Journal of Management Information Systems , year =
-
[3]
Trust and risk in e-government adoption , journal =
B. Trust and risk in e-government adoption , journal =. 2008 , volume =
2008
-
[4]
, title =
Metzger, Miriam J. , title =. Journal of the American Society for Information Science and Technology , year =
-
[5]
Retrieval-augmented generation for knowledge-intensive
Lewis, Patrick and Perez, Ethan and Piktus, Aleksandra and Petroni, Fabio and Karpukhin, Vladimir and Goyal, Naman and K. Retrieval-augmented generation for knowledge-intensive. Advances in Neural Information Processing Systems , year =
-
[6]
ACM Computing Surveys , year =
Ji, Ziwei and Lee, Nayeon and Frieske, Rita and Yu, Tiezheng and Su, Dan and Xu, Yan and Ishii, Etsuko and Bang, Yejin and Madotto, Andrea and Fung, Pascale , title =. ACM Computing Surveys , year =
-
[7]
and Wilder, Esther Isabelle , title =
Walters, William H. and Wilder, Esther Isabelle , title =. Scientific Reports , year =
-
[8]
New Media & Society , year =
Ananny, Mike and Crawford, Kate , title =. New Media & Society , year =
Show all 47 references
-
[9]
, title =
Gwet, Kilem L. , title =. British Journal of Mathematical and Statistical Psychology , year =
-
[10]
Richard and Koch, Gary G
Landis, J. Richard and Koch, Gary G. , title =. Biometrics , year =
-
[11]
, title =
Wilson, Edwin B. , title =. Journal of the American Statistical Association , year =
-
[12]
Krippendorff, Klaus , title =
-
[13]
Larger and more instructable language models become less reliable , journal =
Zhou, Lexin and Schellaert, Wout and Mart. Larger and more instructable language models become less reliable , journal =. 2024 , volume =
2024
-
[14]
Journal of Computational Social Science , year =
Kuznetsova, Elizaveta and Makhortykh, Mykola and Vziatysheva, Victoria and Stolze, Martha and Baghumyan, Ani and Urman, Aleksandra , title =. Journal of Computational Social Science , year =
-
[15]
, title =
Zhou, Yujia and Liu, Yan and Li, Xiaoxi and Jin, Jiajie and Qian, Hongjin and Liu, Zheng and Li, Chaozhuo and Dou, Zhicheng and Ho, Tsung-Yi and Yu, Philip S. , title =. 2024 , note =
2024
-
[16]
2024 , number =
Governing with Artificial Intelligence: Are Governments Ready? , institution =. 2024 , number =
2024
-
[17]
2025 , note =
Content Credentials:. 2025 , note =
2025
-
[18]
and Burke-Moore, Liam and Chan, Ryan Sze-Yin and Enock, Florence E
Williams, Angus R. and Burke-Moore, Liam and Chan, Ryan Sze-Yin and Enock, Florence E. and Nanni, Federico and Sippy, Tvesha and Chung, Yi-Ling and Gabasova, Evelina and Hackenburg, Kobi and Bright, Jonathan and Carrasco-Farr. Large language models can consistently generate hi...
2025
-
[19]
, title =
Schlicht, Erik J. , title =. 2024 , note =
2024
-
[20]
Government Information Quarterly , year =
Ju, Jingrui and Meng, Qingguo and Sun, Fangfang and Liu, Luning and Singh, Shweta , title =. Government Information Quarterly , year =
-
[21]
Larsen, A. G. and F. The impact of chatbots on public service provision: A qualitative interview study with citizens and public service providers , journal =. 2024 , volume =
2024
-
[22]
Government Information Quarterly , year =
Li, Xuesong and Wang, Jian , title =. Government Information Quarterly , year =
-
[23]
and Esnaashari, Saba and Francis, John and Hashem, Youmna and Morgan, Deborah , title =
Bright, Jonathan and Enock, Florence E. and Esnaashari, Saba and Francis, John and Hashem, Youmna and Morgan, Deborah , title =. Digital Government: Research and Practice , year =
-
[24]
2026 , doi =
Majithia, Neil and Shinde, Rajat and Maskey, Manil and Simperl, Elena and Shadbolt, Nigel , title =. 2026 , doi =
2026
-
[25]
2026 , note =
Onweller, Hailey and Lumer, Elias and Huber, Austin and Ramchandani, Pia and Subbiah, Vamse Kumar and Feld, Corey , title =. 2026 , note =
2026
-
[26]
2026 , note =
Germain, Thomas , title =. 2026 , note =
2026
-
[27]
Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) , year =
Einarsson, Hafsteinn , title =. Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) , year =
2026
-
[28]
Proceedings of the RESOURCEFUL Workshop at LREC 2026 , year =
Einarsson, Hafsteinn , title =. Proceedings of the RESOURCEFUL Workshop at LREC 2026 , year =
2026
-
[29]
Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (EACL) , year =
Muennighoff, Niklas and Tazi, Nouamane and Magne, Lo. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (EACL) , year =
-
[30]
2024 , note =
Offenhartz, Jake , title =. 2024 , note =
2024
-
[31]
2026 , note =
Government proposes referendum on whether to return to accession talks with the. 2026 , note =
2026
-
[32]
The 2024 Al
Einarsson, Hafsteinn and Har. The 2024 Al. Icelandic Review of Politics and Administration , year =. doi:10.13177/irpa.a.2025.21.1.1 , url =
2024 doi
-
[33]
2024 , volume =
Polarisation, News Consumption, and Beliefs in Misinformation and Conspiracy Theories: Early Signs of the Fragmentation of the Public Sphere in Iceland , journal =. 2024 , volume =. doi:10.1080/13183222.2024.2383905 , url =
2024 doi
-
[34]
2024 , pages =
Iceland , booktitle =. 2024 , pages =
2024
-
[35]
2021 , volume =
Superficial, Shallow and Reactive: How a Small State News Media Covers Politics , journal =. 2021 , volume =. doi:10.2478/nor-2021-0018 , publisher =
2021 doi
-
[36]
2017 , number =
Wardle, Claire and Derakhshan, Hossein , title =. 2017 , number =
2017
-
[37]
2026 , howpublished =
Icelandic. 2026 , howpublished =
2026
-
[38]
2026 , month =
Kristj. 2026 , month =
2026
-
[39]
The International Journal of Press/Politics , year =
Fletcher, Richard and Cornia, Alessio and Nielsen, Rasmus Kleis , title =. The International Journal of Press/Politics , year =
-
[40]
and Eddy, Kirsten and Nielsen, Rasmus Kleis , title =
Newman, Nic and Fletcher, Richard and Robertson, Craig T. and Eddy, Kirsten and Nielsen, Rasmus Kleis , title =. 2022 , url =
2022
-
[41]
2024 , howpublished =
2024
-
[42]
and Sastry, Girish and Musser, Micah and DiResta, Ren
Goldstein, Josh A. and Sastry, Girish and Musser, Micah and DiResta, Ren. Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations , year =
-
[43]
and Schoch, David and Stier, Sebastian and Yang, JungHwan , title =
Keller, Franziska B. and Schoch, David and Stier, Sebastian and Yang, JungHwan , title =. Political Communication , year =
-
[44]
2025 , note =
A Well-funded. 2025 , note =
2025
-
[45]
Harvard Kennedy School Misinformation Review , year =
Alyukov, Maxim and Makhortykh, Mykola and Voronovici, Alexandr and Sydorova, Maryna , title =. Harvard Kennedy School Misinformation Review , year =
-
[46]
and Mercea, Dan , title =
Bastos, Marco T. and Mercea, Dan , title =. Social Science Computer Review , year =
-
[47]
New Perspectives , year =
Marshall, Hannah and Drieschova, Alena , title =. New Perspectives , year =
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.