{"id":"c9bf6eae-14da-4840-84ea-a026ff03e62a","arxiv_id":"2607.05217","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Web search answered more EU questions than curated RAG but had at least one expert-flagged (mostly untrustworthy) source in 35% of reviewed answers; a trusted-domain prompt list only raised on-list citations from 12% to 21%.","lead":"Expert review of Iceland's public EU-info AI service found open web search answers far more questions than a curated corpus but cites sources experts flag as untrustworthy or irrelevant in 35% of cases. The coverage–trust trade-off, plus weak prompt steering and the total absence of the public broadcaster among citations, is directly relevant to governments deploying LLMs before elections.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The 35% web-flag rate rests on non-blinded, thin-overlap expert judgments whose reliability is unmeasured, so the headline proportion may not be a stable measure of source-trust risk.","rationale":"The reader correctly isolates the thin reliability and non-blinded character of the expert flags as the weakest assumption underwriting the 35% claim. The paper itself flags the sparse overlap and treats flags as individual judgments (Section 5.8), yet still leads with the 35% figure as the primary empirical result. No derivation failure or circularity is present; the measurement instrument and ablation remain useful. The concern therefore does not overturn the CONDITIONAL verdict or the paper’s contribution to measurement practice, but it confirms that the proportion should not yet be treated as a stable benchmark without the reliability check above. Agreement with the reader is full on the load-bearing point; the concrete test simply operationalizes the reliability gap the reader already named.","tokens_in":31625,"tokens_out":587,"duration_ms":5739,"concrete_test":"On the six answers that already carry independent dual flags, compute raw agreement and Gwet AC1 for the binary “any source flagged” decision and for the untrustworthy/irrelevant labels. Separately, re-present a random 40 of the 187 web answers to two new blinded reviewers who see only the cited URLs (no mode label) and recompute the “at-least-one-flagged” rate; if dual-flag AC1 < 0.4 or the blinded rate falls below ~25%, the 35% headline weakens as a stable benchmark.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that “in more than a third of the reviewed web-search answers (35%, 65 of 187) at least one cited source was flagged, almost always as untrustworthy or irrelevant” treats the five experts’ binary flags as a reliable measure of source-trust risk. Section 5.8 and Appendix A show that only 82 of 449 answers were doubly reviewed, only 40 of 128 flags fall on those answers, and only six answers received independent dual flags—too few to estimate flag reliability. Mode was not blinded (reviewers saw local links vs external URLs), reviewers may hold EU-related priors (Section 7), and the instrument applied different reasons by mode, so the 35% figure is a single-reviewer observational rate, not an adjudicated prevalence. If dual-flag agreement is low or mode visibility systematically inflates web flags, the coverage–trust trade-off and the claim that fluency does not predict trustworthiness rest on a soft foundation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper reports a pre-launch expert evaluation of Evrópuvefur, an Icelandic public AI service answering EU-related questions ahead of the 29 August 2026 referendum. Five domain experts produced 551 evaluations of 449 answers under two retrieval modes (curated RAG vs open web search), scoring a seven-criterion answer rubric and flagging individual cited sources. Headline results: at least one source was flagged in 35% (65/187) of reviewed web answers, almost always as untrustworthy or irrelevant, versus 6% for RAG (outdated only); web search answered far more questions while the curated corpus declined when coverage failed; fluency and topical fit did not predict source flags within the web path; a trusted-domain list in the prompt raised on-list citations only from 12% to 21%; and RÚV was never cited across 287 web answers. The authors frame source trustworthiness as a measurable information-quality dimension and discuss transparency-oriented responses.","tokens_in":31908,"tokens_out":1574,"duration_ms":20061,"significance":"If the results hold, the paper makes a timely and practically useful contribution to public-sector AI evaluation. It supplies rare expert-evaluated evidence on source trustworthiness in a low-resource language and high-stakes civic setting, introduces a reusable two-part review instrument that separates answer quality from per-source flags, and documents a coverage–trust trade-off with multiple corroborating strands (matched-question tests, disclaimer analysis, free-text comment audit fully reproduced in Appendix C, Fjölmiðlanefnd audience profiling, and a controlled prompt ablation). The finding that prompt-level domain lists only weakly steer citations, and that surface fluency does not signal source quality, is directly actionable for procurement and governance of public AI information services. Strengths include transparent limitation reporting, Wilson CIs, McNemar/Wilcoxon on the matched subset, and full reproduction of the 77 web flag comments for audit.","major_comments":[{"comment":"Section 5.8 and Appendix A: the headline 35% (65/187) web flag rate rests on thin dual-review support—only 40 of 128 flags fall on doubly-reviewed answers, and only six answers received independent dual flags—so flag reliability is unestimated. The paper already treats flags as individual judgments rather than adjudicated rulings, but the Abstract and Conclusion still present “more than a third” as a prevalence claim. Please qualify the headline proportion more explicitly in the Abstract and Conclusion (e.g., “a reviewer flagged…” / single-reviewer observational rate) and state that dual-flag agreement could not be estimated, so the figure is not an adjudicated prevalence.","section":"§5.8, Abstract, Conclusion"},{"comment":"Sections 4.2, 4.5, and 7: mode was not blinded (reviewers saw local article links vs external URLs), and assignment was a shared queue rather than balanced randomisation (262 RAG vs 187 web evaluations). Cross-mode flag comparisons are already labelled descriptive, but the coverage–trust trade-off narrative still juxtaposes the 35% and 6% rates in Figure 1B and the Abstract. Please either (a) report the matched-subset flag rates only as secondary descriptive statistics and lead with within-web and within-RAG analyses, or (b) add a sensitivity discussion of how mode visibility and unequal review volume could inflate the web flag share, and keep the primary claim as the within-web trust problem plus the coverage gap on ‘answers the question’ (McNemar on 179 matched questions).","section":"§4.2, §4.5, Fig. 1B, Abstract"},{"comment":"Section 5.4: the claim that fluency and topical fit do not predict source trustworthiness is load-bearing and rests on Fisher exact tests within the web path (flagged vs unflagged on language quality, scope, hallucinations, answers-question; all p>0.05). With only 65 flagged web answers and multiple criteria, power is limited and non-significance is not strong evidence of decoupling. Please report effect sizes (e.g., risk differences or odds ratios with CIs) alongside p-values, note the exploratory multiple-comparison setting already flagged in §7, and soften “did not predict” to “showed no statistically detectable association on surface criteria in this sample.”","section":"§5.4"}],"minor_comments":[{"comment":"Figure 1B caption correctly notes that available flag reasons differ by mode, but the y-axis label “Reviewed answers with a flagged source” still invites a like-for-like reading. Consider adding “(descriptive; reasons not comparable)” in the panel title.","section":"Fig. 1B"},{"comment":"Section 5.6 / Figure 6: the production system never cites RÚV, yet the ablation cites RÚV 50 times in each arm. The configuration-sensitivity point is important; a short explicit sentence in the Abstract or Discussion that citation mix is configuration-dependent (structured output vs free-text parse; timing) would help readers avoid over-generalising the RÚV absence.","section":"§5.6–5.7"},{"comment":"Section 4.1: question generation used esbvaktin.is both as seed material and as the trusted-domain classifier. The residual feedback-loop risk is acknowledged in §7; a one-sentence note in Methods that no esbvaktin content entered answer context would make the mitigation easier to find.","section":"§4.1"},{"comment":"Table 1 (Appendix A): report n of pairwise overlaps per criterion or note that AC1 is computed on the 82 doubly-reviewed answers only, so readers can judge precision of the coefficients.","section":"Appendix A, Table 1"},{"comment":"Minor typography: several run-together words appear in the compiled text (e.g., “reportapre-launchexpertevaluation”, “Wecomparedtwo”). Please re-export with correct spacing before camera-ready.","section":"Abstract / body"}],"recommendation":"minor_revision","confidential_remarks":"The reliability concern raised by the stress-test is real but already largely disclosed by the authors; I do not see it as grounds for reject or major_revision if the Abstract/Conclusion qualifications and effect-size reporting are fixed. Fit for a serious cs.CY / digital-government venue is good. No novelty or citation-pattern concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful thing here is a clean pre-launch comparison of curated RAG versus open web search on the same 287 Icelandic EU-referendum questions, with experts scoring answers and separately flagging sources. The coverage–trust trade-off is the real result: web answered far more often (91.5% vs 48.5% on “answers the question”), but 35% of reviewed web answers (65/187) carried at least one flagged source, almost always untrustworthy or irrelevant; RAG flags were rare and only for staleness. Fluency and topical fit did not predict source quality. The prompt ablation is clean: a trusted-domain list only moved on-list citations from 12% to 21%. RÚV never appeared in the 287 production web answers. That package is new for public AI information services, especially in a low-resource language and high-stakes civic setting.\n\nWhat they do well: the two-part instrument (seven-criterion rubric + per-source flags with mode-asymmetric reasons) is reusable and separates answer surface from source trust. Methods are transparent—Wilson CIs, McNemar/Wilcoxon on the 179 matched questions, full reproduction of the 77 web flag comments with codes in Appendix C, and an honest limitations section. The political-audience map from the Fjölmiðlanefnd survey independently corroborates the expert flags. No circularity; the trusted list is external.\n\nSoft spots are real but proportionate. Inter-rater overlap is thin by design (only 82 doubly-reviewed answers; only six answers with independent dual flags), so flag reliability is unmeasured and the 35% figure is a single-reviewer observational rate, not an adjudicated prevalence. Mode was not blinded (reviewers saw local links vs external URLs), questions were generated rather than real-user, and assignment was unbalanced. The authors already flag all of this. It does not sink the central pattern—web sources were frequently ones experts would not endorse, and prompt steering was weak—but it means the exact 35% should not be treated as a stable benchmark yet.\n\nThis is for people building or overseeing public-sector LLM services, election-integrity work, and RAG evaluation. It deserves a serious referee. I would engage with it, cite the measurement approach and the trade-off framing, and push for tighter IRR and real-user queries in revision. Send it out.","headline":"Solid pre-deployment expert study that makes source trustworthiness measurable for public AI services; the 35% web-flag rate is real but rests on thin IRR and non-blinded single-reviewer judgments.","tokens_in":32540,"tokens_out":592,"would_cite":true,"duration_ms":5863,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Open web search answers more civic questions than a curated corpus, but experts flag sources in over a third of those answers.","keywords":["artificial intelligence","large language models","retrieval-augmented generation","information quality","source trustworthiness","public information","expert evaluation","coverage-trust trade-off"],"falsifier":"A replication with balanced, planned multi-reviewer overlap on real user queries that finds the share of web answers carrying an untrustworthy or irrelevant source far below one third, or that finds ordinary surface quality scores reliably predict expert source flags.","tokens_in":32553,"feed_emoji":"🔍","tokens_out":615,"duration_ms":5304,"temperature":0.7,"pith_summary":"Public AI services often ground answers either in a curated knowledge base or in live open-web search. This paper shows, through a pre-launch expert review of an Icelandic government-funded EU information service ahead of a national referendum, that the two paths trade coverage against trust. Web search answered far more questions, yet in 35 percent of the reviewed web answers experts flagged at least one cited source, almost always as untrustworthy or irrelevant. The curated corpus was trusted by design but frequently lacked material, and the model declined rather than inventing. Prompt instructions to prefer trusted domains barely steered citation behaviour, fluency did not predict source quality, and the country's most-used public broadcaster was never cited. The authors argue that source trustworthiness is a measurable, currently invisible dimension of information quality that public services should log, audit, and disclose.","feed_headline":"Web search answers more, but experts flag sources in 35%","feed_subtitle":"Public AI services face a coverage–trust trade-off that fluency alone does not reveal","key_machinery":"A dual review instrument that scores whole-answer quality on a seven-criterion yes/no rubric while separately flagging individual cited sources with mode-specific reasons (outdated for curated; untrustworthy or irrelevant for web), applied by domain experts to paired answers on the same 287 questions.","core_discovery":"In a controlled pre-deployment comparison of RAG over a vetted local corpus versus open web search, web search answered more questions at the cost of source quality: experts flagged at least one cited source in 35 percent of reviewed web answers (65 of 187), nearly always as untrustworthy or irrelevant, while curated sources were flagged far less often and only for being outdated. Answer fluency and topical fit carried no signal of whether the sources underneath were sound.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Web search covers more but experts flag sources in 35% of answers","Curated RAG stays trusted; open web brings untrustworthy citations","Fluency hides it: 35% of web answers had flagged sources","RAG declines more yet sources flag far less than web search","Never cited Iceland's top news: coverage-trust trade-off in public AI"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That five experts' non-blinded, thinly overlapping flags of individual web sources as untrustworthy or irrelevant are a reliable measure of the source-trust risk that would face real citizens.","fun_headline_variants_meta":{"raw":{"variants":["Web search covers more but experts flag sources in 35% of answers","Curated RAG stays trusted; open web brings untrustworthy citations","Fluency hides it: 35% of web answers had flagged sources","RAG declines more yet sources flag far less than web search","Never cited Iceland's top news: coverage-trust trade-off in public AI"]},"model":"grok-4.5","effort":"low","cost_usd":0.005072,"raw_usage":{"total_tokens":1486,"prompt_tokens":913,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":50720000,"prompt_tokens_details":{"text_tokens":913,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":496,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":913,"tokens_out":77,"duration_ms":3825,"temperature":1.0,"reasoning_tokens":496,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T07:43:53.940362+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A replication with balanced, planned multi-reviewer overlap on real user queries that finds the share of web answers carrying an untrustworthy or irrelevant source far below one third, or that finds ordinary surface quality scores reliably predict expert source flags.","supporting_citations":[],"review_version":2}