{"id":"b1fba5ed-365f-4834-bc90-29bd4f82d9b5","arxiv_id":"2411.17375","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"Across five defined operating points, more abstractive LLM answers are rated more useful but are cited less accurately and take longer for users to verify.","lead":"This paper defines a spectrum of citation behavior between search engines and large language models, and measures how perceived usefulness rises while verifiability falls as responses become more abstractive. It gives system builders an explicit menu of operating points, from pure quotation to free synthesis, with human evaluation of the trade-offs.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'up to 3x time-to-verify' result is confounded by citation count and citation passage length; a controlled re-analysis is needed before accepting the abstraction-specific verification burden claim.","rationale":"The paper is a careful empirical study with transparent limitations, released code/data, and consistent trends across four query distributions and a pilot. I find no issue with the utility/coverage trade-off: the reference OPs share grounding quotes (Section 4.2), and the coverage decline is robust. The load-bearing weakness is the T2V metric, specifically the attribution of the 3x effect to abstraction. My proposed regression is feasible with released data and would settle whether T2V tracks abstraction or merely citation quantity and passage length. The reader's weakest assumption points to human-eval validity; I partially agree but make the confound concrete. If the controlled analysis supports the paper's interpretation, acceptance stands; if not, the T2V claim should be tempered. Hence CONDITIONAL rather than REJECT or UNCHANGED.","tokens_in":46929,"tokens_out":11252,"duration_ms":106236,"concrete_test":"Using the released per-sentence T2V data, fit a mixed-effects regression of log(T2V) on operating point (quoted as baseline), number of citations in the sentence, and total character count of cited passages, with annotator random effects. If the operating-point coefficients for paraphrased/entailed/abstractive/GPT-4+Vertex are no longer positive and significant relative to quoted, the abstraction-specific T2V claim is not supported. As a robustness check, restrict to sentences with exactly one citation and compare T2V across OPs; if the monotonic increase disappears, the confound is real.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest component of the central claim is that abstraction alone increases verification time ('users take up to 3 times as long to verify cited information'), but the evidence does not isolate abstraction from citation display characteristics. T2V (Section 5.1) is measured on an annotation screen where users evaluate coverage by reading the cited quotes. Across the seven systems, both the number of citations per sentence and the length of cited passages co-vary with abstraction: Table 1 shows GPT-4+Vertex averages 3.14 citations per cited sentence versus 1.19 for quoted, and Appendix 13.6.3 shows Gemini citations are presented as 1000-character source spans, far longer than the ~14.5-word quotes used in reference OPs (Section 4.2). The paper's rebuttal in Section 6.1.4 (Gemini has few citations yet high T2V) is not decisive because Gemini also differs in passage length, source presentation, and generation pipeline. The reference OPs alone show a more modest increase (quoted-to-paraphrased ~40% with nearly matched citation counts), and Appendix 13.3 acknowledges a selection confounder: T2V is only measured for sentences that have citations, and 'generations that remain cited at more abstractive OPs are those that are the easiest to cite (and thereby verify)' (Appendix 13.3). If T2V is driven by citation count and passage length rather than the abstractive relation between sentence and source, the headline 3x number overstates the verification-burden trade-off.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper defines the extractive-abstractive spectrum with five operating points (extractive, quoted, paraphrased, entailed, abstractive), implements reference instantiations of each, and evaluates seven systems in total across four query distributions using human annotators. It reports that as generations become more abstractive, perceived utility rises by as much as 200%, citation coverage falls by as much as 50%, and time-to-verify rises by up to a factor of three. It also presents a user preference survey, a failure analysis, and task-specific recommendations for choosing operating points.","tokens_in":47232,"tokens_out":8694,"duration_ms":84131,"significance":"The study is potentially valuable because it provides a concrete taxonomy and a systematic human evaluation of the utility-verifiability tradeoff, a topic of central importance for deployed LLM systems. The reference implementations are carefully designed: the paraphrased, entailed, and abstractive generations are revisions of the same quoted generation, so the comparisons within the reference family are unusually well controlled. The paper also ships code and human evaluation data, includes a pilot replication, and reports 95% confidence intervals. The strongest empirical contribution is the mapping of deployed systems such as Google Gemini onto the same tradeoff curve. However, the central claim that abstraction itself causes the observed verification-time increase is threatened by confounds in the time-to-verify measurements, and the paper's significance statements are not backed by formal statistical inference. Because the abstractive operating point is defined as permitting uncited claims, the coverage decline at that endpoint is partly definitional; the paper acknowledges this, but the headline wording does not always separate the definitional from the empirical component.","major_comments":[{"comment":"The headline claim that users take up to three times as long to verify cited information as outputs become more abstractive is not supported by the current measurements because abstraction is confounded with citation count and cited-passage length. Across systems, citations per cited sentence range from 1.00 (Gemini) to 3.14 (GPT-4+Vertex) in Table 1, and Appendix 13.6.3 states that Gemini citations are presented as roughly 1000-character source spans, whereas the reference operating points use quotes averaging 14.5 words (Section 4.2). The rebuttal in Section 6.1.4 that Gemini has few citations yet high T2V is not decisive, because Gemini also differs in passage length, source presentation, and generation pipeline. The quoted-to-paraphrased contrast is suggestive because citation counts are similar (1.19 vs 1.33), but the headline 3x figure comes from comparisons that also vary in citation count and passage length. Please provide a controlled re-analysis, for example a regression of log T2V on operating point, citation count, total cited-passage character length, and sentence length with annotator random effects, or a comparison restricted to sentences with a single citation of matched length.","section":"Section 6.1.4, Table 1, Appendix 13.6.3"},{"comment":"The manuscript itself notes that T2V is measured only for sentences with citations and that 'generations that remain cited at more abstractive OPs are those that are the easiest to cite (and thereby verify)'. This selection effect can bias cross-operating-point T2V comparisons in either direction, and it is not quantified. Because the strongest component of the paper's central claim is that abstraction alone increases verification burden, the authors should address this selection issue directly, for example by modeling the selection step, reporting a sensitivity analysis, or bounding the possible bias. As written, the caveat in Appendix 13.3 undermines the confidence with which the 3x T2V result can be interpreted.","section":"Appendix 13.3"},{"comment":"Several statements of significance and the headline magnitudes rest on 95% confidence intervals without formal hypothesis tests. For example, Section 6.1.1 claims that the quoted OP 'significantly improves' fluency and utility, Section 6.1.2 claims precision is 'significantly lower' for GPT-4+Vertex and Gemini, and Section 6.1.4 claims the quoted OP 'significantly expedites' verification. These load-bearing comparative claims should be supported with appropriate inferential procedures, such as pairwise tests with multiplicity correction or mixed-effects models with annotator and query random effects. In addition, perceived utility is measured on a three-point ordinal scale, so the '200% increase' should be reported as a difference in means or as a distributional shift rather than a percentage change.","section":"Section 6.1, Figure 3"}],"minor_comments":[{"comment":"Because perceived utility is measured on a three-point ordinal scale, the '200% increase' is better expressed as a difference in means or as a shift in the distribution of ratings; percentage changes on an ordinal scale can be misleading.","section":"Section 6.1.1, Figure 3"},{"comment":"The text states that the 'claim too specific' precision failure is common for entailed generations, but Example 8.5 is labeled as a GPT-4+Vertex generation; please clarify the intended mapping between the failure category and the example.","section":"Section 8.2.3, Example 8.5"},{"comment":"The Vertex citation threshold alpha is calibrated separately for each dataset (0.5, 0.25, 0.05, 0.25), and this choice directly affects the number of citations and therefore T2V for the GPT-4+Vertex system; a sensitivity analysis over alpha would help establish that the reported results are not driven by this free parameter.","section":"Section 13.6.2"},{"comment":"The claim that annotator noise in precision judgments preserves the ordering of systems would be strengthened by a quantitative sensitivity analysis that removes the flagged false-positive annotations and recomputes the system ordering.","section":"Section 13.5.5, Table 10"},{"comment":"Only 14 of the 31 original annotators continued into the second batch, and T2V was measured in separate batches for the reference systems and for GPT-4+Vertex/Gemini; the main text should state this as a potential batch-effect limitation even though the quoted condition was re-evaluated for normalization.","section":"Section 13.4.2"}],"recommendation":"major_revision","confidential_remarks":"This is a strong empirical study with careful reference implementations, a pilot replication, and open code and data. My main reservation is the time-to-verify confound: the central '3x' claim is not isolated from citation count and passage length. I would not reject the paper on this basis, because the overall tradeoff pattern is robust and valuable, but the authors should either provide a controlled re-analysis or weaken the causal wording. The lack of formal significance tests is also a fixable issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious empirical paper that gives the community a useful vocabulary and a careful map of the utility-verifiability trade-off. The utility and coverage trends hold up. But the headline \"3x time-to-verify\" number is over-attributed to abstraction; it is confounded with citation count and passage length.\n\nWhat is new: the five operating points, especially quoted and paraphrased cited generation as explicit targets, are a real addition. The reference implementations are sensible and shared. The relative T2V normalization (per-annotator, against a quoted baseline) is a reasonable attempt to control reading speed. Evaluating seven systems including Gemini across four query distributions, with the same source quotes across OPs, is a lot of work. The failure analysis is genuinely useful, and the code and data are released. The paper also flags its own limitations unusually clearly (Appendix 13.3, Section 11).\n\nSoft spots: \n\n1. T2V. In the reference OPs, quoted-to-paraphrased is about 40% longer verification with roughly matched citation counts, which is credible. But the \"up to 3x\" headline comes from GPT-4+Vertex, which averages 3.14 citations per cited sentence vs 1.19 for quoted (Table 1). The Gemini rebuttal does not rescue it: Gemini has 1.0 citations but shows 1000-character source spans, while reference quotes average 14.5 words. So passage length is a confound. The paper acknowledges the selection issue in Appendix 13.3 but does not fix it. A controlled comparison (e.g., matching citation count and span length) is needed before claiming abstraction itself triples verification time.\n\n2. The 50% coverage drop is partly by construction: the abstractive OP permits uncited claims, so lower coverage is baked in. The paper says this, but the abstract and intro do not. The empirical finding is really about how much coverage drops under a particular implementation plus the post-hoc systems.\n\n3. Minor: three-point utility scale, no significance tests, and the per-dataset Vertex threshold is a tuned free parameter. These are annoying but not load-bearing.\n\nBottom line: the spectrum and the utility/coverage trade-off deserve to become common reference points. The T2V claim needs a re-analysis or a softer statement. I would send this to peer review, and I would push for the T2V fix before acceptance. Who is it for: anyone building or evaluating RAG or cited generation systems, and the verifiability evaluation community. I would bring it to reading group.","headline":"A careful, useful map of the utility-verifiability trade-off with a new operating-point vocabulary, but the 3x time-to-verify headline is confounded by citation count and passage length.","tokens_in":47747,"tokens_out":2485,"would_cite":true,"duration_ms":24564,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper establishes a quantitative trade-off along an extractive-abstractive spectrum of cited LLM generations: perceived utility rises by up to 200%, properly cited sentences fall by up to 50%, and verification time triples.","keywords":["extractive-abstractive spectrum","citation coverage","citation precision","time-to-verify","LLM attribution","perceived utility","human evaluation","cited generation"],"falsifier":"Re-run the same five-operating-point comparison with an independent annotator pool and a single within-subject batch, controlling for the calibration shift the authors observed between their two evaluation batches; if paraphrased generations are not slower to verify than quoted generations, or if Google Gemini's citation coverage does not fall well below the entailed generations' coverage, the claimed monotone trade-off would not reproduce.","tokens_in":46725,"feed_emoji":"🤖","tokens_out":5960,"duration_ms":55621,"temperature":0.7,"pith_summary":"This paper asks whether an information tool can be as useful as an LLM and as verifiable as a search engine at the same time. It argues that these goals sit on one extractive-abstractive spectrum with five named operating points, and that moving toward abstraction trades verifiability for perceived quality in a measurable way. In human evaluations of seven systems on 480 queries drawn from four real-world distributions, perceived utility increases by as much as 200%, the share of properly cited sentences falls by as much as 50%, and users take up to three times longer to verify cited information when generations become more abstractive. The result matters because it tells system builders which operating point fits which application, and it suggests that better citation retrieval alone will not restore verifiability if generations keep getting more abstractive.","feed_headline":"Fluent AI answers cost citations: verification slows 3x","feed_subtitle":"Five answer styles show perceived quality rising exactly as verifiability falls.","key_machinery":"The load-bearing object is the extractive-abstractive spectrum, with five formally specified operating points. In an extractive output, each attributable unit is exactly one source quote; in a quoted generation, claims are word-for-word quoted substrings of the cited quotes; in a paraphrased generation, cited quotes and the sentence carry the same information in both directions; in an entailed generation, the cited quotes entail the sentence but the sentence may drop information and contract reasoning; in an abstractive generation, the sentence may add claims that no cited quote supports. The evaluations use the \"According to the source\" test for precision and coverage, plus a wall-clock metric, relative time-to-verify, normalized per annotator against quoted-generation times. The spectrum does the argument's work by turning a diffuse worry about citations into a ranked set of user experiences with measurable gaps between adjacent points.","core_discovery":"The central claim is that the extractive-abstractive spectrum, defined by the semantic relation between a generation and its cited source quotes, organizes the trade-off between answer quality and verifiability. At the extractive end, outputs are source snippets with inherent citations; at the abstractive end, outputs may contain claims that no cited quote supports. The paper defines five operating points along this spectrum—extractive, quoted, paraphrased, entailed, and abstractive—and shows through human ratings of fluency, perceived utility, citation precision, citation coverage, and time-to-verify that abstraction improves perceived utility at the direct expense of verifiability. Deployed post-hoc citation systems fall on the same curve: Google Gemini, despite high perceived utility, properly cites only 15.0% of generated sentences. The paper further argues that no single operating point dominates and recommends matching operating points to the stakes and information needs of the application.","pith_inferences":["The trade-off implies that \"verifiable AI\" should be treated partly as a user-interface and cognitive-load design problem, not only as a retrieval or entailment problem; citation metrics that ignore verification time will overstate how verifiable abstractive answers are.","The spectrum could extend to non-text outputs, where extractive answers are direct recordings and abstractive answers are synthesized plans; the same ratio between perceived quality and verification cost may appear in image, code, and agentic settings.","A testable extension is to measure whether users' ability to detect hallucinations decays along the spectrum exactly as verification time grows, since slower verification may cause users to check fewer citations.","As LLMs become more fluent, the frontier may shift outward for perceived utility but not for verifiability, so routing queries to different operating points or mixing operating points within a single response may become necessary for general-purpose assistants."],"forward_implications":["High-stakes settings with dispersed information, such as legal research or clinical case review, should target extractive or quoted generations because verification burden matters most there.","Low-stakes settings that need recombination or creative reformulation, such as writing assistance or brainstorming, are best served by abstractive generations, where utility gains matter and verifiability costs are acceptable.","High-stakes settings that also require style change or logical transformation, such as simplifying medical records or drafting clinical notes, sit in the hardest region; paraphrased and entailed generations are the current compromise, and better entailed implementations are the stated opportunity.","Post-hoc citation systems, which cite sources after generation, show the largest precision and coverage losses; grounding a response in pre-selected quotes, the \"attribute first, then generate\" paradigm, is the more reliable path.","Because time-to-verify rises even for properly covered sentences, making cited text more abstractive imposes a cognitive cost that no improvement in citation retrieval alone can remove."],"supporting_citations":[{"why":"Supplies the citation precision and coverage framing and the deployed-system baselines whose reported accuracy motivates the study.","marker":"[Liu et al., 2023]"},{"why":"Provides the \"According to the source\" attribution test used to train annotators for citation precision and coverage.","marker":"[Rashkin et al., 2022]"},{"why":"Provides the Natural Questions distribution of real Google Search queries used in two of the four evaluation settings.","marker":"[Kwiatkowski et al., 2019]"},{"why":"Provides the 2WikiMH multi-hop dataset with gold sources, used to test abstractions that draw conclusions across premises.","marker":"[Ho et al., 2020]"},{"why":"Provides the MASH healthcare query distribution with gold WebMD sources for the high-stakes medical setting.","marker":"[Zhu et al., 2020]"},{"why":"Supplies the dense passage retrieval method used to select the snippets that ground the reference operating-point implementations.","marker":"[Karpukhin et al., 2020]"},{"why":"Defines the \"attribute first, then generate\" paradigm that the reference implementations follow and against which post-hoc citation failures are contrasted.","marker":"[Slobodkin et al., 2024]"},{"why":"Contributes the corroborative-versus-contributive attribution distinction used to characterize the reference implementations' citations.","marker":"[Worledge et al., 2024]"}],"fun_headline_variants":["Fluent AI costs citations: utility soars 200%, verifiability falls","LLM fluency up, source trust down: spectrum exposes trade-off","High-stakes queries favor search engines over LLMs, survey finds","Abstractive AI answers 3x slower to verify, citations halved","New spectrum maps AI's utility-verifiability tension"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The measured trade-off rests on the assumption that the annotation task, especially wall-clock time spent on the coverage judgment, faithfully captures how real users experience and pay for verification effort.","fun_headline_variants_meta":{"raw":{"variants":["Fluent AI costs citations: utility soars 200%, verifiability falls","LLM fluency up, source trust down: spectrum exposes trade-off","High-stakes queries favor search engines over LLMs, survey finds","Abstractive AI answers 3x slower to verify, citations halved","New spectrum maps AI's utility-verifiability tension"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000297,"raw_usage":{"total_tokens":1768,"prompt_tokens":1036,"completion_tokens":732,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":653}},"tokens_in":652,"tokens_out":732,"duration_ms":7531,"temperature":1.0,"reasoning_tokens":653,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:11:14.459064+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same five-operating-point comparison with an independent annotator pool and a single within-subject batch, controlling for the calibration shift the authors observed between their two evaluation batches; if paraphrased generations are not slower to verify than quoted generations, or if Google Gemini's citation coverage does not fall well below the entailed generations' coverage, the claimed monotone trade-off would not reproduce.","supporting_citations":[],"review_version":1}