{"id":"d6f59ca2-cbfc-495d-b446-489dc46275a8","arxiv_id":"2502.00881","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Interviews with 11 computing survey authors show that keeping narrative surveys up to date is valued but unrewarded, and that updates fall into empirical, structural, and interpretive types.","lead":"Interviews with eleven authors of computing surveys show that keeping such surveys current is seen as valuable but difficult and unrewarded. The study maps the types of updates authors make and suggests where AI tools could help.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Update typology may be an artifact of a leading interview prompt, not an emergent finding; design implications rest on hypothetical intentions, not observed updates.","rationale":"The manuscript is a careful qualitative study with transparent methods and explicit limitations; I am not disputing the credibility of the authors or the internal consistency of the analysis. My concern is narrower: the interview guide appears to supply the very categories the paper reports as findings. The 'Update Elements' question explicitly lists text, structure, figures and tables, and 'other elements,' which is nearly the same partition as empirical/structural/interpretive. This does not mean the authors were dishonest; it means the central typology may be a researcher-imposed frame rather than an emergent participant frame. The reader flagged retrospective self-report as the weakest assumption; I agree but localize the risk more precisely to (a) the hypothetical nature of the updating task and (b) the priming in the prompt. The study would still support the more modest claim that computing researchers perceive updating as valuable but costly and unrewarded, and that they can, when prompted, categorize imagined updates along text/structure/synthesis lines. The stronger claim — that updates 'concisely fall into three types' as a property of real updating work — is not yet secured by the data as presented. Because the reader's conditional verdict already reflects the need for additional evidence, my assessment leaves the verdict unchanged rather than escalating it.","tokens_in":16129,"tokens_out":3984,"duration_ms":41983,"concrete_test":"Request the anonymized interview transcripts and coding sheets. Re-code the pre-'Update Elements' portion of each transcript (before Appendix A's 'Which parts of the paper do you envision would change...' question) for spontaneously mentioned update categories, and compare with the post-prompt portion. If the three categories appear only after the prompt, the typology is an artifact of the instrument. A second check: run a small validation study (e.g., 8-10 new respondents) with a protocol that omits the 'text, structure, figures, tables' phrasing and instead asks 'What, if anything, would you change in your survey, and why?' If empirical/structural/interpretive do not emerge unprompted, the central typology fails to replicate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution — the three update types (empirical, structural, interpretive) — closely tracks the wording of Appendix A's guide question: 'Which parts of the paper do you envision would change from these updates? For instance, could you speak to potential revisions regarding the existing text, structure, figures and tables, or other elements?' The prompt primes participants to think in terms of text/tables/figures versus structure, and the reported typology maps one-to-one onto those primed dimensions (empirical = text/tables/figures, structural = structure/organization, interpretive = synthesis/framing). Moreover, all data are hypothetical: no participant updated a survey during the study; they were asked to 'walk me through how you would update this paper.' The typology therefore describes intended, prompted future actions, not observed updating practice. Table 2 and the leverage-point discussion are built on this typology, so if the categories are instrument-driven the paper's main design implications weaken. The value/unmanageable/incentives finding is less affected, since it rests on direct statements, but the typology is the paper's primary transferable claim. Sample skew (7/11 PhD students, no senior faculty) further limits the generality of the incentives finding, though this is secondary.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a qualitative interview study of 11 corresponding authors of computing survey articles published in CSUR, CHI, and CSCW. Through semi-structured retrospective interviews, the authors examine how surveys are authored across search, appraisal, synthesis, and interpretation, and how authors envision updating them. The main findings are that authors value up-to-date surveys but find continuous updating unmanageable and misaligned with academic incentives; that envisioned updates fall into three types (empirical, structural, interpretative); and that authors are receptive to AI assistance for routine tasks but skeptical of AI for nuanced synthesis. The paper concludes with design implications for AI-assisted living narrative reviews.","tokens_in":16311,"tokens_out":4443,"duration_ms":41167,"significance":"The study addresses a real gap in the literature on living reviews, which has focused almost exclusively on systematic reviews. Its empirical grounding in interview data, with an explicit thematic-analysis codebook, direct participant quotes, and a full interview protocol in Appendix A, is a strength. The three-type typology, if robust, would provide a useful frame for system design. However, the small self-selected sample and the hypothetical, prompted nature of the updating data mean the findings should be treated as hypotheses about updating practices rather than as established facts. The paper is a useful empirical starting point but needs revision to support its strongest claims.","major_comments":[{"comment":"The interview guide asks participants, \"Which parts of the paper do you envision would change from these updates? For instance, could you speak to potential revisions regarding the existing text, structure, figures and tables, or other elements?\" This prompt strongly primes a partition of updates into text/tables/figures versus structure, and the reported typology of empirical versus structural versus interpretative updates maps almost one-to-one onto that partition, with interpretative serving as a residual category for synthesis and framing. The paper presents this typology as an emergent finding, but it may instead be a coding artifact of the interview instrument. This matters because Table 2 and the design implications in §5.2.1 are organized around the three types. To support the typology, the authors should either provide evidence that participants spontaneously articulated these categories before the prompt was introduced, or explicitly reframe the typology as an analyst-imposed coding scheme rather than a participant-derived discovery.","section":"Appendix A / §4.2"},{"comment":"The sample consists of 11 self-selected corresponding authors, of whom 7 are PhD students, 3 are assistant professors, and 1 is a research engineer; there are no senior or full professors, and all are based in the US or Canada. The claim that continuous updating is \"misaligned with academic incentives\" is stated as a general finding about computing research, but the incentive perceptions of PhD students and early-career researchers may not reflect those of tenured faculty who have different publication pressures and more control over their time. The paper should either restrict the claim to the studied population or add an explicit discussion of this sampling limitation in §5.3, which currently only mentions recall bias.","section":"§3.1 / §4.2"},{"comment":"The updating data are entirely hypothetical: participants were asked to \"walk me through how you would update this paper,\" and no participant actually updated a survey during the study. The findings in §4.2 on approaches and obstacles are therefore reports of intended, prompted actions rather than observed updating practice. The abstract and conclusion say the paper \"identifies three key types of updates for maintaining narrative reviews,\" which overstates the evidentiary status. The authors should revise the language to reflect that these are envisioned update types, and add the hypothetical nature of the update task to the limitations in §5.3.","section":"Appendix A / §4.2"}],"minor_comments":[{"comment":"In the full-text rendering, the title appears as \"Updati ng Survey Articles\" with a stray space; this should be corrected.","section":"Title"},{"comment":"The mean age and standard deviation appear as raw unicode escape sequences (\"/u1D440 = 32, /u1D446/u1D437 = 4.5\"); these should be typeset properly.","section":"§3.2"},{"comment":"The text refers to P4 as \"his own expertise,\" but Table 1 lists P4 as female; the pronoun should be corrected. The same sentence also contains a grammatical error: \"made it challenge to\" should be \"made it challenging to.\"","section":"§4.1 / Table 1"},{"comment":"The paper alternates between \"interpretative\" and \"interpretive\" (e.g., §4.2 uses \"interpretative\" while Table 2 uses \"Interpretative\" but the abstract uses \"interpretative\" and other sections use \"interpretive\"); while both are acceptable, the usage should be consistent.","section":"§4.2 / Table 2"}],"recommendation":"major_revision","confidential_remarks":"The main concern is whether the three-type typology is a genuine empirical finding or an artifact of the interview instrument. In my view this is fixable with a re-analysis or a reframing, so I recommend major revision rather than rejection. The paper is within scope for CHI and makes a useful empirical contribution, but its central transferable claim needs careful revision. The small sample and the absence of senior faculty should also be addressed in the limitations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful empirical study of an under-examined problem, and the core finding about incentives is solid. But the three-type update taxonomy the paper leads with should be read as provisional, because the interview guide likely nudged participants toward exactly those categories.\n\nWhat's new: prior work on living reviews is almost entirely about systematic reviews and clinical evidence updating. This is the first interview study I know that focuses on narrative/survey articles in computing. The method is standard but careful: purposive recruitment from CSUR/CHI/CSCW, 11 corresponding authors, thematic analysis with a two-coder open-coding step and explicit codebook. The four authoring activities (search, appraisal, synthesis, interpretation) and the finding that authors see continuous updating as valuable but unmanageable and unrewarded are well supported by the quotes. The discussion of AI as useful for routine automation but not trusted for interpretive synthesis matches recent empirical work and is a useful design constraint for tool builders.\n\nSoft spots: first, the empirical/structural/interpretive typology. In Appendix A, the 'Update Elements' prompt explicitly asks about changes to 'text, structure, figures and tables, or other elements.' That maps one-to-one onto empirical (text/tables/figures) and structural (structure), with interpretive the leftover 'other.' Participants did describe synthesis and framing changes, so the categories are not fabricated, but I would not treat this as an emergent grounded taxonomy. Second, everything about updating is hypothetical; nobody updated a survey during the study. That is fine for an exploratory study, but Table 2 and the leverage points are built on intended actions, not observed ones. Third, the sample skew: seven PhD students and no full professors means the incentive story is dominated by people at the most precarious career stage. That may understate the incentive problem, but it limits generality.\n\nNone of this is fatal. The authors acknowledge recall bias and the lack of observational data. The study is honest and internally consistent.\n\nWho this is for: researchers designing scholarly communication tools, HCI people working on AI-assisted writing, and anyone interested in publishing-practice reform. Worth a serious referee; the main revision request should be reframing the typology as a hypothesis and providing fuller coding tables or anonymized transcripts. Send to peer review.","headline":"A careful qualitative study whose three-type update taxonomy looks partly shaped by the interview prompt; the incentives finding is the sturdier contribution.","tokens_in":16806,"tokens_out":2598,"would_cite":true,"duration_ms":24245,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Survey authors value living reviews but find continuous updating unmanageable, and interviews with 11 computing researchers show updates cluster into three types.","keywords":["living literature reviews","narrative reviews","survey articles","computing research","qualitative interviews","AI-assisted authoring","scholarly communication","update process"],"falsifier":"A longitudinal or observational study would settle the claim: observe authors updating real survey articles over one to two years, recording live processes rather than recollections. If updates in practice do not cluster into the empirical, structural, and interpretative types, or if authors regularly update continuously when given modest institutional support, the paper's central generalizations would be weakened. A simpler check is a broad survey of survey authors asking whether they have ever published an update and what blocked them, looking for evidence that updating happens at scale without AI or incentive changes.","tokens_in":15925,"feed_emoji":"🔄","tokens_out":5383,"duration_ms":48456,"temperature":0.7,"pith_summary":"The paper reports on interviews with 11 authors of computing survey articles to learn how narrative reviews are created and updated. It claims that authors see keeping surveys current as valuable to the community but find continuous updating impractical because it is costly, triggers cascading revisions, and is not rewarded by academic incentives. The authors argue that updates to a survey fall into three types—empirical, structural, and interpretative—each with distinct challenges and opportunities for AI assistance. If the claim holds, it gives an empirical foundation for building AI tools that monitor new literature, locate where changes are needed, and draft local revisions while leaving interpretive judgment to human experts.","feed_headline":"Survey authors want living reviews but can't sustain them","feed_subtitle":"Interviews with 11 computing researchers show updates cluster into three types, with no reward for effort.","key_machinery":"The paper's central analytical object is a three-part taxonomy of survey updates—empirical, structural, and interpretative—derived from the interview data. The taxonomy organizes the range of modifications authors imagine making: updating quantitative and qualitative evidence (empirical), revising organization and taxonomies (structural), and reinterpreting the synthesis and narrative framing (interpretative). The taxonomy carries the argument because it links each update type to specific leverage points for AI assistance, from recalculating values to proposing new section divisions, while explaining why interpretive synthesis is seen as the hard core that AI should only support, not replace.","core_discovery":"The central discovery is that computing researchers who write narrative survey articles believe in the ideal of the 'living review' but do not practice it, because continuous updating is both a time cost and a career cost. Through thematic analysis of retrospective interviews, the paper identifies three kinds of updates that authors envision: empirical updates that add or revise evidence, numbers, tables, and examples; structural updates that reorganize taxonomies, sections, and frameworks; and interpretative updates that re-synthesize findings, limitations, and future directions. A key finding is that authors expect the original survey's organizational structure to stay stable for years, so most update effort concentrates on empirical and interpretative work, and that adding even one paper can force a cascade of revisions across text, figures, and conclusions. The authors conclude that the 'unmanageable' nature of this process, plus the absence of academic credit for updates, is the main obstacle to living narrative reviews, and they suggest AI could reduce routine costs but not replace the expert synthesis that gives surveys their value.","pith_inferences":["My inference: the three update types likely generalize beyond computing surveys to narrative reviews in other fields that use semi-systematic methods, such as management, environmental science, or education, so the taxonomy could serve as a shared vocabulary across disciplines.","My inference: the paper's finding that authors fear AI would make surveys 'formulaic' suggests a design constraint: AI update assistants should aim at invisible consistency (numbers, references, organization) rather than producing prose that reads as template-generated, preserving the author's voice and the pleasure of reading.","My inference: a testable extension would be to instrument a living-review authoring environment to log actual update events and compare them against the taxonomy, converting the interview-derived categories into a measurable annotation schema.","My inference: the study implies that the unit of 'update' in citation counting and review metrics may need to be redefined; if updates were citable and peer-reviewed as contributions, the incentive barrier the authors identify could be partially dismantled."],"forward_implications":["If the three-type taxonomy holds, future 'living narrative review' tools can be structured by update type: automated recalculation and citation refresh for empirical updates, clustering and section-splitting suggestions for structural updates, and bias and gap detection with the author in the loop for interpretative updates.","Designers of AI support should expect authors to delegate only low-cost, verifiable tasks to automation, while tasks where mistakes cascade into the argument's conclusions will be kept under human control.","Because authors anticipate reusing original taxonomies and workflows, tools that preserve and reuse codebooks, search strings, and scripts from the original survey could lower the cost of updates more than general-purpose summarization.","Peer review and evaluation practices that count articles only at first publication will need to change if living narrative reviews are to become common; technology alone will not solve the incentive mismatch."],"supporting_citations":[{"why":"Supplies the survival-analysis evidence that reviews go out of date within a few years, motivating the need for updates.","marker":"[29]"},{"why":"Introduces the concept of living reviews as documents continually updated with new evidence.","marker":"[8]"},{"why":"Shows that many living systematic reviews are never updated after initial publication, framing the maintenance problem.","marker":"[11]"},{"why":"Provides the review typology that distinguishes systematic, semi-systematic, and integrative reviews, positioning narrative surveys.","marker":"[31]"},{"why":"Describes the thematic analysis methodology used to code and interpret the interview data.","marker":"[3]"},{"why":"Supports the finding that scholars are receptive to AI for narrow objective tasks but hesitant about subjective creative ones.","marker":"[23]"},{"why":"Describes combining human and machine effort for living systematic reviews, giving a prior point of comparison.","marker":"[33]"}],"fun_headline_variants":["Living surveys are a noble ideal researchers can't sustain","Three update types, zero incentive: why surveys go stale","Adding one paper to a survey forces a cascade of edits","Researchers want living reviews but don't get credit for upkeep","The unmanageable dream of the living narrative review"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's findings stand on the assumption that what 11 researchers said in retrospective interviews about their authoring and updating practices matches what they actually did, so the reported update types and obstacles reflect real workflows rather than memory or rationalization.","fun_headline_variants_meta":{"raw":{"variants":["Living surveys are a noble ideal researchers can't sustain","Three update types, zero incentive: why surveys go stale","Adding one paper to a survey forces a cascade of edits","Researchers want living reviews but don't get credit for upkeep","The unmanageable dream of the living narrative review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000563,"raw_usage":{"total_tokens":2646,"prompt_tokens":893,"completion_tokens":1753,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":1674}},"tokens_in":509,"tokens_out":1753,"duration_ms":10800,"temperature":1.0,"reasoning_tokens":1674,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T17:21:21.474372+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A longitudinal or observational study would settle the claim: observe authors updating real survey articles over one to two years, recording live processes rather than recollections. If updates in practice do not cluster into the empirical, structural, and interpretative types, or if authors regularly update continuously when given modest institutional support, the paper's central generalizations would be weakened. A simpler check is a broad survey of survey authors asking whether they have ever published an update and what blocked them, looking for evidence that updating happens at scale without AI or incentive changes.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that many living systematic reviews are never updated after initial publication, framing the maintenance problem."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes combining human and machine effort for living systematic reviews, giving a prior point of comparison."}],"review_version":1}