{"id":"3c8b633b-5775-45d3-b121-c83fb7bdc5c1","arxiv_id":"2509.07819","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Wikipedia editors who used LLMs reported that experienced editors expanded their contributions, while newcomers were pushed into editorial judgment they lacked skills for and saw their LLM-assisted edits rejected.","lead":"This paper interviews 16 Wikipedia editors who have used large language models and finds an expertise-based divide: experienced editors gain confidence and quality, while newcomers face lower entry barriers but higher demands and frequent rejection of their edits. The study challenges standard theories of how newcomers learn to participate in online knowledge communities.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Community-response half of the divide rests on self-reports with no non-LLM baseline; observed newcomer rejection may be ordinary gatekeeping, and Section 5.2 concedes newcomers would skip peripheral tasks regardless.","rationale":"The reader identified the self-selected, self-reported sample as the weakest assumption. That is a real concern, and my analysis agrees with it, which is why I would keep the verdict at CONDITIONAL. However, I believe the more precise load-bearing issue is not only who was interviewed but what was measured: the paper makes a causal claim about how 'other editors' respond to LLM-assisted edits, yet the data are second-hand reports from the contributors themselves, with no baseline from non-LLM newcomers. Because newcomer rejection is a well-documented pre-existing phenomenon on Wikipedia, the absence of a comparison group leaves the central divide potentially confounded. The paper's own Section 5.3 limitation statement calls for future work including non-users' perspectives, and Section 5.2's closing paragraph explicitly questions whether LLM use causes newcomers to bypass peripheral tasks, further limiting the causal interpretation. These manuscript passages support the concern and are in-scope evidence. The proposed log-analysis test would settle whether the community-response asymmetry survives direct measurement; until such evidence exists, the paper should be read as a rich qualitative hypothesis-generating study rather than a demonstrated causal account. I therefore do not change the reader's verdict.","tokens_in":17903,"tokens_out":8419,"duration_ms":74432,"concrete_test":"Use Wikipedia's public revision and talk-page histories to build a matched sample of new-article creations or first edits by users with less than 90 days tenure, split into edits flagged as likely LLM-generated (e.g., via the stylistic detector of Brooks et al. 2024, cited as [10]) and a matched non-flagged or pre-LLM group. For each edit, record whether it was reverted or removed and code the associated talk-page message for AI/LLM mentions or quality criticism, stratified by editor tenure. If LLM-flagged newcomer edits are no more likely to be reverted or challenged than non-flagged newcomer edits, the community-response divide reported in Section 4.3 is not specific to LLM use; if an LLM-specific penalty appears only for newcomers and not for experienced editors, the central claim is strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that LLM use creates an expertise-based participation divide depends on three linked findings, and the most load-bearing is the community-response asymmetry in Section 4.3: other editors are said to reject LLM-assisted edits from newcomers while approving those from experienced editors. The evidence for this third-party behavior consists entirely of retrospective self-reports from the 16 LLM-using interviewees; no talk-page logs, reversion statistics, or interviews with the responding editors are provided. This matters because Wikipedia research has long shown that newcomer edits, especially new-article creations, are reverted at high rates even in the absence of LLMs (Halfaker et al. 2013, cited as [34]). Without a baseline of non-LLM-using newcomers, the rejections reported by P11, P01, P10, and P15 could be standard newcomer gatekeeping rather than an LLM-specific 'participation paradox.' The paper's Section 5.3 acknowledges the absence of non-users' perspectives and the risk of adoption bias, but it does not address the missing behavioral baseline for community responses. A further internal tension appears in the final paragraph of Section 5.2, which concedes that 'newcomers wanted to create articles and LLMs may just be a tool they utilize to achieve their goals. If not LLMs, they would still start contributing complex tasks from day one'; this undercuts the causal reading in Section 5.1.1 that LLMs interrupt the legitimate peripheral participation pathway. Taken together, the evidence supports an exploratory, experience-based account of a divide, but not the stronger claim that LLMs create it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative interview study of 16 Wikipedia editors who have used large language models in their editing work. The authors ask how LLMs affect content contribution (RQ1), what strategies editors use to align LLM output with community norms (RQ2), and how other editors respond to LLM-assisted contributions (RQ3). The central finding is an expertise-based participation divide: experienced editors use LLMs to explore new topics, gain confidence, and improve editing quality, while newcomers experience a 'participation paradox' in which LLMs lower entry barriers but simultaneously demand editorial judgment that newcomers have not yet developed, leading to community rejection. The authors interpret this through Legitimate Peripheral Participation (LPP) and situated learning, arguing that LLM use interrupts newcomers' gradual learning pathways and shifts them from peripheral tasks to high-stakes editorial judgment. They close with design implications for scaffolding, teaching community norms, and expertise-aware LLM interaction, and they acknowledge in Section 5.3 that prevalence was not quantified and that only LLM-using editors were interviewed.","tokens_in":18187,"tokens_out":3751,"duration_ms":35191,"significance":"If the findings hold, the paper makes a useful contribution to CSCW/HCI scholarship on human-AI collaboration and peer production. It articulates a specific mechanism, the participation paradox, that goes beyond 'LLMs help some and hurt others' by linking the divide to a mismatch between the tacit knowledge required for editorial judgment and the capabilities of newcomers. The extension of LPP to a setting where the tool itself changes the trajectory of participation is a plausible and interesting theoretical move. Methodologically, the paper has strengths: the interviews are quoted extensively, the thematic analysis procedure is standard, the coding process is described transparently, and the authors honestly concede the main limitations of self-report and sample composition. The design implications are concrete and grounded in the data. However, the central empirical claim about differential community responses relies on retrospective self-reports from a self-selected sample, and one passage in the Discussion appears to undercut the causal reading of the LPP argument.","major_comments":[{"comment":"The community-response asymmetry, namely that other editors reject LLM-assisted edits from newcomers but approve those from experienced editors, is supported only by retrospective self-reports from the 16 LLM-using interviewees. No talk-page logs, reversion statistics, or interviews with the responding editors are provided. This matters because the paper itself cites Halfaker et al. [34] showing that newcomer edits, especially new-article creations, are reverted at high rates even without LLMs. The rejections reported by P11, P01, P10, and P15 could therefore reflect ordinary newcomer gatekeeping rather than an LLM-specific 'participation paradox.' I recommend either (a) obtaining behavioral data, for example comparing reversion/rejection rates for LLM-assisted versus non-LLM-assisted newcomer edits, or (b) explicitly reframing the finding as participants' perceptions of community response rather than as the community's actual response. The acknowledgment in Section 5.3 that non-users were excluded does not address this missing behavioral baseline.","section":"Section 4.3 / RQ3"},{"comment":"The final paragraph of Section 5.2 states that 'newcomers wanted to create articles and LLMs may just be a tool they utilize to achieve their goals. If not LLMs, they would still start contributing complex tasks from day one.' This concession is in direct tension with the causal claim in Section 5.1.1 that LLMs interrupt the LPP trajectory by enabling newcomers to skip peripheral tasks. If newcomers would attempt complex tasks regardless of LLM availability, then the interruption of gradual learning is not caused by LLM use, and the claim that LLMs 'shortcut' situated learning is overstated. The paper needs to separate two possibilities: LLMs enabling a pre-existing preference for complex tasks, versus LLMs creating that preference. As written, the causal framing in the Discussion is stronger than the evidence and the paper's own closing caveat allow.","section":"Section 5.2 (final paragraph), with Section 5.1.1"},{"comment":"The expertise-based divide that organizes the results is never operationalized. The paper frequently contrasts 'newcomers' with 'experienced editors,' but Table 1 shows that P09 has a 0-2 year tenure while making 5.3K+ edits daily, P10 has 2-5 years of tenure with monthly edits, and P14 and P15 have 0-2 years with 200+ edits. These participants are quoted as newcomers in the results, yet by edit count some of them are highly active. The paper should define the newcomer/experienced boundary (tenure, edit count, or a combination) and show that the qualitative patterns are stable under that definition. Without this, the central 'participation divide' claim is vulnerable to circular categorization.","section":"Section 3.1 / Table 1"}],"minor_comments":[{"comment":"There is a duplicated word in the sentence beginning 'However, However, even experienced editors occasionally faced false accusations...'.","section":"Section 4.3.2"},{"comment":"The word 'satisfication' appears in the sentence 'This cognitive support, in turn, increased confidence and satisfication for experienced editors' and should be 'satisfaction'.","section":"Section 4.1.3"},{"comment":"The caption reads 'Participant summary, adopted from [75].' If the table is adopted from a prior paper, the source should be clarified; if it is original to this study, the wording should be corrected.","section":"Table 1 caption"},{"comment":"The paper alternates between 'WikiMedia project page' and 'Wikimedia'; please standardize the spelling.","section":"Section 3.1"},{"comment":"There are duplicate reference entries for the Wikipedia policy pages: NPOV appears as [85] and [86], Verifiability as [89] and [90], and Notability as [88]. These should be consolidated.","section":"References"},{"comment":"The term 'WikiCotent' appears in the sentence describing super-labels and should be 'WikiContent' (or similar).","section":"Section 2.2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for CSCW and the qualitative data are reported in good faith, but the authors should be pushed to align the Discussion's causal claims with the evidence. The missing baseline for newcomer rejection and the internal tension in Section 5.2 are fixable through reframing and additional analysis; I would not recommend rejection, but the current manuscript overstates the community-response finding relative to what the data can support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a worthwhile qualitative study, and the 'participation paradox' framing is genuinely useful, but the paper's strongest causal phrasing runs ahead of its evidence—and the authors themselves concede part of the gap in Section 5.2.\n\nWhat's new: this appears to be the first interview study of active Wikipedia editors who use LLMs. The finding that LLM assistance helps experienced editors expand their contributions while pushing newcomers into high-stakes editorial judgment they aren't ready for is a valuable synthesis of prior work on bots, newcomer socialization, and the second-level digital divide. The quotes are well chosen, the thematic analysis is standard for CSCW, and the design implications (scaffolding, teaching norms, expertise-aware responses) follow from the data. Credit where due: the authors explicitly flag in Section 5.3 that they didn't quantify prevalence and didn't include non-users.\n\nWhere it gets soft: the community-response half of the divide (Section 4.3) rests entirely on self-reports from the 16 LLM-using interviewees. There are no talk-page logs, reversion statistics, or interviews with the editors doing the rejecting. Given the well-documented gatekeeping against newcomers in Wikipedia (Halfaker et al. 2013, which they cite), the rejections could be ordinary newcomer friction rather than an LLM-specific effect. The paper doesn't provide a baseline. More damaging, the final paragraph of Section 5.2 concedes that newcomers would 'still start contributing complex tasks from day one' without LLMs—which undercuts the causal reading in 5.1 that LLMs interrupt the LPP trajectory. If newcomers bypass peripheral tasks anyway, LLMs may accelerate a pre-existing tendency, not create it.\n\nThat said, this doesn't sink the paper. The cautious read—that LLM assistance may widen an existing expertise gap, and that newcomers face real difficulties judging AI output—is well supported by the interview data. The authors are honest about the limits, and the theory application (LPP as lens, not tautology) is reasonable.\n\nWho it's for: anyone working on AI-mediated participation in peer production, Wikipedia governance, or newcomer onboarding. I'd send this to referees with a request to push for tempered causal claims and ideally a follow-up with behavioral data. It deserves peer review, not desk rejection.","headline":"A timely, honest qualitative study of LLM use by Wikipedia editors; the 'participation paradox' is a useful framing, but the causal claim about community responses rests on self-reports and is partly conceded in Section 5.2.","tokens_in":18702,"tokens_out":3004,"would_cite":true,"duration_ms":24564,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM use on Wikipedia helps expert editors thrive while pushing newcomers into a participation paradox that gets their edits rejected.","keywords":["large language models","Wikipedia","knowledge communities","participation divide","newcomer onboarding","legitimate peripheral participation","human-AI collaboration","qualitative interviews"],"falsifier":"A quantitative audit of Wikipedia's edit history: flag contributions whose edit summaries or text indicate LLM use, split by editor tenure, and compare reversion rates. If newcomers' LLM-assisted edits are reverted at the same rate as experienced editors' once content quality is controlled, the claimed expertise-based participation divide would not be supported.","tokens_in":17719,"feed_emoji":"🤖","tokens_out":4664,"duration_ms":40062,"temperature":0.7,"pith_summary":"This paper asks what happens to participation in a knowledge community when members use large language models to write and edit. Drawing on interviews with 16 Wikipedia editors who used LLMs, it argues that the outcome splits along expertise: experienced editors use LLMs to explore new topics, check their work, and earn praise, while newcomers find that the same tools lower the barrier to entry but raise the bar for what counts as acceptable participation. The key claim is a 'participation paradox': LLMs let newcomers attempt central tasks such as drafting articles before they have the editorial judgment those tasks require, and the community responds by rejecting their edits. If correct, the paper implies that simply giving people access to AI writing tools does not democratize knowledge production; it can reproduce or even widen existing inequalities.","feed_headline":"LLMs widen Wikipedia's gap between new and expert editors","feed_subtitle":"16 editor interviews show LLMs lower entry barriers but demand judgment newcomers lack, so their edits get rejected.","key_machinery":"The load-bearing mechanism is the 'participation paradox': LLMs simultaneously lower barriers to entry and raise the demands of participation by requiring editorial judgment that newcomers lack. The paper frames this through legitimate peripheral participation, the classic model in which newcomers start with peripheral tasks and gradually move toward central responsibilities while learning community norms through social interaction. The paper also identifies three concrete judgment strategies editors must apply to LLM output - evaluation, verification, and modification - and argues that these demand tacit knowledge about Wikipedia's policies and culture that newcomers have not yet acquired. This combination - interrupted learning pathways, a shift from peripheral tasks to editorial judgment, and community sensitivity to AI origin - is what produces the expertise-based divide.","core_discovery":"The paper's central claim is that LLM use in Wikipedia creates a participation divide mediated by expertise. Experienced editors enhance their participation: they venture into new topics, treat LLMs as sources of fresh perspectives, gain confidence in unfamiliar areas, and receive positive responses from other editors. Newcomers, by contrast, face a paradox: LLMs lower the barriers to entry, but they also demand that newcomers act as editorial judges who evaluate, verify, and modify AI output before publishing it, a role that presupposes the very wiki knowledge newcomers have not yet developed. As a result, LLM-assisted contributions from newcomers are more likely to be flagged, criticized, and reverted. The authors argue that this dynamic breaks the traditional trajectory of legitimate peripheral participation, in which newcomers earn legitimacy by starting with small, low-risk tasks and gradually absorbing community norms; LLMs skip that pathway and push newcomers straight into central, high-stakes judgment.","pith_inferences":["If the divide is real, edit-reversion data should show a measurable tenure-by-LLM interaction: newcomers' LLM-flagged edits are reverted more often and faster than experienced editors' LLM-flagged edits, which is testable in Wikipedia's public revision history.","The design proposals imply a concrete test: an LLM assistant that decomposes article creation into source-finding, summarizing, and structuring steps should improve newcomer retention and edit acceptance relative to a direct-draft assistant in a randomized field experiment.","The paper's logic also predicts that community norms will harden around process rather than outcome: editors will be sanctioned for using LLMs without demonstrated review, regardless of content quality, which could suppress good-faith contributions from non-native speakers who rely on LLMs for language support.","Because the sample is self-selected LLM adopters, the same mechanisms might look different for editors who tried LLMs and quit; surveying that group could separate tool effects from adopter effects."],"forward_implications":["LLM-assisted edits from newcomers will continue to be flagged and reverted quickly unless tools or norms change, because the community treats AI origin as a signal of low quality.","Experienced editors will keep gaining more from LLM assistance, including broader topics, faster drafting, and higher confidence, which widens the gap between them and newcomers.","LLM tools that simply generate finished articles are the wrong design; assistants should scaffold tasks, teach community norms, and adapt to the editor's expertise level.","Wikipedia's norm-formation problem is urgent: the absence of shared rules leaves individual editors to improvise, and the resulting ambiguity feeds mistrust of all LLM-assisted work.","If LLMs are integrated without attention to learning, they can undermine the social processes that have sustained knowledge communities, not just the quality of individual edits."],"supporting_citations":[{"why":"Supplies the account of newcomers' gradual path from peripheral tasks to fuller participation that the paper argues LLMs interrupt.","marker":"[11]"},{"why":"Grounds the situated-learning model whose learning pathway LLM use is claimed to short-circuit.","marker":"[49]"},{"why":"Documents how Wikipedia's quality control rejects good-faith newcomers, the baseline pattern the paper extends.","marker":"[34]"},{"why":"Provides the access-versus-skills distinction the paper uses to frame the participation divide.","marker":"[37]"},{"why":"Shows LLM-generated content is already entering Wikipedia articles, motivating the study.","marker":"[10]"},{"why":"Gives prior evidence that LLMs misapply Wikipedia neutrality norms, supporting participants' NPOV concerns.","marker":"[3]"},{"why":"Frames peripheral participation as a route to legitimacy that newcomers are argued to bypass.","marker":"[35]"},{"why":"Serves as an example of a prior norm-aligned tool whose design contrasts with LLMs delegating judgment to editors.","marker":"[18]"}],"fun_headline_variants":["LLMs boost Wikipedia experts but trip up newcomers","Wikipedia newcomers face LLM judgment gap","LLM edits: experts win, newcomers reverted","AI divides Wikipedia: experts thrive, novices fail"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The strongest assumption is that the 16 volunteer interviewees' recollections accurately represent how Wikipedia editors generally use LLMs and how the community reacts; if the sample skews toward enthusiastic adopters or particular editor types, the reported divide may be an artifact of who agreed to talk.","fun_headline_variants_meta":{"raw":{"variants":["LLMs boost Wikipedia experts but trip up newcomers","Wikipedia newcomers face LLM judgment gap","LLM edits: experts win, newcomers reverted","AI divides Wikipedia: experts thrive, novices fail"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1390,"prompt_tokens":966,"completion_tokens":424,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":365}},"tokens_in":582,"tokens_out":424,"duration_ms":4293,"temperature":1.0,"reasoning_tokens":365,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:10:07.628074+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A quantitative audit of Wikipedia's edit history: flag contributions whose edit summaries or text indicate LLM use, split by editor tenure, and compare reversion rates. If newcomers' LLM-assisted edits are reverted at the same rate as experienced editors' once content quality is controlled, the claimed expertise-based participation divide would not be supported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents how Wikipedia's quality control rejects good-faith newcomers, the baseline pattern the paper extends."},{"cited_title":"Second-Level Digital Divide: Mapping Differences in People's Online Skills","cited_arxiv_id":"cs/0109068","evidence_quote":"Provides the access-versus-skills distinction the paper uses to frame the participation divide."},{"cited_title":"Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms","cited_arxiv_id":"2407.04183","evidence_quote":"Gives prior evidence that LLMs misapply Wikipedia neutrality norms, supporting participants' NPOV concerns."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames peripheral participation as a route to legitimacy that newcomers are argued to bypass."}],"review_version":2}