{"id":"0b110b25-8952-45a4-a142-a2fe6b1cecc6","arxiv_id":"2411.13032","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Writers define authenticity in human-AI co-writing through their internal experiences and creative process, and personalized AI should target writer growth beyond text production; readers respond positively and largely cannot tell AI-assisted from solo work.","lead":"The paper interviewed 19 professional writers co-writing with personalized and generic AI tools, plus 30 readers, to understand how people think about authenticity in AI-assisted writing. It argues that writers care most about the internal process of expressing their authentic selves, and that personalization should support writer growth, not just match text style.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing manipulation check for the personalization condition: writers were told which tool was personalized only after both sessions, so the Section 4.4 preferences may reflect labels or demand characteristics rather than actual voice-matching; re-analyzing the writing logs with style-similarity…","rationale":"The reader's weakest_assumption identifies the same load-bearing concern I would raise: the personalization condition lacks a check that writers perceived the personalized tool as reflecting their own voice. This is the most consequential gap because the paper's RQ3 conclusions and its design implications for personalized AI writing tools depend on the comparison between personalized and non-personalized conditions, and Section 4.4.2 reports no behavioral difference between them. Without a manipulation check, the preference results could reflect demand characteristics rather than the actual effect of personalization. I also considered the abstract's statement that readers 'could not distinguish' AI-assisted work from solo work. The paper's own Table 7 shows significant differences in likelihood-of-human-writing ratings (solo vs. personalized AI: beta = -0.73, p = 0.004; solo vs. non-personalized AI: beta = -0.64, p = 0.012 in Round 2), so the abstract overstates the null result. However, that is primarily a wording issue in the reader-facing summary and does not directly threaten the writer-centered conceptual contribution. The missing manipulation check is more load-bearing because it undermines an entire research question and the associated design claims. I agree with the reader's assessment and would keep the conditional verdict: the qualitative findings are credible, but the personalization claims need either a manipulation check or explicit softening.","tokens_in":34525,"tokens_out":4834,"duration_ms":49746,"concrete_test":"Re-analyze the existing CoAuthor writing logs from Appendix D: for each writer, compute a style-similarity score (e.g., embedding cosine similarity or n-gram/stylometric overlap) between the writer's 200-word pre-study sample and the suggestions accepted in the personalized versus non-personalized condition, then test whether similarity is significantly higher in the personalized condition. If it is not, the personalization manipulation was too weak to support the Section 4.4 preference claims.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2.3 builds the personalized tool via in-context learning from one ~200-word writing sample collected in the pre-interview survey (Section 3.2.2), but the paper never verifies that writers actually experienced the personalized tool as reflecting their own voice. The debrief questions in Appendix B.3 ask whether the tool picked up the writer's voice only after the personalization was revealed, and no pre-reveal rating of perceived stylistic closeness is reported. Section 4.4.2 shows no significant behavioral difference between conditions in suggestion request frequency or acceptance rate, so the only evidence for a personalization effect is post-hoc self-report collected after condition labels were known. Without a manipulation check, the preference results in Section 4.4.1 and the design implication that personalization should target writer growth rather than text production are not firmly tied to the personalization manipulation; they could be driven by expectation, the label, or the fact that the writer's own sample was in the prompt. This is load-bearing because RQ3 is explicitly about whether personalization can support authenticity, and the paper's contributions include claims about personalized AI writing tools. The qualitative core of the paper is otherwise well supported: the interview protocol, codebook, and extensive quotes provide transparent grounding for the process-oriented authenticity theme, and the writing-log data add behavioral detail. The particular vulnerability is the link between the manipulation and the reported preferences.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper examines how professional writers and avid readers conceptualize authenticity in human-AI co-writing, and whether personalization of AI writing assistance supports writers' authentic voice. Part 1 uses semi-structured interviews with 19 professional writers, who each completed two writing sessions with a GPT-4-based CoAuthor interface, one personalized via in-context learning from a 200-word writing sample and one non-personalized. Part 2 is an online survey with 30 avid readers who rated and compared solo-written, personalized-AI, and non-personalized-AI passages produced by six of the writers. The paper reports three main contributions: a writer-centered definition of authenticity emphasizing internal experiences and process, design implications for co-writing tools targeting writer growth, and an account of readers' largely positive and undifferentiating responses to AI-assisted creative writing.","tokens_in":34788,"tokens_out":6029,"duration_ms":55582,"significance":"If the findings hold, the paper makes a useful contribution to CSCW/HCI by grounding authenticity in writers' lived experiences and by distinguishing process-oriented authenticity from the category/source/value frameworks in prior literature. The qualitative core is a strength: the study provides a transparent interview protocol, a detailed codebook, a participant table, and extensive verbatim quotes, giving readers direct access to the evidence behind the thematic claims. The mixed-methods design, combining writing logs with interviews and a reader survey, adds behavioral texture. The authors also showed methodological care in screening AI-generated survey responses. The central qualitative claim about process-oriented authenticity is well supported by the interview data; the main weaknesses concern the strength of the claims about reader discrimination and the lack of a manipulation check for the personalization condition.","major_comments":[{"comment":"The abstract states that \"Readers could not distinguish AI-assisted work, personalized or not, from writers' solo-written work,\" but the paper's own statistical models contradict this. In Table 7 (Round 2), the likelihood-of-human-writing rating differs significantly for solo vs. personalized AI (β = −0.73, SE = 0.25, t = −2.94, p = 0.004) and for solo vs. non-personalized AI (β = −0.64, SE = 0.24, t = −2.57, p = 0.012). Section 4.4.3 correctly reports that reader participants rated the solo human work as more likely to have been written independently. The abstract and the concluding section therefore overstate the finding. The supported claim is that likeability, enjoyment, and creativity ratings did not differ significantly across conditions and readers could not reliably identify which text segments were AI-generated, while judgments of how likely a piece is to have been written by a human did distinguish solo from AI-assisted work. Please revise the abstract, Section 5 discussion, and Section 6 conclusion to align with the actual statistical results.","section":"Abstract; Section 4.4.3; Table 7"},{"comment":"The paper lacks a manipulation check for the personalization condition. The personalized tool was built by in-context learning from a single ~200-word writing sample, but no pre-reveal measure verifies that writers actually experienced this tool as reflecting their own voice. The only direct question (\"Does the personalized AI tool pick up your unique voice, tone, style, etc. in writing?\") appears in Appendix B.3, after the experimenter disclosed which of the two tools was personalized. Behavioral logs (Section 4.4.2) show no significant difference between conditions in the frequency of requesting AI assistance or in the rate of accepting suggestions. As a result, the subjective preference findings in Section 4.4.1 and the design implication in Section 5.2 that personalization should support writers' growth rather than merely text production are not firmly tied to the personalization manipulation; they could reflect the condition label, demand characteristics, or the mere presence of the writer's own sample in the prompt. Because RQ3 asks whether personalization can support authenticity, this gap is load-bearing. Please add an appropriate manipulation check (e.g., a pre-reveal rating of perceived stylistic closeness, or a style-similarity analysis of the writing logs using the authors' own writing samples) and temper the causal language in Sections 4.4 and 5.2 accordingly.","section":"Section 3.2.3; Section 4.4.2; Appendix B.3"}],"minor_comments":[{"comment":"In the Creativity row for round 1 (\"Solo vs. Personalized AI\"), the reported values are β = 0.91, SE = 0.25, t = 0.01, p = 0.999. A t value of 0.01 is inconsistent with β/SE ≈ 3.6, and the p value also does not correspond to that t. Please correct the entry and re-verify the reported statistics.","section":"Table 7"},{"comment":"The survey question in G.2 reads \"How much do you creative the writing?\" and should be \"How creative do you find the writing?\".","section":"Appendix G"},{"comment":"The text says \"the majority of writers preferred working with personalized writing tools\" but does not state the number or proportion; given the quantifier convention defined in Section 4.1 (a few, some, most, nearly all), please report the explicit count for this claim.","section":"Section 4.4.1"},{"comment":"The statement \"the rate of correct identification has no significant difference between the personalized and non-personalized conditions\" reports only the absence of a difference between conditions; please also report the actual rates and, if available, compare them to chance to support the claim that readers could not identify AI-assisted segments.","section":"Section 4.4.3"},{"comment":"In the codebook, the code \"Expedite internationalization\" appears; the intended word is presumably \"Expedite internalization.\"","section":"Appendix C"},{"comment":"The relationship among the three data columns (\"Round 1\", \"Round 1 (removed 7 responses)\", and \"Round 2\") is unclear: it is not stated which column corresponds to the primary analysis reported in the text, what the sample size of each column is, and why multiple columns are presented as co-equal. Please clarify the analysis pipeline and designate the primary dataset.","section":"Section 3.3.2 and Table 7"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the venue and the qualitative findings are likely to be of interest to the CSCW/HCI community. The two major issues are fixable in a revision: aligning the abstract with the reported statistics and adding a manipulation check for personalization. I would not recommend rejection, as the core qualitative contribution is solid and the missing manipulation test does not invalidate the interview-based themes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the thing to know: the conceptual heart of this paper is real. The writers' own definitions of authenticity—centered on internal experience, the authentic self, and the process of constructing that self—extend the standard source/category/value frameworks in a useful way. The qualitative methods are transparent: interview protocol, codebook, participant table, and extensive quotes all make the thematic claims checkable. That part deserves serious engagement.\n\nThe personalization comparison is the soft spot. The personalized condition was built from one ~200-word writing sample via in-context learning, and there is no manipulation check showing writers actually experienced the tool as voice-matched. The debrief questions about picking up the writer's voice come after the conditions were revealed, so the Section 4.4.1 preferences may reflect the label or demand characteristics rather than genuine personalization. That matters because RQ3 is explicitly about personalization, and the paper's design implications lean on it. A pre-reveal rating of perceived stylistic closeness, or a content-based similarity check on the writing logs, would firm it up. The stress-test note has this right.\n\nAlso, the abstract overstates the reader results. Table 7 shows readers rated solo work as significantly more likely to be human-written than either AI-assisted condition. That is not 'could not distinguish.' The qualitative point about readers not caring much and reacting positively is still supported, but the sentence needs correcting.\n\nMinor: the reader survey is small (N=30) with many unadjusted tests. The authors acknowledge the sample limitations; it doesn't sink anything, but it's another reason to keep the claims close to the qualitative data.\n\nBottom line: this is a useful paper for HCI/CSCW researchers working on AI writing support. The writer-centered authenticity concept and the design implication (support writer growth, not just text production) are worth taking seriously. It deserves peer review, not a desk reject. I'd send it to referees with a request to add a manipulation check or soften the personalization claims, and to fix the abstract.","headline":"Solid qualitative core on writer-centered authenticity, but the personalization finding lacks a manipulation check and the abstract overstates the reader data.","tokens_in":35314,"tokens_out":2346,"would_cite":true,"duration_ms":22857,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Writers see authenticity as process, not output, when co-writing with AI.","keywords":["authenticity","human-AI co-writing","large language models","personalization","professional writers","reader perceptions","content gatekeeping","AI writing assistance"],"falsifier":"A replication with a manipulation check, asking writers after each session how much the suggestions sounded like their own voice, would settle the personalization claims: if writers rate the personalized tool no higher than the non-personalized one, the preference and null behavior findings cannot be attributed to personalization; if, with stronger personalization, writers' acceptance rates rise and readers can distinguish the output, the paper's claim that personalization does not change behavior or reader perception would fail.","tokens_in":34324,"feed_emoji":"✍️","tokens_out":5321,"duration_ms":51615,"temperature":0.7,"pith_summary":"The paper argues that authenticity in human-AI co-writing is primarily a process of constructing and expressing the writer's authentic self, not a property of the finished text. Based on interviews with 19 professional writers who co-wrote passages with personalized and non-personalized large language model tools, it finds writers define authenticity through the source of content, their internal experiences and identities, and the writing outcomes, with the internal-experience facet receiving the most emphasis. Writers preserve authenticity by starting from a clear vision, using AI in the middle \"fuzzy\" stage of idea development, and acting as content gatekeepers who decide what enters the text. The paper further claims that personalization built from a 200-word writing sample is preferred by writers but changes neither their requesting and accepting behavior nor readers' ability to distinguish the results, and that readers react positively to writers experimenting with AI.","feed_headline":"Writers see authenticity as process, not output, when co-writing with AI","feed_subtitle":"Writers guard authenticity through judgment, not word count—and readers can't tell AI-assisted work apart from solo writing.","key_machinery":"The central object is the writer-centered conception of authenticity, a three-part construct of source authenticity, authentic self, and content authenticity, in which the second part, authentic self, defined as the writer's lived experience, identity, passion, and autonomy expressed through the process of writing, carries most of the weight. The argument runs on content gatekeeping: the practice by which writers decide what goes into the text, which the paper identifies as the main way writers claim contribution and preserve authenticity when co-writing with AI. The study's comparative machinery is a within-subject writing session using the CoAuthor interface with GPT-4, personalized by in-context learning from each writer's 200-word sample versus a non-personalized version, followed by a reader survey that assessed likeability, enjoyment, creativity, attribution, and perceived authenticity.","core_discovery":"On the paper's own terms, the discovery is that writers' conception of authenticity extends beyond the three classic themes of category, source, and value to a fourth, dominant theme: the authentic self. Writers treat authenticity as grounded in lived experience, emotions, passion, autonomy, and the ability to justify putting one's name on the work; they see the work as authentic when it could only have been written by them. Co-writing with AI does not change this definition, but it changes how writers can protect it: the practices that matter include holding a clear vision before writing, calling on AI mainly in the \"fuzzy area\" between an initial idea and a developed draft, and performing content gatekeeping, which means actively selecting, rejecting, and revising AI suggestions. The paper also claims that readers in a leisurely reading setting cannot reliably distinguish AI-assisted passages from solo work, rate both kinds of AI-assisted work as preserving the writer's authentic voice, and respond positively rather than negatively to writers' use of AI.","pith_inferences":["If authenticity is process-oriented, then tools that maximize output efficiency may undermine the internalization writers value; a testable extension is to compare feedback-only tools against text-generation tools on writers' sense of authenticity.","The paper's null behavioral results may reflect a weak personalization manipulation rather than the absence of an effect; a replication with stronger personalization and a manipulation check could change the design implications.","The reader positivity may be specific to casual, leisurely reading; professional editing, journalistic, or academic contexts with stricter provenance norms may produce different reactions, which the paper does not test.","The \"80% me\" framing suggests writers track a quantitative contribution ratio; future work could study how that ratio is negotiated when AI suggestions are adopted verbatim versus revised, and whether tools that display such a ratio change behavior."],"forward_implications":["AI writing tools should be judged by whether they support writers' growth and internal experience, not just by the quality of generated text.","Personalization that only mimics a writer's surface style is insufficient; support is most valuable before and around text production, in early ideation and feedback.","Writers' fear of reader backlash may be overstated in leisurely-reading contexts, since readers in this study could not distinguish AI-assisted writing and viewed AI experimentation positively.","Contribution and authorship in co-writing should be understood through content gatekeeping decisions rather than by the proportion of text a writer typed."],"supporting_citations":[{"why":"Supplies the AI Ghostwriter Effect finding on perceived ownership that this study extends to authenticity and personalization.","marker":"[13]"},{"why":"Provides the CoAuthor interface used for the co-writing sessions with GPT-4.","marker":"[40]"},{"why":"Defines the classic authenticity framework of category, source, and value that the paper's writer-centered definition extends.","marker":"[21]"},{"why":"Contributes the anthropological framing of authenticity as source that the writers' source-of-content theme echoes.","marker":"[23]"},{"why":"Offers the theory of authenticity in creative work against which the authors compare writers' internal-experience emphasis.","marker":"[42]"},{"why":"Establishes the role of originals and source in judgments of value, used to frame content and source authenticity.","marker":"[51]"},{"why":"Defines AI-mediated communication and grounds the comparison between interpersonal AIMC and one-way creative writing.","marker":"[22]"},{"why":"Provides evidence of readers' negative perceptions of AI-written profile text that the paper's positive reader results contrast with.","marker":"[31]"},{"why":"Describes in-context learning, the method used to build each writer's personalized tool from their writing sample.","marker":"[3]"},{"why":"Shows that co-writing with opinionated language models shifts users' views, the key prior effect the paper's design and discussion reference.","marker":"[29]"}],"fun_headline_variants":["Authenticity in AI writing: It's the process, not the output","Readers can't spot AI-assisted writing, study finds","Writers: AI is 20% of me, but my voice stays mine","Co-writing with AI? Your authentic voice is in the judgment","It was 80% me, 20% AI: authenticity in co-writing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central comparison between personalized and non-personalized AI assumes that one 200-word writing sample is enough for writers to experience the tool as reflecting their voice, but the study never checks whether writers actually did experience it that way.","fun_headline_variants_meta":{"raw":{"variants":["Authenticity in AI writing: It's the process, not the output","Readers can't spot AI-assisted writing, study finds","Writers: AI is 20% of me, but my voice stays mine","Co-writing with AI? Your authentic voice is in the judgment","It was 80% me, 20% AI: authenticity in co-writing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000682,"raw_usage":{"total_tokens":3109,"prompt_tokens":967,"completion_tokens":2142,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":2048}},"tokens_in":583,"tokens_out":2142,"duration_ms":27770,"temperature":1.0,"reasoning_tokens":2048,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:54:08.822670+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication with a manipulation check, asking writers after each session how much the suggestions sounded like their own voice, would settle the personalization claims: if writers rate the personalized tool no higher than the non-personalized one, the preference and null behavior findings cannot be attributed to personalization; if, with stronger personalization, writers' acceptance rates rise and readers can distinguish the output, the paper's claim that personalization does not change behavior or reader perception would fail.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the anthropological framing of authenticity as source that the writers' source-of-content theme echoes."},{"cited_title":"Lehman, Kieran O’Connor, Balázs Kovács, and George E","cited_arxiv_id":null,"evidence_quote":"Offers the theory of authenticity in creative work against which the authors compare writers' internal-experience emphasis."}],"review_version":1}