{"id":"8bb27224-56cd-4f6b-b8cf-8597aa3ae36f","arxiv_id":"2502.06560","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Fine-tuned open LLMs can imitate individual writing styles from small samples, evade detection tools, and are not yet addressed by current safeguards or law.","lead":"This position paper argues that fine-tuning open-source LLMs on a person's writing samples makes it cheap and easy to imitate their style in text, enabling phishing, fraud, and fake accounts. It urges researchers and regulators to treat text-based impersonation as a distinct risk from image, audio, and video deepfakes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Feasibility premise (75 emails suffice; detectors evaded) rests on one self-cited ENRON study and a 2-model Copyleaks test; Section 6 concedes data to gauge vulnerable users is lacking.","rationale":"After reading the paper, I find the policy argument timely and the writing clear, and I credit the authors for explicitly reporting their small-scale demonstrations and for flagging data limitations in Section 6 and the 'preliminary' nature of Table 2 in Section 5. My stress-test pass did not reveal an internal inconsistency; the concern is empirical support for the threat-model premise. The paper's central claim as articulated in the abstract and Section 4 is that a person's writing samples (roughly 75 emails per the reproduced Panza figure) suffice to create a model whose output is mistaken by acquaintances and by AI-text detectors. The evidence for the acquaintance part is a single self-cited human study in ENRON; the evidence for detector evasion is the authors' own two-model, one-tool Copyleaks test. Neither is sufficient to establish the claimed generality across domains, users, and detectors, and Section 6 concedes that the field lacks the cross-modal datasets needed to estimate how much text is needed. A replication with multiple domains, more users, forced-choice human judgments, and a multi-detector benchmark would settle whether the concern lands. If it does not replicate, the 'time to act' urgency is weakened but not eliminated, since even the possibility of such attacks may justify policy attention; hence the reader's CONDITIONAL verdict is appropriate and unchanged.","tokens_in":18815,"tokens_out":5376,"duration_ms":48107,"concrete_test":"Run an independent replication using Panza-style fine-tuning on at least 10 individuals from at least 3 domains (e.g., ENRON, Reddit history, personal blog/email). For each user, fine-tune Llama-3-8B on 75 emails/posts, generate 100 emails across the Section 4 prompt types, and have acquaintances plus independent judges classify genuine vs generated text in a forced-choice design. Simultaneously, run the same 100 generated texts (plus 100 genuine texts and 100 base-model texts) through at least 4 detectors (Copyleaks, GPTZero, Originality, RadNet) with repeated trials. Report credibility rates and detector evasion rates with confidence intervals. If fine-tuned models are not significantly more credible or evasive than base models, or if results do not generalize beyond ENRON, the central feasibility premise is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that personalized fine-tuning enables credible impersonation and evades AI-text detection is the load-bearing empirical premise, but the evidence offered in this preprint is minimal. Figure 1 reproduces a single self-cited human study (Nicolicioiu et al., 2025) conducted in one email domain (ENRON); the 'accepted by acquaintances' result lacks independent replication, confidence intervals, and cross-domain validation. The detector-evasion demonstration in Section 4 is even thinner: one online tool (Copyleaks), two fine-tuned models (Jeff and Kay), and the handful of outputs in Table 1, with no human-text or base-model false-positive controls. The paper itself flags the underlying gap in Section 6: it knows of no dataset linking an author's text across modalities and concedes it is 'difficult to estimate the amount and type of text necessary to build a compelling personalized model.' Without a broader evidence base, the position's urgency rests on a threat model whose generality is unestablished; a skeptic could attribute the results to a single corpus, a single detector, or chance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that efficient fine-tuning of open-weight language models makes personalized text generation—credible imitation of a specific person's writing style—practically feasible, and that this creates novel safety risks distinct from image/audio/video deepfakes. The paper presents illustrative outputs from fine-tuned models on ENRON emails (Table 1), reproduces human-study results from the same group's earlier Panza work (Figure 1), reports a two-model test suggesting evasion of a commercial AI-text detector (Section 4), probes five chatbots with four malicious prompts (Table 2), and reviews legal and policy gaps (Section 5). It concludes with research, model-governance, and policy recommendations (Section 6) and replies to anticipated objections (Section 7).","tokens_in":19002,"tokens_out":3764,"duration_ms":34475,"significance":"If the central feasibility claim holds, the paper identifies a genuinely underexplored risk: text impersonation that is cheap, local, hard to audit, and complementary to other deepfake modalities. The paper's strengths include an original (though small) prompt probe, a concrete detector-evasion test, a specific legal-analysis contribution (EU AI Act, FTC rule, California statutes), and an explicit acknowledgment of underlying evidence gaps. The authors are appropriately self-critical about missing datasets and measurement difficulties. However, the load-bearing empirical evidence for 'credible imitation from ~75 emails' and 'evasion of AI-text detectors' is thin, partly self-cited, and not yet independently replicated; the urgency claims therefore currently outrun the evidence base. The paper is a useful position piece that would benefit from either stronger evidence or more carefully qualified claims.","major_comments":[{"comment":"The central feasibility claim—that as few as 75 emails suffice for a model whose output is mistaken for the author by acquaintances—rests entirely on Figure 1, a reproduction of Nicolicioiu et al. (2025), co-authored by two of the present authors and the senior author. The figure reports no confidence intervals, no per-participant variability, and no independent replication or cross-domain validation. Since this is the load-bearing empirical premise of the paper, the manuscript should either add independent experiments with statistical uncertainty and significance tests (ideally outside the ENRON email domain), or explicitly reframe the claim as suggestive evidence from a single corpus rather than an established feasibility result.","section":"Section 2 / Figure 1"},{"comment":"The detector-evasion demonstration uses a single online tool (Copyleaks), two fine-tuned models (Jeff and Kay), and the handful of outputs in Table 1, with no genuine-human-text false-positive control, no repeated sampling, and no distribution of detector scores. The claim that personalized fine-tuning 'evades AI text detection' is therefore not supported at the level stated. I recommend either a systematic evaluation—multiple detectors, several users, matched human and base-model controls, reported error rates and score distributions—or a strictly hedged wording such as 'initial evidence that one commercial detector failed to flag these outputs.'","section":"Section 4"},{"comment":"The authors concede in Section 6 that no datasets link the same author across modalities and that it is 'difficult to estimate the amount and type of text necessary to build a compelling personalized model.' This concession is in tension with the abstract and introduction, which describe practical feasibility of impersonating specific individuals based on small amounts of text as an established premise. The paper should calibrate the scope of its claims to this acknowledged uncertainty, for example by distinguishing 'possible in one studied corpus' from 'prevalent across populations,' and by using that distinction in the policy recommendations.","section":"Section 6"},{"comment":"The prompt probe in Table 2 covers four prompts and five models and is presented as evidence that 'safeguards against unsafe messages are largely missing in current models.' The authors state that the prompts are not cherry-picked, but the sample is too small and insufficiently systematic to support the general claim: there is no taxonomy of attack types, no human rating of whether the responses are harmful, no variation of phrasing, and no discussion of how representative the four scenarios are. I recommend either expanding the probe or explicitly limiting the conclusion to the four tested scenarios.","section":"Section 5 / Table 2"}],"minor_comments":[{"comment":"The sentence 'Personalized text generation has also attracted industrial applications. across modalities.' contains a punctuation/capitalization error; 'applications' should be followed by a comma or the fragment should be merged into the previous sentence.","section":"Section 2"},{"comment":"In the scholastic-dishonesty paragraph, 'whether a text is is AI-generated' contains a duplicated 'is'.","section":"Section 3"},{"comment":"The text contains a typo: 'Howerver' should be 'However'.","section":"Section 3"},{"comment":"The phrase 'the International AI Safety Report ... who in Section 2.1.1' should use 'which' instead of 'who' for a report, or be rephrased to avoid the grammatical mismatch.","section":"Section 5"},{"comment":"In the cryptographic-authentication paragraph, 'adoption of has been relatively slow' is missing a word; it should read 'adoption of these methods has been relatively slow'.","section":"Section 7"},{"comment":"The main text refers to 'ChatGPT-4o' and 'AI's Claude-3.5 Haiku,' while the appendix labels the same responses as 'GPT-4-Turbo' and 'Claude 3.5 Sonnet'; the model names should be consistent and, where possible, include version and access date information.","section":"Table 2 / Appendix"},{"comment":"The Copyleaks reference URL is missing the protocol slashes: 'https:www.copyleaks.com' should be 'https://www.copyleaks.com'.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The near-total reliance on the authors' own Panza study for the core feasibility claim is a concern for a position paper arguing for urgency; the paper should be reviewed with attention to whether the evidence is sufficiently independent. The detector-evasion result is also quite preliminary relative to the strength of the wording. These issues are fixable within the manuscript's scope, so I do not recommend rejection, but the revision should either supply stronger evidence or visibly temper the central claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should read this one if you care about AI safety or policy. It's a position paper, and it makes a good case that personalized text impersonation—fine-tuning an open LLM on someone's writing to imitate them—is a distinct, under-regulated risk, not just a footnote to image/audio/video deepfakes. The authors show the practical barrier is low (consumer GPU, small data), and they document a real legal gap: the EU AI Act defines deepfakes as image/video/audio only, California's recent bills protect voice and likeness but not writing style, and the FTC business-impersonation rule doesn't cover individuals. That part is well researched and current.\n\nThe best original piece is Table 2: four malicious prompts across five models, with full transcripts in the appendix. They state these are the only four prompts they tried, no prompt engineering. That's honest and reproducible—only one model (Claude) pushed back meaningfully, and even it produced a variant. Also, they got a 100%-to-0% Copyleaks detection flip on two personalized models versus the base model. Small, but the direction is consistent with what you'd expect.\n\nSoft spots, in order. First, the human-credibility evidence is borrowed entirely from one self-cited study (Panza, Nicolicioiu et al. 2025, which includes all three of these authors), on ENRON email data only. The paper is transparent about the source, and the examples in Table 1 help, but independent replication would matter a lot. Second, the Copyleaks test is two models, one tool, no false-positive controls. They call it a 'simple demonstration,' so the framing is fine—just don't let a reviewer treat it as a systematic result. Third, they admit in Section 6 that no dataset links the same author across modalities, so they can't estimate how many people are actually vulnerable. That's a real limitation, but it's stated, not buried.\n\nMy bottom line: this is a solid position piece. The central claim—that this risk is real, cheap, local, and under-addressed—survives the thin empirical base, because the argument is about plausibility and policy, not prevalence. The paper would be stronger with a few independent detector tests and some gesture toward replication, but those are revision asks, not reasons to desk reject. Yes, send it to peer review. If the reviewers take it for what it is (a call to action with preliminary evidence), it would move the conversation forward.","headline":"A legitimate position paper on a real and underappreciated risk; the empirical legs are thin and partly self-cited, but the argument doesn't depend on them, and it deserves a serious referee.","tokens_in":19521,"tokens_out":2705,"would_cite":true,"duration_ms":24881,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cheap, locally run fine-tuning of open-weight LLMs already lets attackers write convincingly in a specific person's voice, and this 'text deepfake' risk is largely absent from AI safety discourse.","keywords":["LLM personalization","text deepfakes","impersonation","phishing","AI-generated text detection","parameter-efficient fine-tuning","AI safety policy","open-source LLMs"],"falsifier":"A preregistered study across diverse authors and genres (chat logs, essays, social media) in which acquaintances identify personalized-model outputs at chance rates, or a commercial detector flags them at rates comparable to base-model outputs, would undercut the paper's claim that personalized imitation is a practical, general threat.","tokens_in":18625,"feed_emoji":"🖋️","tokens_out":8978,"duration_ms":68184,"temperature":0.7,"pith_summary":"This paper argues that cheap, locally run fine-tuning of open-weight language models has made it practically feasible to impersonate a specific person in writing, and that this 'text deepfake' risk is distinct from image, audio, and video deepfakes and largely neglected. The authors point to a reproduced personalization method that trains a Llama-3-8B model on roughly 75 emails of a person and produces messages that acquaintances attribute to that person, and to a demonstration that the same outputs are rated 0% AI-generated by a commercial detector. They contend this capability enables personalized phishing, character assassination through fake accounts, and self-imitation to escape AI detection or academic integrity systems, and they show that several major chatbot services will readily write such harmful messages. The paper therefore calls on the research community and policymakers to treat written-style impersonation as its own safety problem rather than folding it into other deepfake categories.","feed_headline":"Finetuned LLMs impersonate people from 75 emails","feed_subtitle":"A position paper argues text-based deepfakes are overlooked, cheap to run locally, and slip past detectors and friends.","key_machinery":"The load-bearing machinery is the personalization pipeline: parameter-efficient fine-tuning (low-rank adaptation) of an open-weight base model such as Llama-3-8B, guided by instruction back-translation, in which an LLM writes synthetic prompts for which a person's genuine texts are the desired answers. This pipeline converts a small corpus of someone's writing into a style-matched text generator that runs on consumer hardware, with inference at about 16 tokens per second on a laptop CPU. The MAUVE metric, measuring distributional distance between corpora, is the evaluation device that makes the style match visible, while the human-acquaintance studies and the Copyleaks detector tests are the evidence that the match holds up in practice.","core_discovery":"Its central claim is that efficient personalized text generation is already a working tool for impersonation: an individual's style can be captured by fine-tuning a small open-weight LLM on their own writing, and the resulting model produces text that people who know the author cannot reliably distinguish from the genuine article. The paper supports this with the Panza study's human evaluations, in which about 68% of emails from a model fine-tuned on 75 emails were judged credible by an acquaintance, versus 76% for genuine human emails and 36% for the base model, and with the authors' own Copyleaks test, in which personalized outputs scored 0% AI-generated while the same prompts through the un-fine-tuned model scored 100%. From this, the paper argues that textual impersonation differs fundamentally from visual and audio deepfakes: the medium is low-bandwidth, training and inference can be done entirely locally away from centralized auditing, watermarking is easy to evade, and style imitation can be combined with other modalities. It also documents that current legislation and influential policy reports define deepfakes around image, video, and audio, excluding written style, and that four popular LLM interfaces, with one partial exception, supplied useful harmful drafts for all four malicious prompts tested.","pith_inferences":["Beyond the paper: the same pipeline should transfer from email to chat logs, social-media posts, and forum writing, which would enlarge the vulnerable population well beyond the ENRON-based study.","Beyond the paper: the MAUVE-style distributional distance used to prove imitation could be inverted into a defense, making stylometric authorship verification a tool for detecting machine-written impersonation.","Beyond the paper: if written impersonation becomes routine, message-signing protocols and style-based provenance may become standard trust infrastructure for personal correspondence.","Beyond the paper: the 75-email threshold hints at a measurable scaling law; a corpus-size versus identification-accuracy curve across genres could predict when any given individual becomes practically imitable."],"forward_implications":["Spear-phishing can be upgraded with a trusted sender's style, making fraudulent requests dramatically more persuasive while costing only consumer-grade compute.","Statistical AI-text detectors, the dominant defense today, fail on fine-tuned personalized output, so detection alone cannot be the safety net.","Fine-tuning can break a base model's safety alignment, potentially giving attackers access to functionality that was blocked at release.","Legal definitions of deepfakes in the EU AI Act and US state laws omit written style, so victims of text impersonation have limited legal recourse.","Model distributors could propagate access controls and watermark individual weight copies to give forensic accountability for downstream misuse."],"supporting_citations":[{"why":"Supplies the Panza personalization method and the human studies showing acquaintances attribute outputs of models fine-tuned on ~75 emails to the author.","marker":"Nicolicioiu et al. (2025)"},{"why":"Demonstrates per-user parameter-efficient fine-tuning on the LAMP benchmark, establishing that cheap personalization outperforms RAG alone.","marker":"Tan et al. (2024)"},{"why":"Introduces the LAMP benchmark that makes task-level LLM personalization measurable, grounding feasibility claims.","marker":"Salemi et al. (2024)"},{"why":"Introduces LoRA, the low-rank adaptation technique that makes consumer-grade fine-tuning practical.","marker":"Hu et al. (2021)"},{"why":"Introduces QLoRA quantized fine-tuning, the mechanism that allows personalization on a free or low-cost GPU.","marker":"Dettmers et al. (2023)"},{"why":"Provides the ENRON email corpus used for the reproduced personalization experiments and Table 1 example outputs.","marker":"Cohen (2015)"},{"why":"The detector used in the paper's direct demonstration: 100% AI-generated for base-model output, 0% for both personalized models.","marker":"Copyleaks (2025)"},{"why":"Shows that fine-tuning aligned models compromises safety even without malicious intent, supporting the alignment-brittleness claim.","marker":"Qi et al. (2023)"},{"why":"The International AI Safety Report, cited to show major policy frameworks discuss only speech, video, and audio while omitting written-style impersonation.","marker":"Bengio et al. (2025)"}],"fun_headline_variants":["AI clones your writing style from just 75 emails","Text deepfakes are cheap, local, and hard to catch","Personalized LLMs impersonate you from 75 samples","Fine-tuned models fake your writing from 75 emails"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The threat assessment depends on the assumption that a small sample of a person's writing—roughly 75 emails from one email domain, per the reproduced study—is enough to train a model that acquaintances and detectors accept as that person.","fun_headline_variants_meta":{"raw":{"variants":["AI clones your writing style from just 75 emails","Text deepfakes are cheap, local, and hard to catch","Personalized LLMs impersonate you from 75 samples","Fine-tuned models fake your writing from 75 emails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000438,"raw_usage":{"total_tokens":2253,"prompt_tokens":1002,"completion_tokens":1251,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":1183}},"tokens_in":618,"tokens_out":1251,"duration_ms":25217,"temperature":1.0,"reasoning_tokens":1183,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:03:48.146569+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A preregistered study across diverse authors and genres (chat logs, essays, social media) in which acquaintances identify personalized-model outputs at chance rates, or a commercial detector flags them at rates comparable to base-model outputs, would undercut the paper's claim that personalized imitation is a practical, general threat.","supporting_citations":[],"review_version":1}