{"id":"72187673-4074-4342-9f9d-48c43fd73e86","arxiv_id":"2508.16385","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"ChatGPT-generated texts show a distinct linguistic fingerprint, preferring nouns over verbs and varying less across registers than human writers.","lead":"The paper reports that ChatGPT-written texts can be told apart from human writing by their grammar: the model leans on nouns and shows less variation across formats. It matters because it offers a potential new signal for detecting AI text and a clue about how machines and people use language.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claim depends on noun/verb asymmetry surviving topic- and register-matched baselines; abstract supplies no matching or controls.","rationale":"The reader's UNVERDICTED verdict stems from lack of the full manuscript, and the abstract alone indeed leaves the core claim methodologically unverified. My stress-test identifies a specific, concrete version of that gap: the noun/verb asymmetry and limited register variation could arise from comparing intrinsically different text types (encyclopedic vs. essayistic) without controlling for topic and content. Wikipedia-style prose is legitimately noun-heavy, and essay writing is verb/TAM-heavy; if the human and model text sets are not matched for topic, genre conventions, and prompt constraints, the reported 'fingerprint' may simply be a prompt-register artifact. This is not an accusation of error or fraud; it is the natural null hypothesis that must be excluded. The proposed matched-design experiment would settle whether the effect is genuinely attributable to the model's generation process. Because the full paper may already include such controls, the verdict should not move to ACCEPT or REJECT on the abstract alone; UNVERDICTED/UNCHANGED remains appropriate until the manuscript is inspected. The abstract is clear and the research question is valuable, but a strong claim about a 'litmus test for AI' demands direct falsification of the confound.","tokens_in":675,"tokens_out":2981,"duration_ms":39071,"concrete_test":"Recompute the noun/verb and register analyses with a matched-design control: for each register (e.g., Wikipedia entry, college essay), sample human texts and generate ChatGPT texts on the same 50 topics, using fixed prompt templates and constant decoding parameters (temperature, top-p, max length). Then fit a mixed-effects model with random intercepts for topic and register and a fixed effect for author type (human vs. model), plus register × author interaction. If the noun/verb preference or the register-variation effect loses significance after topic matching, the fingerprint is an artifact of content/register confounds. Additionally, compute the same metrics on a second LLM (e.g., Claude or LLaMA) and on an older GPT version; if noun/verb ratios differ across models, the trait is not a universal non-human linguistic backbone.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that ChatGPT-generated texts show limited register variation and a noun-over-verb preference, implying a distinctive non-human linguistic backbone. The load-bearing premise is that this imbalance is inherent to the model rather than an artifact of comparing noun-heavy registers (e.g., Wikipedia entries) against verb/TAM-heavy essays without matching content, length, prompt wording, or socio-demographic covariates. The abstract reports only aggregate stylometric/register comparisons; it does not state which registers were used, what prompts generated the model texts, which human corpus served as baseline, how part-of-speech tags were assigned, or whether topic/register variance was statistically controlled. If the human Wikipedia baseline is itself unusually noun-dense, and the model was asked to mimic that genre, the noun/verb ratio may largely reflect the targeted register rather than a stable model fingerprint. The further claim that this pattern is a 'litmus test for AI' assumes generalizability across model versions, decoding parameters, languages, and prompt phrasings; without evidence, the observed effect could be a transient property of one model version or corpus pairing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims that stylometric and multidimensional register analyses of ChatGPT-generated texts reveal a detectable non-human fingerprint: model outputs show more limited register variation than human texts, and they favor nouns over verbs, whereas human writing relies on tense, aspect, and mood. The abstract proposes that this grammatical difference could serve as a 'litmus test' for AI-generated text. My assessment is based solely on the abstract, as the full-text portion of the review package was not supplied. The claim is plausible but not verifiable from the abstract alone, which omits details of the corpus, prompts, register matching, statistical methods, and effect sizes.","tokens_in":907,"tokens_out":4115,"duration_ms":50442,"significance":"If substantiated, the result would be a notable contribution to stylometry and AI-text detection, offering a potentially falsifiable marker of machine authorship. The proposed contrast between nominal and verb/TAM-centered language is interesting and could motivate further work. However, the abstract provides no quantitative evidence or methodological detail, so the significance cannot be evaluated at this stage. The paper would benefit from reporting matched comparisons and effect sizes, and from tempering the 'litmus test' generalization.","major_comments":[{"comment":"The central claim of a noun-over-verb preference in model texts depends on the comparison being matched across registers. The abstract does not state whether the human and model texts are matched for topic, genre, length, or formality. If the model was prompted to produce Wikipedia-style entries and compared with human Wikipedia articles, the noun density may be an artifact of the encyclopedic register rather than a model-specific trait. Without a matched or statistically controlled design, the observed asymmetry cannot be attributed to the model.","section":"Abstract"},{"comment":"No quantitative evidence is reported: no effect sizes, confidence intervals, or test statistics for the claimed differences in register variation or part-of-speech distributions. The phrase 'more limited variation' is not operationalized. For example, is the model's standard deviation across registers significantly smaller than the corresponding human standard deviation, and by how much? Without this information, the abstract does not establish that the difference is diagnostic.","section":"Abstract"},{"comment":"The 'litmus test for AI' claim overgeneralizes from what appears to be a single model (ChatGPT) and a limited set of registers. No evidence is given that the noun/verb asymmetry persists across model versions, decoding temperatures, languages, or prompt phrasings. The abstract provides no support for treating this as a stable marker of non-human authorship, and the speculative link to 'a mode of thought unique to humans' goes beyond the data presented.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase 'linguistic fingerprint' is used in a nonstandard way: in authorship attribution it usually identifies an individual, but here it appears to mean a group-level model/human distinction. This should be clarified.","section":"Abstract"},{"comment":"The sentence 'It is possible that the more complex domains of grammar reflect a mode of thought unique to humans' is an interpretive leap; it should be explicitly labeled as speculation, not a finding.","section":"Abstract"},{"comment":"The abstract would benefit from citing at least one dataset or benchmark used, and from naming the model version and date of evaluation, to aid reproducibility.","section":"General"}],"recommendation":"uncertain","confidential_remarks":"The referee was provided with only the abstract; the 'FULL TEXT' section in the review package was blank. As a result, I cannot determine whether the missing methodological details (e.g., matched corpora, statistical controls, effect sizes) are already present in the body of the paper. If the full text is available, the recommendation should be revisited; otherwise, the paper as presented cannot be verified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nI only have the abstract, so this is provisional. The paper claims ChatGPT texts show narrower register variation than human texts and prefer nouns over verbs, while humans anchor in tense, aspect, and mood. That is a concrete, testable claim, and the noun/TAM axis is a genuinely fresh angle on AI-text detection. Using multidimensional register analysis is the right approach.\n\nBut the abstract gives me no reason to trust the result yet. No corpus sizes, no prompts, no model version, no statistical tests, no mention of matching registers or topics across human and machine texts. The stress-test worry is valid: if the human baseline is heavy on noun-dense Wikipedia prose and the model was prompted to imitate that genre, the noun preference could simply reflect genre fidelity, not a model-specific trait. The same goes for the claim of limited register variation, which needs a controlled comparison with matched content and length.\n\nThe 'litmus test for AI' line is too strong. Even if the pattern holds for this model and these registers, generalizing to all LLMs, languages, and decoding settings is a leap. That would be a presentation issue if the data is solid, but it is a substantive flaw if the controls are missing.\n\nI can't give a soundness verdict from the abstract. But the question is important for forensic stylometry and LLM evaluation, and the claim is specific enough to evaluate with proper methods. A serious referee could check whether the POS tagging is reliable, whether the registers are truly matched, and whether the effect sizes are meaningful. The paper deserves that scrutiny, not a desk reject.\n\nBottom line: send it to peer review, but I would not cite it until the methods and data are visible.\n\nBest","headline":"A plausible but unverified claim that ChatGPT has a noun-heavy, low-variation grammatical fingerprint; worth refereeing, not citing yet.","tokens_in":1298,"tokens_out":2029,"would_cite":false,"duration_ms":25299,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that ChatGPT-generated texts carry a detectable non-human linguistic fingerprint: narrow register variation and a preference for nouns over the tense-aspect-mood grammar where human writing anchors.","keywords":["authorship attribution","stylometry","multidimensional register analysis","large language models","ChatGPT","linguistic fingerprint","tense aspect mood","noun-to-verb ratio"],"falsifier":"Collect a matched corpus where the same writing tasks are given to ChatGPT and to humans, varying prompts and topics widely. If the model's spread across registers matches or exceeds humans once prompt difficulty and topic are controlled, the narrower-variation fingerprint disappears.","tokens_in":647,"feed_emoji":"🤖","tokens_out":3371,"duration_ms":39575,"temperature":0.7,"pith_summary":"This paper asks whether ChatGPT has a linguistic fingerprint, just as individual people do. Comparing human- and model-written texts across registers with stylometric and multidimensional register analysis, it finds that the model can adapt its style (for example, a Wikipedia entry versus a college essay) but not to the point of indistinguishability. Model outputs vary less across registers than human writing and show a noun-over-verb preference. Human writing, by contrast, anchors in tense, aspect, and mood. If correct, this gives a grammatical cue for detecting non-human text and raises the possibility that these domains of grammar reflect a human-specific mode of thought.","feed_headline":"AI text shows a noun-heavy fingerprint humans lack","feed_subtitle":"Across registers, model text varies less and leans on nouns; human writing anchors in tense, aspect, and mood.","key_machinery":"The central object is multidimensional register analysis, a method that locates texts on continua of situational language use through co-occurring grammatical features, combined with stylometric comparison. The decisive quantities are the model's compressed register range and its noun/verb balance, contrasted with human reliance on tense, aspect, and mood markers. These features together form the proposed non-human fingerprint.","core_discovery":"The paper sets out to determine whether a large language model has a recognizable 'linguistic fingerprint' in the same way individual people do. Comparing ChatGPT output with human writing across several registers using stylometric and multidimensional register analysis, the authors find that ChatGPT adapts its style to the requested register, but the adaptation is incomplete. Model-generated texts show narrower variation across registers than human texts, and they are marked by a preference for nouns over verbs. Human writing, in contrast, is anchored in the highly grammaticalized categories of tense, aspect, and mood. The authors interpret this grammatical divergence as a possible marker o","pith_inferences":["I would predict the same noun-over-verb signature appears in other instruction-tuned LLMs if the mechanism is the next-token prediction objective; a cross-model replication would show whether this is ChatGPT-specific.","The tense/aspect/mood axis suggests a language-dependent test: in morphologically rich languages the human-model gap may be larger, while in isolating languages it may shrink.","The paper leaves open whether the noun preference comes from the model's default output distribution or from the prompt style itself; prompting with verb-centered tasks would help disentangle the two."],"forward_implications":["Texts generated by ChatGPT can be flagged as non-human through a grammatical profile, even when the model is explicitly trying to match a register.","Register adaptation in large language models is real but bounded: the model's stylistic range is narrower than a human writer's.","Noun preference and reduced tense/aspect/mood marking can be operationalized as features in authorship or AI-detection tools.","If the grammatical backbone is stable, written language becomes a test bed for probing whether LLMs organize meaning differently from humans."],"supporting_citations":[],"fun_headline_variants":["ChatGPT's noun-heavy grammar fails to mimic human style","AI text shows limited register variation and noun bias","Linguistic fingerprint reveals ChatGPT's non-human writing","Noun preference outs AI-generated text across registers","Human grammar complexity acts as AI litmus test"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The starting premise is that the noun preference and reduced register variation are properties of the model itself, not consequences of the specific prompts, the human texts chosen for comparison, the model version, or the registers sampled.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT's noun-heavy grammar fails to mimic human style","AI text shows limited register variation and noun bias","Linguistic fingerprint reveals ChatGPT's non-human writing","Noun preference outs AI-generated text across registers","Human grammar complexity acts as AI litmus test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000155,"raw_usage":{"total_tokens":1033,"prompt_tokens":705,"completion_tokens":328,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":255}},"tokens_in":449,"tokens_out":328,"duration_ms":4277,"temperature":1.0,"reasoning_tokens":255,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:18:45.056009+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a matched corpus where the same writing tasks are given to ChatGPT and to humans, varying prompts and topics widely. If the model's spread across registers matches or exceeds humans once prompt difficulty and topic are controlled, the narrower-variation fingerprint disappears.","supporting_citations":[],"review_version":1}