{"id":"6ba9f21b-487f-49a9-9272-a992231a2d8a","arxiv_id":"2505.01800","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A conceptual mapping from stylometric features to psycholinguistic theories is presented, but no empirical validation is provided.","lead":"This paper proposes a framework that maps 31 writing-style metrics to cognitive processes such as memory and self-monitoring, aiming to explain why human and AI text differ. It does not include experiments or data showing the framework actually works.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim is asserted, not demonstrated: the paper reports no classification experiment, dataset, or accuracy, so the claimed psycholinguistic framework is unsupported.","rationale":"The reader's weakest_assumption correctly identifies the hand-assigned feature-to-theory mapping as a serious weakness. I agree that the attributions in Tables 1 and 2 are not validated. However, I would place the decisive weight even earlier: the paper never establishes that the feature set discriminates between human and AI text at all. Without that empirical baseline, the psycholinguistic mapping has nothing to explain, and the interpretability claim is moot. The internal count inconsistency (31 vs. 18 vs. 29 features) and the unsupported cognitive attributions are secondary symptoms of this missing evidentiary core. A single evaluation on a public benchmark would resolve the question: if the features do not separate the classes, the framework fails; if they do but the assigned labels do not predict feature importance, the contribution reduces to ordinary stylometric classification with decorative theory. Since this concern is essentially the same basis as the reader's REJECT verdict, I recommend no change to that verdict.","tokens_in":7296,"tokens_out":6011,"duration_ms":66752,"concrete_test":"Run a held-out binary classification experiment on a public human/AI text corpus (e.g., AuTexTification or the StyloAI dataset from [17]) using the 29 features listed in Table 2. Report accuracy and F1, then perform an ablation in which the psycholinguistic category labels are randomly permuted while keeping the feature set identical, and compare classification performance and per-feature discriminative power (e.g., univariate AUC ranks) against the original labeling. If no classification result can be supplied, or if permuted labels predict feature rankings as well as the hand-assigned labels, the central interpretability claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The conclusion (Section 4) states that this study 'demonstrates' an interpretable framework for distinguishing AI-generated from human-authored texts. For that claim to hold, two empirical preconditions are necessary: (i) the 31 stylometric features actually separate human and AI text in a measurable way, and (ii) the assigned psycholinguistic categories in Tables 1 and 2 explain the discriminative power. Neither is established anywhere in the manuscript. There is no dataset description, no classifier, no accuracy or F1, no baseline, and no ablation. The only empirical grounding is a citation to the prior StyloAI paper [17] and two unquantified excerpts in Section 3.1. The feature-to-theory mapping is also hand-assigned: for example, contraction_count is tied to 'metacognitive self-monitoring' and first_person_count to 'self-referential awareness' without derivation or validation. The numerical inconsistency (31 features claimed, 18 mapped in Table 1, 29 listed in Table 2) further weakens confidence in the framework's precision. Because the psycholinguistic interpretability claim requires a measured link between features, cognitive constructs, and classification outcomes, and no such link is reported, the paper offers a proposal rather than a demonstration.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a framework that maps 31 stylometric features (the count varies across the manuscript) to psycholinguistic constructs such as cognitive load, metacognition, lexical access, and discourse planning, with the stated goal of making AI-generated-text detection more interpretable. The framework extends the author's earlier StyloAI model, and Section 3.1 illustrates the mapping with two short text excerpts. The conclusion asserts that the study 'demonstrates' an interpretable framework for distinguishing AI-generated from human-authored text. However, the manuscript reports no experiments: there is no dataset description, no classifier, no evaluation metric, and no quantitative evidence that the proposed features separate the two text types or that the assigned psycholinguistic explanations are valid.","tokens_in":7456,"tokens_out":3853,"duration_ms":36062,"significance":"If empirically validated, a framework that provides cognitive explanations for stylometric differences between human and AI text could be genuinely useful in educational integrity applications, where black-box detectors are difficult to trust. The paper identifies a plausible and interesting direction: connecting feature-level stylometry to theories of human writing processes. That said, the current manuscript is only a conceptual proposal. The central claim is empirical, and the paper provides no empirical support beyond citation to prior work and two anecdotal excerpts. The value of the contribution therefore depends entirely on future validation, which the conclusion prematurely claims to have delivered.","major_comments":[{"comment":"The conclusion states that this study 'demonstrates' an interpretable framework for distinguishing AI-generated from human-authored texts, but the manuscript contains no experiment, dataset description, classifier, baseline, accuracy measure, or error analysis. The only empirical elements are two unquantified excerpts in Section 3.1 and a citation to the prior StyloAI model [17]. The central claim is therefore asserted rather than supported, and this is a load-bearing omission, not a presentation issue.","section":"Section 4 (Conclusion)"},{"comment":"The number of stylometric features in the framework is inconsistent. The abstract and Section 1.1 state that 31 features are mapped to cognitive processes; Section 3 says that 'Table 1 summarises the mapping of 18 out of the 31 of these features'; Table 1 actually lists 19 distinct feature names; and Table 2 lists 29 features. The paper must define a single, consistent feature set and explain the relationship between the claimed 31 features, the 18 (or 19) mapped in Table 1, and the 29 listed in Table 2.","section":"Section 3 and Tables 1 and 2"},{"comment":"The assignment of stylometric features to psycholinguistic processes is hand-specified without empirical derivation or validation. For example, Table 2 states that contraction_count reflects metacognitive self-monitoring and that first_person_count reflects self-referential awareness, but no data, prior study, or formal argument establishes these links. Since the paper's interpretability claim rests on these feature-to-theory mappings, they need justification or empirical testing; otherwise the framework's explanatory component is arbitrary.","section":"Section 3, Tables 1 and 2"},{"comment":"The paper presents as fact claims about AI systems lacking cognitive processes, such as 'AI systems do not experience cognitive load in the human sense' and 'AI models, by contrast, lack fundamental metacognitive capabilities.' These claims are unsupported and are used as premises for why particular features should discriminate human from AI text. They should be reframed as hypotheses with testable implications, rather than established facts.","section":"Section 2"}],"minor_comments":[{"comment":"The title contains an extra space: 'T ext' should be 'Text'.","section":"Title"},{"comment":"The text attributes the cohesion discussion to Halliday and Hasan (1976), but reference [3] is Carrell (1982), 'Cohesion is not coherence.' Please correct the citation or add the Halliday and Hasan reference.","section":"Section 2.4 and References"},{"comment":"'Open ai' should be written as 'OpenAI'.","section":"Section 1, Introduction"},{"comment":"The feature name 'ﬂesch_reading_ease' in Table 2 should be 'flesch_reading_ease'.","section":"Table 2"},{"comment":"The sentence referring to 'Table 2 in the Appendix' is confusing because Table 2 appears in the main body of the manuscript, not in an appendix; please clarify the intended organization.","section":"Section 3, Table 2"}],"recommendation":"reject","confidential_remarks":"The paper is a proposal for a psycholinguistically grounded stylometric framework, but its conclusion claims empirical demonstration. The absence of any experiments, the inconsistent feature counts, and the unsupported feature-to-theory mappings are fundamental issues that cannot be fixed by minor revision. If the author plans to submit a revised version, it would need to include a concrete evaluation using a dataset, a defined feature set, a classifier, and an analysis of whether the psycholinguistic attributions correspond to any measurable patterns. I also note that the framework relies heavily on the author's prior StyloAI paper [17], but the current manuscript does not provide enough detail about that system for the mapping to be independently assessed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The punchline: this is a taxonomy with a conclusion that overstates it. Opara maps stylometric features from her earlier StyloAI work onto psycholinguistic theories and then says the framework “demonstrates” interpretable AI-text detection. Nothing in the paper demonstrates anything in the empirical sense. There’s no dataset, no classifier, no accuracy, no comparison, no ablation.\n\nCredit where it’s due: the paper is clearly written and the idea is not silly. If you want a structured way to think about why certain features—contractions, first-person pronouns, hapax legomena—might separate human and machine text, Tables 1 and 2 are a reasonable starting point. The conceptual organization into lexical, syntactic, sentiment, readability, named entity, and uniqueness categories is tidy, and the citations to cognitive load theory, metacognition, and lexical access are relevant. As a proposal for future work, it’s a coherent research agenda.\n\nThe soft spots are large. The central claim depends on two things the paper never measures: that the 31 features actually discriminate, and that the assigned cognitive labels explain the discrimination. Neither is established. The only empirical grounding is a citation to the author’s own StyloAI paper and two unlabeled excerpts from its dataset. The feature count is inconsistent: 31 claimed, Table 1 maps 18 (or 19 if you count unique features), and Table 2 lists 29. That kind of sloppiness undercuts confidence in the framework’s precision. The psycholinguistic attributions are hand-assigned; “contraction_count reflects metacognitive self-monitoring” is asserted, not derived. The conclusion switches from “proposes” to “demonstrates” without any new evidence. Also, claims like “AI lacks fundamental metacognitive capabilities” are stated as fact when they’re at best reasonable hypotheses.\n\nThe citation pattern isn’t a problem on its own—[17] is the author’s own prior work and self-citation is normal here—but the paper leans on it rather than re-evaluating it, so the reader can’t tell whether StyloAI’s features work in the first place.\n\nWho is this for? Someone looking for a quick overview of stylometric features with cognitive labels might find it useful. For an editor, it’s not ready for peer review as a research paper. I’d desk reject with an invitation to resubmit after either (a) adding real experiments or (b) rewriting it explicitly as a position paper with the counts fixed and the conclusions softened.","headline":"A tidy taxonomy of stylometric features with psycholinguistic labels, but the central claim of a demonstrated interpretable framework has no empirical support.","tokens_in":8031,"tokens_out":3428,"would_cite":false,"duration_ms":33037,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A psycholinguistic map explains why human and AI writing differ.","keywords":["AI-generated text detection","stylometry","psycholinguistics","cognitive load theory","metacognition","authorship verification","large language models","academic integrity"],"falsifier":"A direct test: measure the features assigned to cognitive load (average sentence length, complex sentence count) in human essays written under high versus low cognitive load, and measure features assigned to self-monitoring (contraction count, first-person count) in human and LLM text; if load does not shift the predicted features, or if LLM output is indistinguishable from human on the self-monitoring features, the framework's causal interpretation collapses.","tokens_in":7009,"feed_emoji":"🧠","tokens_out":8770,"duration_ms":78378,"temperature":0.7,"pith_summary":"This paper proposes that the stylistic differences between human-written and AI-generated text are not arbitrary statistical quirks but traces of cognitive processes. It maps 31 stylometric features—word counts, sentence complexity, sentiment, readability, named entities, and lexical uniqueness—onto psycholinguistic mechanisms such as cognitive load, metacognitive self-monitoring, lexical retrieval, and discourse planning. The intended result is an interpretable framework for AI-text detection: instead of a black-box score, a detector could explain that a text looks machine-written because it lacks the self-monitoring or planning signatures human writers leave behind. The paper's contribution is this mapping and the argument that human writing carries measurable cognitive signatures AI systems do not share.","feed_headline":"Psycholinguistic map explains why human and AI writing differ","feed_subtitle":"If the mapping holds, detectors can say why a text looks machine-written, not just flag it.","key_machinery":"The carrying mechanism is a two-table mapping: six stylometric feature categories (lexical, syntactic, sentiment, readability, named entity, uniqueness) are linked to four psycholinguistic theories (Cognitive Load Theory, metacognition and self-monitoring, lexical access and retrieval, discourse planning and cohesion), with individual features assigned to processes in Table 1 and Table 2. The mapping does the explanatory work: it converts each quantitative feature into a claimed cognitive cause, allowing the paper to propose why that feature should discriminate human from AI writing. The framework leans on the 31-feature stylometric detection model from the author's earlier work as its empirical base.","core_discovery":"The central claim, stated in the paper's own terms, is that stylometric features in human writing are surface indicators of underlying psycholinguistic processes, and that these processes are absent in AI generation. The paper assigns each of 31 features to one or more of four processes: cognitive load management, metacognition and self-monitoring, lexical access and retrieval, and discourse planning and cohesion. For example, contraction count is presented as a sign of metacognitive self-monitoring, first-person pronoun count as a sign of self-referential awareness, and average sentence length as a sign of cognitive load management. On this account, AI-generated text differs from human text not just in feature values but in the absence of the cognitive states those features index; AI produces fluent output through statistical prediction without experiencing load, monitoring, retrieval effort, or planning. The conclusion the paper draws is that this feature-to-cognition mapping gives an interpretable basis for distinguishing AI-generated from human-authored text.","pith_inferences":["One extension the paper does not pursue: induce cognitive load in human writers (e.g., dual-task typing) and check whether the features assigned to load move as predicted; if they do not, the causal mapping would need revision.","A further implication is that an LLM explicitly trained or prompted to imitate these cognitive traces—inserting contractions, varying sentence length under simulated load—could erode the discriminative value of the mapped features, pushing detection toward content-level or process-based signals.","The mapping is best read as an interpretive overlay rather than a quantitative model, since several features appear under multiple cognitive processes and no causal magnitudes are specified."],"forward_implications":["If the mapping is correct, AI-text detectors can report a cognitive rationale for each decision—for example, that a low contraction count and low first-person pronoun count indicate missing self-monitoring—rather than an opaque score.","Feature selection for authorship verification can be guided by psycholinguistic theory, concentrating on features whose cognitive basis is strongest rather than all available statistics.","The framework predicts that human writing under high cognitive load should exhibit the features the paper associates with load (e.g., shorter sentences, lower complexity), giving a natural experimental check.","In educational settings, a detector built on this framework could distinguish machine-like statistical fluency from the imperfect, self-monitored patterns the paper treats as markers of genuine authorship."],"supporting_citations":[{"why":"Supplies the 31-feature detection model and dataset excerpts that the framework extends.","marker":"[17]"},{"why":"Defines cognitive load theory, used to justify load-related feature mappings.","marker":"[22]"},{"why":"Defines metacognition and self-monitoring, assigned to features like contraction count and first-person count.","marker":"[7]"},{"why":"Provides the lexical-access and retrieval account behind vocabulary-based features.","marker":"[6]"},{"why":"Cited for discourse cohesion and planning, grounding cohesion-related feature mappings.","marker":"[3]"},{"why":"Supports the claim that language and thought dissociate in large language models, so AI text lacks intentional cognition.","marker":"[12]"},{"why":"Supports the claim that large language models produce fluent text without human-like cognitive load or conceptual coherence.","marker":"[21]"},{"why":"Used to argue that human language processing involves brain mechanisms AI cannot replicate, underpinning the framework's core contrast.","marker":"[24]"}],"fun_headline_variants":["AI writing lacks cognitive fingerprints humans leave behind","Why AI text reads fluent but lacks human thinking traces","Psycholinguistic features unmask AI writing without buzzwords","Detecting AI by the missing mind behind the words","Human text bears cognitive traces AI can't fake"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework hinges on the assumption that each style feature is a reliable outward sign of a particular mental process in human writers, and that AI text lacks those signs; if any of those mappings is wrong, the explanation for why the features separate human and AI writing fails.","fun_headline_variants_meta":{"raw":{"variants":["AI writing lacks cognitive fingerprints humans leave behind","Why AI text reads fluent but lacks human thinking traces","Psycholinguistic features unmask AI writing without buzzwords","Detecting AI by the missing mind behind the words","Human text bears cognitive traces AI can't fake"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000388,"raw_usage":{"total_tokens":2009,"prompt_tokens":873,"completion_tokens":1136,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":1062}},"tokens_in":489,"tokens_out":1136,"duration_ms":8088,"temperature":1.0,"reasoning_tokens":1062,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:10:07.120428+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test: measure the features assigned to cognitive load (average sentence length, complex sentence count) in human essays written under high versus low cognitive load, and measure features assigned to self-monitoring (contraction count, first-person count) in human and LLM text; if load does not shift the predicted features, or if LLM output is indistinguishable from human on the self-monitoring features, the framework's causal interpretation collapses.","supporting_citations":[{"cited_title":"In: International conference on artiﬁcial intelligence in education","cited_arxiv_id":null,"evidence_quote":"Supplies the 31-feature detection model and dataset excerpts that the framework extends."},{"cited_title":"Learning and instruction 4(4), 295–312 (1994)","cited_arxiv_id":null,"evidence_quote":"Defines cognitive load theory, used to justify load-related feature mappings."},{"cited_title":"American Psychologist 34, 906–911 (1979)","cited_arxiv_id":null,"evidence_quote":"Defines metacognition and self-monitoring, assigned to features like contraction count and first-person count."},{"cited_title":"Cognition 42(1-3), 287– 314 (1992)","cited_arxiv_id":null,"evidence_quote":"Provides the lexical-access and retrieval account behind vocabulary-based features."},{"cited_title":"TESOL quarterl y 16(4), 479–488 (1982)","cited_arxiv_id":null,"evidence_quote":"Cited for discourse cohesion and planning, grounding cohesion-related feature mappings."},{"cited_title":", Tenenbaum, J.B., Fedorenko, E.: Dissociating language and thought in large language models","cited_arxiv_id":null,"evidence_quote":"Supports the claim that language and thought dissociate in large language models, so AI text lacks intentional cognition."},{"cited_title":"-ive,” “-ous","cited_arxiv_id":null,"evidence_quote":"Used to argue that human language processing involves brain mechanisms AI cannot replicate, underpinning the framework's core contrast."}],"review_version":1}