{"id":"e4544fad-d356-4114-9209-bc62152e408e","arxiv_id":"2604.03136","paper_version":5,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Discourse-level narrative features, extracted by LLMs across 10 dimensions, distinguish AI-generated from human-written stories at 93.2% macro-F1 and attribute authorship at 68.4% macro-F1.","lead":"StoryScope extracts 304 narrative-level features from 61,608 human and AI-written stories and shows these features alone separate human from AI fiction at 93.2% macro-F1. It is a new interpretable detection route that targets plot, agency, and temporality rather than surface style.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 93.2% narrative-only result rests on LLM-assigned feature values; if Gemini 3 Flash's annotations embed stylistic or source-specific priors, the features are not objective narrative measures and the central claim is unsupported.","rationale":"The reader's weakest assumption correctly identifies LLM annotation validity as the load-bearing point. I agree, and I sharpen it by noting that the annotation model is Gemini 3 Flash, which is also one of the six sources, increasing the risk that the features encode the annotator's own stylistic or source-recognition biases. The human validation in Appendix C is too small (12 stories, 240 items) to rule this out, and the human–human κ of 0.74 suggests the feature definitions themselves are not fully objective. This is not an internal inconsistency; it is a correctness-risk concern about whether the empirical separation generalizes beyond this specific annotation pipeline. The proposed cross-annotator replication directly tests the central interpretation. I would keep the paper's conditional status: the empirical result is plausible and well-engineered, but the 'underlying narrative construction' interpretation should not be accepted as established until the annotation instrument is shown to be reproducible across independent annotators.","tokens_in":26373,"tokens_out":4850,"duration_ms":58360,"concrete_test":"Select a stratified sample of 300 held-out test stories (50 per source). Re-run the complete feature-assignment protocol with a different annotation model (e.g., GPT-5.1 or Claude Sonnet 4.6) using the same 304 feature definitions, and also collect human annotations on a 30-story subset. Compute per-feature agreement (Cohen's κ/ICC) between the original Gemini-3-Flash vectors and the alternative vectors. Then retrain the XGBoost binary and 6-way classifiers on the alternative-annotator training features and evaluate on the same held-out test set. If macro-F1 drops by more than ~5 points, or if the top-20 SHAP features show κ < 0.6, the headline result is not robust to the annotation instrument and the 'narrative construction' interpretation is unsupported. If cross-annotator agreement is high and classifier performance holds, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — that 93.2% macro-F1 reflects 'underlying narrative construction' rather than surface style — requires that the 304 feature values are valid, objective measurements of narrative choices. The least secure link in that chain is the annotation stage: every feature value is assigned by Gemini 3 Flash, the same model family that generated one of the five AI sources. Feature definitions were themselves induced by GPT-5.1 from LLM-generated comparative analyses, and the style/non-style boundary was drawn by a GPT-5.4 audit. Human validation is tiny: 12 stories, 240 feature-items, mean human–model κ=0.84, with human–human κ only 0.74 (Appendix C). If Gemini 3 Flash answers narrative questions by picking up on stylistic regularities, source-typical phrasing, or its own generation priors — rather than genuinely measuring plot, agency, and temporality — then the 93.2% F1 is an artifact of the annotation instrument, not of narrative structure. The edit-robustness result (Section 4.2) does not resolve this, because LAMP edits surface prose and would not remove an annotator's latent stylistic sensitivity. The main empirical result is therefore contingent on cross-annotator reproducibility.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"StoryScope proposes an LLM-driven pipeline for extracting discourse-level narrative features from fiction. From 10,272 Books3 short stories, the authors generate mirror stories with five LLMs (61,608 total). A template representation is built with GPT-5.1, comparative analyses are distilled, and 304 features are discovered, then assigned to every story by Gemini 3 Flash. XGBoost on these features achieves macro-F1 93.2% for human/AI detection and 68.4% for 6-way authorship attribution. The paper identifies 30 core features and per-model fingerprints, and uses feature-space distance to argue that AI stories are more convergent and less rare than human stories. Code, prompts, and AI-generated narratives are released.","tokens_in":26720,"tokens_out":6206,"duration_ms":72720,"significance":"The contribution is substantial if the central interpretation holds. A narrative-level representation that is robust to surface style would be a valuable tool for AI authorship analysis and an interpretable complement to raw-text detectors. The evaluation is unusually careful in several respects: prompt-level train/test grouping, length-matched controls, topic-sensitivity tests, a memorization audit, repeatability measurements, and human validation of a subset of annotations. These design choices support the headline classification numbers as empirical facts about the constructed feature space. The main weakness is that the feature space is built and populated entirely by LLMs, which ties the interpretation to the annotator's objectivity. If the authors can address the annotation-validity concern, the paper would be a strong candidate for publication.","major_comments":[{"comment":"The central claim that 93.2% macro-F1 reflects \"underlying narrative construction\" rather than surface style (Abstract, §4) depends on the validity of the feature assignments, but every feature value is generated by Gemini 3 Flash, the same model family as one of the five AI sources, and the feature definitions are induced by GPT-5.1 from LLM-written comparative analyses. Human validation covers only 12 stories and 240 feature-items (mean human–model κ=0.84; human–human κ=0.74, Table 7), which is too small to rule out that the annotator is answering from source-typical stylistic priors or its own generation preferences. The LAMP edit experiment (§4.2) targets surface artifacts and would not remove such latent sensitivity. Please add a cross-annotator experiment (e.g., annotate a held-out sample with a different LLM or with human raters) and report both annotation agreement and downstream","section":"§2.1–2.2, Appendix C"},{"comment":"The \"Narrative\" variant is defined by excluding 47 features rated as style-related by a GPT-5.4 audit, but this boundary is itself one of the load-bearing modeling choices. The paper does not report inter-auditor agreement for this rating, nor a sensitivity analysis around the boundary. If even a fraction of the excluded features are actually narrative (or if some retained features are stylistic in disguise), the 93.2% number changes. Please provide the audit prompt, run an independent second audit, and report results for a wider and more restrictive narrative set (e.g., ±10 features around the boundary).","section":"§2.1, §B (style boundary)"},{"comment":"The 30 core features and the rarity/divergence analyses are selected from the same data and feature space that was constructed by discriminative discovery. This makes the interpretive statements in §4.1 and §5 largely descriptive of the induced space, not independent evidence that human stories are inherently rarer or more complex. The feature-selection thresholds in §D are justified post hoc, and the Core-Only F1 may be optimistically biased by selection on the validation set. I do not dispute the held-out classification result for the full 257-feature model, but the paper should avoid essentialist language (\"AI over-explains\", \"human authors subvert linearity\") or support it with out-of-space validation—for example, human annotations of the core concepts on a larger sample.","section":"§2.2, §5, §G"}],"minor_comments":[{"comment":"For the ModernBERT baseline, specify how long stories are truncated or downsampled; with max length 512 tokens and stories averaging 4,753 words, the comparison may not be apples-to-apples.","section":"Table 2"},{"comment":"The figure says \"~60k stories\" but the exact count is 61,608; the \"N\" in the figure is also ambiguous. Please use precise numbers.","section":"Figure 1"},{"comment":"Gemini and DeepSeek fail the length instruction by about 3,000 words on average. The paper shows length-matching for the binary task, but it would be useful to discuss whether such systematic length differences could interact with narrative feature values in the attribution task.","section":"Table 5"},{"comment":"The edit experiment uses only 278 Gemini-generated stories. Please clarify why only this subset was used and whether results generalize to the other four AI sources.","section":"§4.2"},{"comment":"The term \"rarity percentile\" is used in Section 5 before it is formally defined in Section G. A forward reference would help the reader.","section":"§G"},{"comment":"Centroid-distance statements such as \"6.6 vs. 4.3\" and \"closest human-AI pair vs. farthest AI-AI pair\" are presented without uncertainty intervals. Given the sensitivity of such claims, bootstrap confidence intervals should be reported.","section":"§5"},{"comment":"The phrase \"grounded in NarraBench\" may overstate the grounding: the feature taxonomy is not a direct instantiation of NarraBench's twelve aspects but a new LLM-derived set. Please clarify the relationship and avoid implying that NarraBench provides a validated measurement instrument for these specific 304 features.","section":"§2.1"}],"recommendation":"major_revision","confidential_remarks":"The empirical machinery is strong and the dataset release is valuable. My main concern is that the headline claim — that narrative features reflect 'underlying narrative construction' — rests on an LLM-built and LLM-annotated feature space. If the authors supply a second annotator or a substantially enlarged human validation, I would support acceptance. As it stands, the interpretation outpaces the validation evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a well-built, unusually honest empirical paper that makes a real case for narrative-level AI detection, and the 93.2% F1 is likely a true effect. But the 'narrative' claim rests on an LLM annotation instrument, and the validation of that instrument is thin. If you read it as a study of what Gemini 3 Flash reliably reports about stories, it's solid; if you read it as proof about underlying narrative construction, it needs another round of evidence.\n\nWhat's genuinely new: the pipeline that induces 304 interpretable narrative features via cross-source comparison, the parallel corpus of 10k prompts × six sources, and the demonstration that these features alone achieve 93.2% macro-F1. The paper does the right empirical work: prompt-level grouping, length-matched evaluation, topic sensitivity, memorization audit, and a careful style/non-style boundary. The core feature set is compact and plausible—AI over-explaining themes, humans using nonlinear time—and the rarity analysis is a nice operationalization of originality.\n\nThe soft spots are real but not fatal. The biggest one is the measurement loop: GPT-5.1 defines the features, Gemini 3 Flash assigns the values, and the same model family generated one of the five AI sources. The human validation is 12 stories and 240 items, with human–human kappa at 0.74. That's not enough to rule out the possibility that the annotator is picking up on stylistic regularities or its own generation priors. The LAMP robustness test edits surface prose but wouldn't remove that kind of latent sensitivity. Separately, the core-feature selection uses training labels, so the 84.8% core-only F1 is partly a fitted number, and the abstract's GPT-dream-sequence fingerprint doesn't show up in the reported fingerprint tables.\n\nNone of this kills the central result. The dataset is a contribution in itself, and the main finding—that discourse-level features carry signal independent of style—is worth taking seriously. But the paper should be pushed to validate the annotation instrument (e.g., a second annotator model or a larger human sample) and to soften the causal language. I'd send it to review; it's the kind of work that makes a subfield better.","headline":"A well-built empirical case for narrative-level AI detection, but the headline F1 is only as clean as the LLM annotator behind it.","tokens_in":81,"tokens_out":2269,"would_cite":true,"duration_ms":54113,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Narrative structure alone separates AI fiction from human fiction with 93.2% macro-F1, even after style is removed.","keywords":["AI-generated fiction","narrative features","authorship detection","discourse analysis","interpretable classification","LLM storytelling","narrative originality","story structure"],"falsifier":"Take a held-out set of stories from the same six sources, have human narrative analysts score the 304 features (or use a non-LLM protocol to do so), retrain the classifier, and compare macro-F1 to the LLM-annotated result. If the separation collapses or drops far below 93.2%, the claim that narrative construction itself drives detection is falsified. A complementary check: classify stories from a new, unseen language model not in the training set; if narrative features fail to separate it from humans, the 'shared AI narrative space' is not general.","tokens_in":26296,"feed_emoji":"🤖","tokens_out":4028,"duration_ms":43090,"temperature":0.7,"pith_summary":"This paper sets out to show that AI-generated fiction differs from human-written fiction not just in surface style but in deeper narrative decisions—how the plot is arranged, how themes are stated, how characters' choices are framed. It introduces StoryScope, a pipeline that turns more than sixty thousand stories into 304 interpretable narrative features, and finds that these features alone distinguish human from AI authorship at 93.2% macro-F1, retaining 97% of the performance of a model that also includes stylistic cues. The separation survives stylistic editing, and human stories occupy a rarer, more dispersed region of narrative space while the five AI models cluster together. If the claim holds, it gives readers and publishers a more durable, explainable basis for assessing whether a story was originally conceived by a human.","feed_headline":"Narrative choices alone expose AI fiction 93% of the time","feed_subtitle":"Why it matters: structural story decisions survive style edits, giving a durable, explainable authorship signal.","key_machinery":"The load-bearing mechanism is the StoryScope pipeline: stories are first converted into structured templates that abstract away surface wording along ten narrative dimensions; a comparative analysis stage identifies where sources diverge on the same prompt; a discovery stage turns those observations into 304 closed-form feature questions; and an LLM annotator scores every story on every feature. The resulting vectors feed gradient-boosted tree classifiers with per-feature attribution, which lets the authors isolate a small set of stable core features. What the machinery does is transfer the burden of detection from lexical surface to structural choices—causal continuity, chronological orderi","core_discovery":"The paper's central claim is that narrative construction carries a systematic, learnable signature of authorship. Using a parallel corpus in which each of 10,272 prompts was written by a human author and five language models, StoryScope induces 304 discourse-level features across ten narrative dimensions—character, plot, setting, time, revelation, perspective, and others—and classifies stories from these features alone. Narrative features reach 93.2% macro-F1 for human-versus-AI detection and 68.4% for six-way attribution; a compact set of 30 core features captures most of the binary signal. The paper also reports that AI stories over-explain themes, favor tidy single-track plots with protag","pith_inferences":["Editorial inference: because the feature values come from LLM judges, part of the gap may reflect the judges' own stylistic or prompt-following biases rather than properties of the stories; a human-annotation or non-LLM replication would test this.","Editorial inference: if narrative features are learnable, adversarial authors could 'humanize' structure by adding subplots, time jumps, and ambiguous endings; the durability claim would then erode, and the method's value would be as a moving target rather than a fixed detector.","Editorial inference: the rarity measure could be used as a generation-time reward to push AI outputs toward less typical narrative combinations, effectively testing whether the observed cluster is an inevitable property of LLMs or just a current default.","Editorial inference: the same template-plus-features approach could be applied to other long-form creative domains such as screenplays or narrative nonfiction to see whether the human/AI structural divide generalizes beyond short fiction."],"forward_implications":["A detector built only on narrative features keeps working after surface artifacts (clichés, purple prose) are edited out, suggesting it targets structure rather than style.","Because narrative choices are more expensive to alter than wording, such features may stay diagnostic longer as language models update.","The 30 core features give an interpretable checklist of how AI storytelling currently defaults: explicit themes, linear plots, protagonist-driven endings, embodied emotion.","Per-source fingerprints allow attributing a story to a specific model with 68.4% macro-F1, and the human class is the most separable.","Rarity in narrative feature space offers a quantitative proxy for originality, which could inform discussions of authorship and creative control."],"fun_headline_variants":["Story structure alone outs AI writers 93% of the time","AI fiction detected by narrative choices, not style","Plot decisions reveal AI authorship with 93% accuracy","Narrative fingerprints: how AI stories differ from human ones"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Everything rests on the assumption that the 304 narrative features, scored entirely by an LLM, are valid measurements of real narrative structure; if those annotations encode annotator-specific biases, the observed human/AI separation may not generalize to features measured another way.","fun_headline_variants_meta":{"raw":{"variants":["Story structure alone outs AI writers 93% of the time","AI fiction detected by narrative choices, not style","Plot decisions reveal AI authorship with 93% accuracy","Narrative fingerprints: how AI stories differ from human ones"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1314,"prompt_tokens":831,"completion_tokens":483,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":418}},"tokens_in":575,"tokens_out":483,"duration_ms":5857,"temperature":1.0,"reasoning_tokens":418,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T05:32:20.716579+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a held-out set of stories from the same six sources, have human narrative analysts score the 304 features (or use a non-LLM protocol to do so), retrain the classifier, and compare macro-F1 to the LLM-annotated result. If the separation collapses or drops far below 93.2%, the claim that narrative construction itself drives detection is falsified. A complementary check: classify stories from a new, unseen language model not in the training set; if narrative features fail to separate it from humans, the 'shared AI narrative space' is not general.","supporting_citations":[],"review_version":2}