{"id":"5c6d2ed0-44a4-4931-b91d-174618c16abc","arxiv_id":"2508.15794","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Language models can classify suspenseful stories but cannot reproduce human ratings of suspense magnitude or arc across story segments, and their judgments diverge from humans under text permutation.","lead":"Researchers asked language models to rate suspense in stories the way human readers did in four classic psychology experiments, and compared the answers. The models could tell suspenseful from non-suspenseful text but did not match human ratings of how suspense builds and falls, and scrambled text changed their judgments in ways humans' were not.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Likert-scale commensurability is the load-bearing assumption: before concluding LMs cannot estimate relative suspense, the analysis must show the negative result is robust to monotonic calibration and rank-based agreement.","rationale":"The reader's weakest assumption—scale commensurability between LM outputs and human ratings—is exactly the load-bearing point. The paper's external-replication design is a strength, and the abstract is appropriately hedged, but the central negative claim depends on comparing numbers produced by different systems as if they were on the same measurement scale. Since the full text is unreadable in the supplied material, I cannot verify whether the authors already used rank correlations or calibration; hence the correct verdict remains UNVERDICTED rather than moving to accept/reject. My concern does not change the reader's verdict because the reader already identified this assumption as the weakest point. The concrete test is feasible because the four human datasets are from published studies and the LM outputs should be reproducible from the prompt protocol.","tokens_in":15011,"tokens_out":5085,"duration_ms":58361,"concrete_test":"Using the authors' datasets from the four replicated studies, recompute agreement between LM ratings and human ratings with Spearman/Kendall rank correlations per story or text segment, and repeat the main comparative analyses after per-model monotonic calibration (e.g., quantile-matching each LM's ratings to the human marginal distribution). If rank-order agreement is high, or if calibration substantially reduces the divergence, then the headline negative result fails. If the divergence persists under both rank-based and calibration-invariant analyses, the scale-commensurability concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that LMs can classify suspense-intended text but cannot accurately estimate relative suspense magnitude or trajectory compared with humans. This is a quantitative comparison between LM outputs and human ratings. The abstract describes substituting human responses with LM responses, and the visible table fragments ('Human Study Result' vs. 'LM Result') suggest direct numeric comparison. If LM ratings were elicited on a Likert-type scale and compared as raw scores or with Pearson-style agreement, documented LM scale-use biases—central tendency, range compression, anchor sensitivity—could produce large divergence even when LMs rank-order the same segments as humans. The binary classification success does not control for this, because category discrimination is far less sensitive to scaling than graded magnitude comparison. The final inference that LMs 'do not process suspense in the same way as human readers' is even stronger than rating divergence alone can support. Without a rank-based or calibration-invariant analysis, the reported dissociation may be an artifact of the prompt-and-scale setup rather than a genuine failure of suspense understanding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper replicates four published psychological studies of narrative suspense by substituting human participants with large language models (open-weight and closed-source). The central claim is a dissociation: LMs can discriminate whether a text is intended to induce suspense in humans, but they cannot accurately estimate the relative amount of suspense within a text sequence, nor can they reproduce human-rated rise-and-fall trajectories across segments. Adversarial permutation of story order is used to probe why LM suspense judgments diverge from human judgments, and the authors conclude that LMs do not process suspense the way human readers do.","tokens_in":15063,"tokens_out":4495,"duration_ms":51879,"significance":"If the dissociation holds, it would be a useful empirical boundary for affective narrative understanding in LLMs, and the design is commendable: replicating four external studies with multiple models reduces the risk of benchmark-specific artifacts, and the binary-vs-graded contrast provides an internal control. The permutation experiments are a plausible way to probe order sensitivity. The paper also has the virtue of making a falsifiable negative claim about model capabilities. However, verification is currently blocked both by a corrupted full-text rendering and by the absence, in the readable portions, of any calibration-invariant analysis linking LM outputs to the original human Likert scales. The central claim is therefore plausible but not yet established.","major_comments":[{"comment":"The headline claim 'LMs cannot accurately estimate the relative amount of suspense' is a quantitative comparison between LM outputs and human Likert ratings from the four replicated studies. The visible comparison block (columns 'Human Study Result' vs 'LM Result') shows raw values but no rank-based or calibration-invariant agreement scores. LLM Likert responses are known to exhibit central-tendency bias, range compression, and anchor sensitivity; a low raw Pearson correlation could occur even when the LM rank-orders segments identically to humans. Please report Spearman/Kendall correlations and a monotonic-calibration robustness check (e.g., isotonic regression mapping LM scores to human ratings) and show that the graded-magnitude dissociation remains.","section":"Abstract and Results tables"},{"comment":"The final inference that LMs 'do not process suspense in the same way as human readers' goes beyond the behavioral evidence reported. Divergent ratings under scrambled text could arise from lower-level surface statistics, recency effects, or prompt-induced local coherence biases, rather than a difference in suspense-specific processing. The permutation results would be more convincing if tied to a priori predictions about what features each model class would use; otherwise the conclusion should be hedged to 'their judgments depend on different aspects of text order/structure.'","section":"Conclusion"},{"comment":"The supplied manuscript text is largely corrupted mojibake, and the Limitations section is not readable. If the four seminal studies' stimulus texts are published in accessible sources, they are very likely present in the pretraining corpora of the evaluated open-weight and closed-source LMs. The abstract does not describe a contamination check. The authors should explicitly state whether the exact stimuli (or near-duplicates) appeared in pretraining data and discuss the direction of bias: contamination would inflate binary classification success and would obscure any clean interpretation of the graded-magnitude failure. This must be legible in the final text.","section":"Limitations section"}],"minor_comments":[{"comment":"The full text appears to be encoded incorrectly; several paragraphs, including the Limitations and Conclusion sections, are mojibake. This must be fixed before any substantive review.","section":"Throughout"},{"comment":"The phrase 'to identify what cause human and LM perceptions of suspense to diverge' needs grammatical revision: 'what causes' or 'the causes of divergence'.","section":"Abstract"},{"comment":"The repeated table rows (e.g., 'Reading Time', 'MC' and 'SAT' rows) are visually cluttered and not fully labeled. Provide clear column definitions and mark which statistics are LM-generated versus taken from the original studies.","section":"Tables"},{"comment":"The text should name the four seminal studies in the Abstract or Introduction; due to corruption I could not verify they are identified in the body. If they are not, please add them.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The version under review is unreadable in large parts (mojibake), so I had to rely on the abstract and visible table fragments. Even within those readable parts, the central claim depends on a scale-comparison between raw LM outputs and human Likert ratings; no rank-based or calibration-invariant analysis is described. If the full text also lacks such an analysis, the headline result may be an artifact of LLM scale-use biases rather than a genuine failure of suspense understanding. I would send the manuscript back with a request to fix the encoding and to add a calibration robustness section. The paper is not rejectable on the merits; the empirical design is strong and the dissociation is interesting if it survives."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. This paper (the abstract at least) makes a sharp, falsifiable claim: language models can tell whether a story is intended to be suspenseful, but they fail to reproduce human graded suspense ratings and trajectory arcs across four replicated psychology studies. That is a useful boundary result for the LLM-as-participant program. The adversarial permutation probe is a nice addition — testing whether flipping text order breaks LM judgments more than human judgments is a clear way to look for sensitivity to surface structure.\n\nWhat's new: the specific dissociation has not been documented at this scale before, and the external human benchmark means the result isn't circular. The authors are appropriately hedged in the abstract.\n\nWhere the soft spot is: the headline negative result compares LM scalar outputs to original human Likert ratings. That's a measurement-equivalence assumption. LLMs are known to compress ranges and anchor to scale wording. If the analysis is only raw-score agreement or Pearson-style correlation, a large divergence could emerge even when the LMs rank-order the same segments as humans. We need to see a rank-based or calibration-invariant analysis (Spearman, monotonic transform, rank agreement on trajectory) before believing the claim that LMs can't estimate relative suspense. The binary classification result doesn't control for this, because category discrimination is far less sensitive to scaling than graded comparison. Also, the final sentence — 'LMs do not process suspense in the same way as human readers' — goes beyond rating divergence; rating differences alone can't tell you how an LM processes text internally.\n\nOne unfortunate note: in the version I saw, the body renders as mojibake, so I could only assess the abstract and table fragments. That forces an abstract-level verdict, low confidence.\n\nWho's it for: anyone running LLM-as-participant studies or building narrative evaluation benchmarks. It will likely be cited as a cautionary case. It deserves serious refereeing — the question is timely and the data is external — but I'd instruct the referee to focus on the scale-equivalence and robustness analyses. If the authors haven't run rank-based checks, they should.","headline":"A clear, testable dissociation claim from the abstract — LMs pass binary suspense detection but fail graded magnitude and trajectory — deserves peer review, but the key comparison hinges on scale calibration and the methods were not readable in this rendering.","tokens_in":15701,"tokens_out":2183,"would_cite":false,"duration_ms":22049,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Language models can recognize suspense as a category but cannot reproduce human judgments of how much suspense a story passage contains or how suspense rises and falls across a text.","keywords":["suspense","language models","affective understanding","narrative comprehension","human-AI agreement","replication study","natural language processing"],"falsifier":"Take the same stories and segment boundaries used in the paper, but replace numeric ratings with a ranking task: ask human readers and LMs to order the segments from least to most suspenseful, then compare the rank orders. If LMs reproduce the human rank order across many stories, the claim that LMs cannot estimate the relative amount of suspense is falsified; if their rank orders diverge, the claim is supported.","tokens_in":14752,"feed_emoji":"📖","tokens_out":8293,"duration_ms":83041,"temperature":0.7,"pith_summary":"The paper attempts to establish a boundary on language models' affective understanding: LMs can tell when a story is meant to be suspenseful, but they do not perceive suspense the way human readers do, either in degree or over time. The authors reach this conclusion by replicating four established psychological studies of human suspense perception, substituting human participants with open-weight and closed-source language models and comparing model ratings with the original human ratings. The models succeed at the coarse task of separating suspenseful from non-suspenseful texts, yet their numeric estimates of how much suspense a passage carries diverge from human estimates, and their suspense curves across story segments do not show the human-shaped rise and fall. Adversarial permutation of story text shows that the ordering of events matters differently for LMs than for human readers. A reader should care because the result undermines the use of LM ratings as a stand-in for reader experience in narrative studies, automated story evaluation, or engagement prediction.","feed_headline":"Language models spot suspense but can't grade it like humans","feed_subtitle":"They can tell a suspenseful story, yet their ratings miss how human tension rises and falls.","key_machinery":"The load-bearing mechanism is an LM-as-participant replication design. The authors take four established human studies of suspense perception—their stories, segment boundaries, and rating tasks—and replace the human responses with ratings produced by LMs, so that model judgments can be compared directly with the human judgments the original studies recorded. The second mechanism is adversarial permutation: story segments are reordered to test whether LMs and humans are sensitive to the same ordering of narrative events. Together these mechanisms let the paper separate recognition of suspense as a topic from perception of its magnitude and trajectory.","core_discovery":"The paper's central claim is that current language models have only a superficial grasp of narrative suspense. Given the same story segments that human participants rated, LMs correctly identify which passages are designed to induce suspense, but they cannot accurately estimate the relative amount of suspense within a text sequence compared with human judgments, and they fail to reproduce the human perception of suspense rising and falling across multiple segments. The divergence is systematic: when the authors adversarially permute the order of story text, LM suspense responses move in ways human perceptions would not. The paper concludes that LMs can superficially identify and track certai","pith_inferences":["The paper's negative result could partly reflect how LMs use numeric rating scales—central tendency, range compression, or anchor wording—rather than a true absence of human-like suspense perception. A calibration pass that matches LM score distributions to the original human rating distributions would separate these explanations.","A natural extension is to replace absolute Likert ratings with forced-choice ranking: if LMs can rank story segments by suspense as humans do, then the reported deficit lies in scale use rather than perception; if ranking also fails, the deficit is perceptual.","The permutation results point to a broader diagnostic for narrative understanding: models that genuinely track story structure should show human-like changes in suspense when the resolution is moved before the buildup, and the paper's method makes that test straightforward.","For applied systems, the result implies that any pipeline using LM affect scores to edit stories, generate reader-engagement predictions, or summarize narrative tension will inherit misplaced suspense peaks, so human validation remains necessary at the point of use."],"forward_implications":["If the claim is right, LMs can serve as coarse binary filters for 'is this text intended to be suspenseful' but not as continuous annotators of suspense intensity.","Automated systems that use LM affect ratings to predict reader engagement, locate a story's climax, or evaluate pacing will systematically misplace where and how much suspense a reader would feel.","Because the deficit appears in graded magnitude and trajectory rather than in category detection, benchmarks of LM affective understanding should separate classification accuracy from agreement with human continuous ratings.","The permutation experiments imply that LM suspense recognition depends on local textual cues more than on global narrative order, so order-sensitive story understanding remains a gap.","The limitation appears across both open-weight and closed-source LMs tested, making it a property of current LM behavior rather than a quirk of one model family."],"supporting_citations":[],"fun_headline_variants":["AI sees suspense but misses human tension curves","Chatbots detect suspense, but can't score its rise and fall","Language models flunk suspense grading tests","When stress rises in stories, AI loses track","AI knows a thriller, but not how it thrills"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the number a language model produces for 'how suspenseful is this?' sits on the same measuring scale as the human ratings from the original studies, so a gap between the two numbers is interpreted as a gap in suspense perception rather than a difference in how models use rating scales.","fun_headline_variants_meta":{"raw":{"variants":["AI sees suspense but misses human tension curves","Chatbots detect suspense, but can't score its rise and fall","Language models flunk suspense grading tests","When stress rises in stories, AI loses track","AI knows a thriller, but not how it thrills"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000625,"raw_usage":{"total_tokens":2690,"prompt_tokens":666,"completion_tokens":2024,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":410,"completion_tokens_details":{"reasoning_tokens":1950}},"tokens_in":410,"tokens_out":2024,"duration_ms":15171,"temperature":1.0,"reasoning_tokens":1950,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T21:04:18.045635+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same stories and segment boundaries used in the paper, but replace numeric ratings with a ranking task: ask human readers and LMs to order the segments from least to most suspenseful, then compare the rank orders. If LMs reproduce the human rank order across many stories, the claim that LMs cannot estimate the relative amount of suspense is falsified; if their rank orders diverge, the claim is supported.","supporting_citations":[],"review_version":1}