{"id":"09f0ee57-571e-477b-b666-f92ba4573c12","arxiv_id":"2411.12907","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Defines entropy and Jensen-Shannon divergence based narrative metrics and applies them to emotion dynamics in TV shows, arguing they capture complexity and pivots.","lead":"The paper proposes an information-theoretic framework for narratives, defining complexity as entropy and pivots as divergence between consecutive story states, and demonstrates it on emotion trajectories from TV shows. The authors argue these measures can help benchmark AI-generated stories, but the empirical evidence is preliminary and the prediction-based metrics are not tested.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that genre differences in entropy/JSD are narrative properties is untested because the per-show values are computed from one aggregate emotion distribution per show, so the reported between-genre separation may reflect show identity or sampling noise rather than the framework's…","rationale":"The reader identifies the deepface emotion-state operationalization as the weakest assumption; I agree that the state definition is the most fundamental conceptual load-bearing choice, and the paper explicitly concedes that the state could be anything and that only the emotion-based illustration is tested. However, the more immediately decisive weakness for the demonstrated claim is statistical: the entire empirical demonstration consists of two scalar summaries per show (one pooled entropy, one mean JSD) plotted against genre for 9 shows, with no measure of uncertainty, no episode-level variation for entropy, and no check that genre differences survive within-show or within-genre replication. This is a correctness risk in the demonstration rather than in the mathematical definitions, which are internally consistent and clearly stated (Eqs. 1-5). The prediction-focused metrics are explicitly reserved for future work and not part of the central demonstrated claim, so the claim in question is the genre-comparison claim in Section 2.1. The paper's own appendix describes the pipeline but stops short of any inferential statistics, and the small sample (as few as 3-8 episodes in several shows, 1 show per genre) means the plotted genre separation could be driven by a single show. A mixed-effects or bootstrap reanalysis would settle this directly, and the paper's stated goal of benchmarking AI-generated stories would genuinely require such a check, since a benchmark based on show idiosyncrasies would not generalize. Thus the reader's CONDITIONAL verdict is appropriate, but the condition should be expanded from 'validate the state mapping' to also include 'show that the genre differences are not artifacts of show identity or sampling noise.' I do not see internal inconsistency in the framework itself, and the paper is appropriately cautious about the prediction-based metrics, so I am not moving toward rejection. The one additional note is that the paper states 'we note that our framework is agnostic to the modality of the story told' while only demonstrating the emotion-from-faces pipeline; that claim is an extrapolation, but the reader already flagged the generalizability issue, and my concern is a more precise version of it. I would keep CONDITIONAL and state the additional statistical condition explicitly.","tokens_in":5431,"tokens_out":1838,"duration_ms":18426,"concrete_test":"Recompute per-episode entropy and per-episode mean JSD for every episode, then fit a mixed-effects model with genre as a fixed effect and show and episode as random intercepts (or run a nested bootstrap resampling episodes within shows and shows within genres). If the 95% intervals for genre coefficients overlap substantially or the intraclass correlation for show exceeds genre variance, the claimed between-genre separation is not supported by the framework's measures.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's demonstrated capability claim in Section 2.1 ('These results demonstrate our framework's ability to quantify narrative structures and emotional dynamics...') rests on between-genre differences in two scalar quantities: per-show entropy (computed from a single average emotion distribution per show, Appendix A.2) and per-show mean JSD (averaged over frames within each episode, then over episodes per show). With 1-16 episodes per show and only 9 shows, the scatter in Fig. 2D/2F conflates within-genre, between-episode, and between-show variation with the genre label. No confidence intervals about the per-show estimates, no episode-level replication check, and no statistical test (e.g., mixed-effects model with genre as fixed effect and show/episode as random effects) are reported. Because the entropy is computed on one pooled distribution per show, the measure cannot distinguish a genre signature from the idiosyncrasy of the particular shows selected, so the headline result that the framework 'quantifies narrative structures... across genres' is currently an uncontrolled descriptive observation rather than a demonstrated capability. Additionally, the pipeline gates everything on the deepface emotion distribution (Appendix A.2), so even the scalar inputs carry an unvalidated operationalization; however, the more decisive and immediately checkable gap is the absence of any error or replication structure around the genre comparisons.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an information-theoretic framework for measuring narratives, decomposing a story into states s_t and defining entropy-based complexity (Eq. 1), Jensen-Shannon divergence between consecutive states as a pivot metric (Eq. 2), and three prediction-based quantities: predictability (Eq. 3), suspense (Eq.  ąd4), and plot twist (Eq. 5). The empirical illustration uses emotion distributions extracted from actor faces via deepface in a corpus of nine TV shows totaling about 3,000 minutes, reporting per-show entropy and per-episode pivot means, and interpreting between-genre differences as evidence that the framework can quantify narrative structure and emotional dynamics. The prediction-focused measures are introduced formally but not computed.","tokens_in":5724,"tokens_out":3487,"duration_ms":39235,"significance":"If the empirical claims were fully supported, this would be a useful contribution: the definitions are simple, principled, and modality-agnostic in principle, and the framework offers concrete quantities for benchmarking AI-generated stories and for human-in-the-loop narrative tools. The paper deserves credit for making the mathematical definitions transparent and for choosing bounded, symmetric divergence measures. However, the central empirical assertion—that the framework 'demonstrate[s] its ability to quantify narrative structures and emotional dynamics across genres'—is not yet established. The reported genre separations rest on a single pooled entropy value per show, lack confidence intervals or significance tests, and depend entirely on an unvalidated deepface emotion state. The framework's theoretical core is sound, but the demonstration needs substantial statistical strengthening before the claims can be accepted.","major_comments":[{"comment":"The claim that the framework 'demonstrate[s] its ability to quantify narrative structures and emotional dynamics ... across different genres' is not supported by the reported analysis. Per-show entropy is computed on a single emotion distribution formed by averaging all states per show (Appendix A.2), so each show contributes exactly one scalar with no confidence interval or episode-level replication. With only 9 shows and 1–16 episodes per show, the separation in Fig. 2D may simply reflect show identity or sampling noise rather than a genre-level narrative property. Please report per-episode entropies and their uncertainty, add bootstrap or Bayesian intervals, and, if possible, a mixed-effects model with genre as a fixed effect and show/episode as random effects; alternatively, explicitly reframe this section as a descriptive illustration rather than a demonstrated capability.","section":"Section 2.1; Appendix A.2"},{"comment":"The pivot metric's per-show means are plotted with standard errors of the mean computed at episode level, but no statistical test is provided for whether genre differences in mean JSD are significant. With repeated frames within episodes and repeated episodes within shows, the effective sample size is much smaller than the number of frames, and naive comparisons across per-show means would have very low power. Please add an explicit repeated-measures or mixed-effects analysis, report effect sizes with confidence intervals, and state the number of independent units underlying the comparison. This is load-bearing because the genre-difference claims in Section 2.1 rest on Fig. 2D and Fig. 2F.","section":"Fig. 2F; Appendix A.2"},{"comment":"All empirical results depend on the deepface emotion distribution as the operationalization of narrative state, yet the manuscript provides no validation that this mapping is faithful or sufficiently rich for comparing genres. This is a correctness risk for the genre-difference claims: if the emotion-from-face mapping is noisy or biased across genres (e.g., due to camera framing, makeup, or acting style), the reported differences may be artifacts of the pipeline. A concrete robustness check would be re-running the main comparisons with an alternative state definition, such as text-based sentiment from subtitles or manual emotion annotations, and showing that the qualitative patterns persist. Without such a check, the empirical demonstration remains coupled to an unvalidated choice.","section":"Appendix A.2; Section 2.1"}],"minor_comments":[{"comment":"There are typographical errors: 'scence' should be 'scene', and 'comparitively' should be 'comparatively'.","section":"Section 2.1"},{"comment":"'suprising' should be 'surprising' in the sentence about the absence of a change being surprising.","section":"Section 2.2"},{"comment":"The appendix refers to 'Fig. 1A', 'Fig. 1B', 'Fig. 1C', 'Fig. 1D', and 'Fig. 1F', but the corresponding panels appear as Figure 2 in the main text. Please renumber or adjust the cross-references for consistency.","section":"Appendix A.2"},{"comment":"The sentence 'Our metrics may also assists in identifying systemic biases' has a subject-verb agreement error; it should be 'may also assist'.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The paper's theoretical proposal is reasonable and likely of interest to the computational media community, but the empirical section is currently overinterpreted relative to its statistical strength. The stress-test concern is valid and should be addressed directly: either the authors add rigorous uncertainty quantification and a mixed-effects structure, or they clearly reposition the TV-show analysis as an illustrative case study rather than a demonstration of genre-level capability. I would also encourage the authors to make data and code available, since the pipeline is central to the empirical claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a modest, clearly written proposal paper, not a breakthrough. The measures are Shannon entropy and JSD applied to emotion distributions extracted from TV frames via deepface. That's about it. The novelty is the application, not the math. The paper is honest that the prediction metrics are theoretical only.\n\nWhat it does well: the writing is clean, the definitions are explicit, the JSD choice is justified, and they cite the relevant prior work (Piper, Barron, Murdock & DeDeo). The example episode plots in Fig. 2 are illustrative. The authors don't oversell the mathematical contribution.\n\nThe soft spot is the empirical support. The claim in Section 2.1 that the results \"demonstrate our framework's ability to quantify narrative structures and emotional dynamics across genres\" rests on between-genre differences in two scalars computed from nine shows. For entropy, they pool all frames per show into one distribution (Appendix A.2), so the per-show value has no error bar and the scatter conflates genre with show identity. For JSD they give per-show means and SEM, but no significance test and no mixed-effects model with show as random effect. The stress-test note is right: with 3-16 episodes per show and no episode-level replication check, the genre separation is an uncontrolled descriptive observation, not a demonstrated capability. The deepface operationalization is another unvalidated link, but the missing error/replication structure is the immediately checkable gap.\n\nI also note they don't ship data or code, which matters for a paper whose main empirical contribution is a demo on an internal dataset. The \"heartbeat of the story\" language is fine, but the strong \"demonstrate\" language should be toned down to \"illustrate.\"\n\nThe paper is honest about its scope in the appendix, and the math is correct as definitions. It's a legitimate proposal for media analytics or AI storytelling evaluation, but it needs revision before publication: add uncertainty estimates, run a proper mixed-effects analysis, release code or at least per-show/per-episode values, and soften the claims.\n\nWho should read this: people working on narrative quantification in computational media or on benchmarks for AI-generated stories. It's a reasonable workshop or short-conference paper after those changes.\n\nMy call: send it to serious peer review, not desk reject, because the framework is clearly stated and the failure modes are fixable. I'd probably not cite it in my own work, but I'd bring it to a reading group to discuss operationalization pitfalls.","headline":"A clean but thin proposal: textbook entropy and JSD applied to TV emotion states, with an uncontrolled genre comparison; worth a serious referee but needs statistical work.","tokens_in":6201,"tokens_out":1891,"would_cite":false,"duration_ms":20618,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes that storytelling devices—complexity, pivots, suspense, plot twists—can be measured by entropy and divergence computed over narrative states, and demonstrates the first two on TV shows.","keywords":["narrative information theory","narrative state","entropy as complexity","Jensen-Shannon divergence","suspense and plot twist metrics","TV genre emotion analysis","face emotion distributions","AI-generated story benchmarking"],"falsifier":"Recalculate the entropy and pivot values on the same episodes after replacing the face-emotion state with a text-based emotion distribution from the scripts; if the genre ordering shifts or disappears, the reported results depend on the face pipeline rather than on narrative structure itself.","tokens_in":5272,"feed_emoji":"📺","tokens_out":11227,"duration_ms":103360,"temperature":0.7,"pith_summary":"The paper sets out to make narrative analysis quantitative by translating storytelling devices into information-theoretic quantities. It claims that the entropy of a story state measures complexity, the Jensen-Shannon divergence between consecutive states measures pivots or story beats, and prediction-based quantities formalize predictability, suspense, and plot twists. The empirical demonstration on a corpus of TV shows finds genre-consistent differences: reality and dating formats show broader emotion mixtures and more frequent large shifts, while dramas and thrillers are more tonally focused and change more gradually. If these measures hold up, they give creatives and AI researchers a common ruler for comparing human- and machine-made stories.","feed_headline":"Information theory puts numbers on cliffhangers and plot twists","feed_subtitle":"A story-state framework measures emotional complexity and genre shifts, putting human and AI stories on one ruler.","key_machinery":"The load-bearing object is the narrative state $s_t$: a probability distribution over whatever features describe the story at time $t$ (in the demonstration, the emotions of visible characters). On this object the paper places five information-theoretic measures: complexity $H(s_t)$, pivot $\\mathrm{JSD}(s_t, s_{t-1})$, predictability $I(s_{t+1}; S_t)$, suspense $H(P(s_{t+1} \\mid S_t))$, and plot twist $\\mathrm{JSD}(P(s_{t+1}), s_{t+1})$. The machinery works because entropy and divergence output dimensionless, comparable numbers for any state definition, so the same formulas transfer across modalities—video, text, audio—without any change to the math.","core_discovery":"The paper's central claim is that a narrative can be turned into a probability distribution over time—one state per moment—and that standard information-theoretic quantities then capture what audiences experience as story structure. It defines complexity as the entropy $H(s_t)$ of the current state, a pivot as the Jensen-Shannon divergence $\\mathrm{JSD}(s_t, s_{t-1})$ between consecutive states, and, for prediction-based devices, predictability as $I(s_{t+1}; S_t)$, suspense as the entropy $H(P(s_{t+1} \\mid S_t))$ of the predicted next state, and plot twist as $\\mathrm{JSD}(P(s_{t+1}), s_{t+1})$ between prediction and realization. Empirically, with states given by emotion distributions read from actors' faces in over 3000 minutes of TV, the paper finds that reality and dating shows have higher entropy and larger average pivots, while dramas and thrillers are lower on both. On these grounds the paper argues the framework quantifies narrative structures and emotional dynamics across genres and can therefore compare human-created and AI-generated stories.","pith_inferences":["Replacing the face-emotion state with a text-based emotion model of the same scripts and checking whether the genre rankings replicate is a testable way to show the results are about narratives rather than about faces.","The paper defines but never computes the suspense and plot-twist metrics; an immediate implementation would use an LLM's next-token distribution over scene summaries and check whether cliffhanger endings raise the suspense entropy.","If the measures stabilize across studies, they could be inverted into a generation objective—steering a story generator toward a target suspense or twist profile—an application the paper does not claim."],"forward_implications":["The entropy of a narrative state gives a quantitative complexity score, so scenes dominated by one emotion score low and emotionally mixed scenes score high.","The Jensen-Shannon divergence between consecutive states localizes story beats, allowing editors and summarizers to find pivotal moments automatically.","Genre profiles from the demonstration—reality shows more emotionally mixed and shift-prone, dramas and thrillers more tonally focused and gradual—can serve as baselines for judging whether an AI-generated story fits a requested genre.","Because the framework treats states as arbitrary distributions, the same definitions transfer from video to text or audio without changing the math.","Once a generative model of story continuation is supplied, the prediction-based metrics provide formal, computable definitions of predictability, suspense, and plot twist."],"supporting_citations":[{"why":"Supplies the state-revelation view of narrative that the framework's decomposition of a story into a sequence of states builds on.","marker":"[7]"},{"why":"Provides the information-theoretic treatment of sequences and predictions over states that underlies the pivot and prediction metrics.","marker":"[9]"},{"why":"Motivates novelty and surprise as the storytelling ingredients the framework is designed to capture.","marker":"[8]"},{"why":"Supports identifying entropy with the complexity of a state, the definition used for the complexity measure.","marker":"[17]"},{"why":"Justifies the choice of Jensen-Shannon divergence over Kullback-Leibler because symmetry and boundedness make it suitable as a metric.","marker":"[18]"},{"why":"Anchors the pivot and cliffhanger application by showing that pivotal moments can be detected in multimodal video content.","marker":"[20]"}],"fun_headline_variants":["Entropy quantifies plot twists and cliffhangers","Info theory measures narrative complexity","Story structure scored by information theory","Cliffhangers and twists on a probability scale"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The empirical argument assumes that the emotion mix detected on actors' faces and averaged over rolling windows is a faithful and sufficiently rich proxy for the narrative state at each moment, so all genre differences inherit whatever this mapping gets wrong.","fun_headline_variants_meta":{"raw":{"variants":["Entropy quantifies plot twists and cliffhangers","Info theory measures narrative complexity","Story structure scored by information theory","Cliffhangers and twists on a probability scale"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000931,"raw_usage":{"total_tokens":3922,"prompt_tokens":817,"completion_tokens":3105,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":433,"completion_tokens_details":{"reasoning_tokens":3051}},"tokens_in":433,"tokens_out":3105,"duration_ms":26373,"temperature":1.0,"reasoning_tokens":3051,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:03:03.723669+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recalculate the entropy and pivot values on the same episodes after replacing the face-emotion state with a text-based emotion distribution from the scripts; if the genre ordering shifts or disappears, the reported results depend on the face pipeline rather than on narrative structure itself.","supporting_citations":[{"cited_title":"Modeling Narrative Revelation,","cited_arxiv_id":null,"evidence_quote":"Supplies the state-revelation view of narrative that the framework's decomposition of a story into a sequence of states builds on."},{"cited_title":"Individuals, institutions, and innovation in the debates of the French Revolution,","cited_arxiv_id":null,"evidence_quote":"Provides the information-theoretic treatment of sequences and predictions over states that underlies the pivot and prediction metrics."},{"cited_title":"Exploration and exploitation of Victorian science in Darwin’s reading notebooks,","cited_arxiv_id":null,"evidence_quote":"Motivates novelty and surprise as the storytelling ingredients the framework is designed to capture."},{"cited_title":"Relating objective complexity, subjective complexity and beauty,","cited_arxiv_id":null,"evidence_quote":"Supports identifying entropy with the complexity of a state, the definition used for the complexity measure."},{"cited_title":"RADio* – An Introduction to Measuring Normative Diversity in News Recommendations,","cited_arxiv_id":null,"evidence_quote":"Justifies the choice of Jensen-Shannon divergence over Kullback-Leibler because symmetry and boundedness make it suitable as a metric."},{"cited_title":"Find the Cliffhanger: Multi- modal Trailerness in Soap Operas,","cited_arxiv_id":null,"evidence_quote":"Anchors the pivot and cliffhanger application by showing that pivotal moments can be detected in multimodal video content."}],"review_version":1}