{"id":"db6e9d5b-ae75-49e8-9e12-b033bdec6615","arxiv_id":"2508.16190","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"ComicScene154 is a new 154-page comic dataset with manual scene-boundary annotations; human agreement is moderate (pk=0.17) and the provided baseline only slightly beats random.","lead":"The authors present ComicScene154, a manually labeled dataset marking scene boundaries in 154 pages of public-domain comics. It is the first scene segmentation resource for comics, but the labels are subjective and the baseline system barely outperforms chance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Annotation reliability is the load-bearing assumption: pk=0.17 without chance-adjusted agreement leaves the benchmark's gold standard underjustified.","rationale":"Good-faith reading: The paper is a dataset paper. Its claim is not that scene segmentation is solved, but that the resource is useful. The authors are transparent about subjectivity and include a limitations section. The baseline experiments are honest. The weakest link is the ground truth itself: the value of a benchmark is bounded by the reliability of the labels. The reader flagged this same issue, and I agree. However, the available numbers do suggest signal: random-vs-ground truth pk is about 0.46, while author-vs-tester pk is 0.17, so testers are much closer to authors than random is. That is evidence in the dataset's favor. The gap is that this comparison is informal and no chance-adjusted metric is reported; the vague annotation protocol adds a construct-validity risk. These are validation gaps, not known errors. Therefore the reader's ACCEPT verdict remains appropriate, but the tests above would materially strengthen the paper.","tokens_in":6960,"tokens_out":11771,"duration_ms":134153,"concrete_test":"Re-annotate a random subset of at least three excerpts with three annotators using the explicit scene definition from Section 2.1 plus worked examples, then compute pairwise pk and Cohen's kappa on panel-boundary indicators, and compare against a Monte Carlo random-baseline with matched scene counts (e.g., 10,000 draws). If the improved-protocol pairwise pk is still >0.15 or the kappa is below ~0.4, the annotations are too unreliable to serve as a benchmark; if kappa is >0.6 and pk is far below the random distribution, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"ComicScene154's central value as a benchmark depends on the reliability of the authors' scene annotations. Section 3.3 reports only pk=0.17 average agreement between authors and testers (and 0.21 tester-tester) on one-third of the data, with one excerpt at 0.37. No chance-adjusted measure (Cohen's kappa/Krippendorff's alpha) or boundary-level precision/recall is given, so pk=0.17 is not interpretable: whether this is above chance depends on the boundary distribution and window size. The random baseline in Table 4 is described only as 'scenes we defined randomly' and is not used as a benchmark for human agreement. Additionally, the annotation guideline in Appendix 7.2 does not contain the formal scene definition from Section 2.1, so labels may reflect intuitive segmentation rather than the intended narrative-arc construct. None of this implies bad faith; it means the gold-standard quality is under-validated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ComicScene154, a manually annotated dataset for scene-level narrative segmentation in comics. It comprises 154 pages from four public-domain comics (1942–1962), 34 stories, and roughly 954 panels, with panel-level binary labels marking scene starts. Inter-annotator reliability is measured on one-third of the data using the pk metric (average tester-vs-author pk = 0.17; tester-tester pk = 0.21). The authors also propose a two-stage baseline using Gemini 2.0 Flash Thinking to predict scene boundaries and then refine them; the baseline performs close to random (average pk 0.42 vs random 0.46), with refinement improving the best iterations (average 0.34) but not the average. The paper argues the dataset is a valuable resource for multimodal narrative understanding despite subjectivity and the Golden-Age skew.","tokens_in":7210,"tokens_out":6371,"duration_ms":66072,"significance":"If the gold annotations are accepted as reliable, ComicScene154 fills a clear gap: no existing comic dataset provides narrative scene segmentation with panel-level boundaries, and the task connects to video scene segmentation and semantic text segmentation. Strengths include public-domain reproducible sourcing, an explicit inter-annotator reliability effort, honest reporting of near-random baseline results, and no parameter tuning to inflate scores. The dataset and code are released. The main risk is that the reliability evidence is too thin to support the benchmark claim, and the construct validity of the annotation guidelines is questionable. These concerns are addressable and do not require collecting a fundamentally different dataset.","major_comments":[{"comment":"The central claim that ComicScene154 is a reliable benchmark rests on the inter-annotator agreement, but only pk is reported (avg 0.17 tester-vs-author; 0.21 tester-tester). pk is not chance-adjusted; without the boundary density of the gold annotations and a random-segmentation baseline computed against the same gold, pk=0.17 cannot be interpreted. The random baseline in Table 4 is for model outputs vs human gold, not for human annotators vs human gold, so the Discussion's statement that human agreement is a 'notable improvement compared to randomly defined scenes' (§5) is not directly supported. Please report chance-level pk, a chance-corrected coefficient, boundary precision/recall, and a human-vs-human random baseline.","section":"Section 3.3, Tables 3–4"},{"comment":"The annotation guideline given to testers does not contain the formal scene definition of §2.1 (a plot-based semantic unit pursuing an overarching task with a certain cast, analogous to Cohn's narrative arcs). It only asks participants to mark perceived transitions and says deviations are acceptable. Consequently, the gold-standard labels may operationalize intuitive page/panel segmentation rather than the intended narrative-arc construct. At minimum, include the formal definition in the guideline and describe how testers were briefed; ideally add a validation of the construct, e.g., whether scene boundaries correlate with changes in characters or goals.","section":"Appendix 7.2 / Section 3.3"},{"comment":"The 'random' baseline is under-specified. 'Scenes we defined randomly' does not state how many boundaries were sampled, whether the number of scenes matched the human annotations, or what distribution was used. Since the model's advantage over random is only 0.03–0.04 pk, the conclusion that the model is 'marginally better than random' depends entirely on this baseline. Provide the exact random-generation procedure and, ideally, multiple random draws with confidence intervals.","section":"Section 4, Table 4"}],"minor_comments":[{"comment":"The formula '0.17 = 0.15 + 0.19/2' should read '(0.15 + 0.19)/2'; as written the arithmetic is incorrect.","section":"Section 3.3"},{"comment":"Define what the 'In-between' column denotes (author–tester vs tester–tester) in the caption; currently the reader must infer it from §3.3.","section":"Table 3"},{"comment":"The phrase 'three different groups of two annotators' is confusing in light of Table 3's '2 of 6 tester'; clarify how many annotators participated and how excerpts were assigned.","section":"Section 3.3"},{"comment":"The window size k=3 is derived from the same annotations used for evaluation. Report the sensitivity of the reported pk values to k (e.g., k=2,4).","section":"Section 3.3"},{"comment":"Typos: 'Details on of' in the Table 2 caption; 'comptutational' in the Introduction; 'V olume' in the Kamath et al. reference.","section":"Table 2 / References"},{"comment":"The dataset link is a GitHub repository; for archival stability, provide a DOI or a permanent repository entry.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a dataset contribution; the main concern is reliability validation, not scope or novelty. The authors should be encouraged to add chance-corrected agreement and to clarify the random baseline. I see no ethical concerns beyond the acknowledged annotator-bias issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is the first labeled scene-segmentation dataset for comics, and that is the real contribution. The paper is honest about its limits and reports baseline performance that is barely above random. Second thing: the annotations that are the whole point of the dataset are not as solid as they need to be. pk=0.17 is reported without any chance-adjusted measure, so I can't tell whether that's actually good agreement or just above noise. That's the soft spot to worry about.\n\nWhat's new and done well: the dataset fills a gap. Table 1 shows no prior comic dataset has scene-level annotations. They use public-domain sources, extract panels in reading order, and provide binary scene-start labels. The construction is transparent. They measure inter-annotator agreement with pk, and the average 0.17 is clearly better than the random baseline of 0.42 they report, though that comparison isn't apples-to-apples. Baselines using a multimodal model plus LLM refinement are reported as near-random and they don't oversell them. No sign of circular reasoning; the only mildly circular bit is computing k from the same annotations used for evaluation, but that's a minor issue.\n\nSoft spots: the load-bearing assumption is annotation reliability. Only one-third of the data has multi-annotator labels. One excerpt hit 0.37, which is close to random. There's no Cohen's kappa or Krippendorff's alpha, and no boundary-level precision/recall. The pk=0.17 number is not interpretable without knowing the chance level for this boundary distribution and window size. Also, the annotation guideline in Appendix 7.2 doesn't state the formal scene definition from Section 2.1; it asks annotators to mark transitions intuitively. That raises a question about whether the labels actually reflect the 'narrative arc' construct the paper claims. The small size and Golden Age skew are acknowledged limitations, not hidden. The refined baseline's best-iteration results are cherry-picking, but they report averages too.\n\nWho should read this: anyone working on comic analysis, multimodal narrative, or segmentation evaluation. The dataset is a useful seed even with small scale. It deserves a serious referee; the right fix is for the authors to add chance-adjusted agreement metrics, report per-boundary agreement, and align the annotation guideline with the formal definition. I'd take it in with revisions, not reject it.","headline":"A genuinely new but small comic scene dataset with honest baselines; the main risk is that the gold-standard annotations are under-validated.","tokens_in":7659,"tokens_out":2340,"would_cite":true,"duration_ms":24196,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ComicScene154 introduces manually labeled scene boundaries for 154 pages of public-domain comics, making narrative segmentation in comics a measurable task.","keywords":["comic scene segmentation","narrative arcs","multimodal narrative understanding","dataset benchmark","public-domain comics","inter-annotator agreement","pk metric","scene boundary detection"],"falsifier":"Recruit independent annotators to label all 154 pages and compute pairwise pk among them, then compare model-to-label pk against this human-human agreement. If human-human pk on the full dataset approaches the random baseline of 0.4 rather than the reported 0.17, the scene boundaries are not a stable ground truth; if model scores match human-human agreement, the benchmark is already at the ceiling set by annotation subjectivity.","tokens_in":1509,"feed_emoji":"💬","tokens_out":1477,"duration_ms":54350,"temperature":0.7,"pith_summary":"The paper introduces ComicScene154, a manually annotated dataset of 154 pages from four public-domain comic magazines in which each panel is marked as starting a new scene or not. Scenes are defined as plot-based semantic units where a cast pursues an overarching task, borrowing a definition from movie scene segmentation and treating scenes as narrative arcs. The authors argue this fills a gap: existing comic datasets annotate pages, objects, or characters, not narrative structure, and page-level divisions miss story boundaries. To show the dataset is usable, they build a two-stage baseline—a multimodal model that proposes scene boundaries and a reasoning language model that refines them—and evaluate with the pk segmentation metric, which measures how often two segmentations disagree on sliding windows. Their results show human annotators agree moderately (average pk 0.17 versus roughly 0.4 for random boundaries), while the baseline remains near random, which they read as evidence that scene segmentation is a hard, still-subjective task that the dataset can now make measurable.","feed_headline":"New dataset maps comic-book scenes for narrative AI","feed_subtitle":"Manual scene boundaries across 154 public-domain pages give multimodal models a benchmark beyond page-level cuts.","key_machinery":"The panel-level scene-start tag: each panel in the 154 pages is numbered in reading order and labeled with a Boolean indicating whether it begins a new scene, or narrative arc. This single mechanism converts the abstract notion of 'scene' into a countable segmentation task, measurable by the pk sliding-window agreement metric adapted from text segmentation, and comparable across human annotators and model outputs.","core_discovery":"On its own terms, the paper's central claim is that ComicScene154 provides a comic dataset labeled at the level of narrative scenes rather than pages or panels, and that this level is the right granularity for multimodal narrative understanding. The discovery is the annotation resource itself: 34 full stories from four Golden Age comics, totaling 154 pages, with each panel numbered in reading order and carrying a Boolean scene-start label. The paper does not claim its baseline solves scene segmentation; it claims the dataset makes the task well-defined and benchmarkable, and documents that both humans and a state-of-the-art multimodal-plus-reasoning pipeline settle only slightly above random","pith_inferences":["Because human annotators disagree substantially, the dataset's reliability could be improved by reporting multiple annotation views or a consensus labeling rather than a single ground truth; the authors leave this implicit.","The comic-as-video-compression analogy suggests a direct transfer test: sample movie frames into comic-like panel sequences and see whether scene boundaries learned on comics predict movie scenes.","The measured inter-annotator agreement of pk 0.17 sets an upper ceiling on how closely any model can match the official labels; future evaluations should compare model scores against human-human agreement, not just against random segmentation.","The Golden Age source material limits the benchmark to older storytelling and art styles, so extending it to modern comics would test how well the scene construct generalizes; the authors acknowledge this limitation."],"forward_implications":["Scene segmentation in comics can be evaluated quantitatively with the pk metric, giving future work an objective benchmark rather than page-level heuristics.","Annotating full stories rather than random pages is necessary for scene construction, because narrative arcs cross page boundaries.","Human annotators agree moderately (average pk 0.17 against roughly 0.4 for random boundaries), but the lack of an intersubjective scene definition remains the core obstacle.","A multimodal model plus reasoning-based refinement performs only slightly better than random scene boundaries (average pk 0.39-0.46), despite high self-consistency across iterations.","The dataset provides a foundation for story summarization, character identification, and entity tracking at narrative scale, and for applying comic-style compression to movie scene segmentation."],"supporting_citations":[{"why":"Supplies the scene definition and the movie-scene analogy that the dataset adopts.","marker":"(Rao et al., 2020)"},{"why":"Defines narrative arcs in comics, the structural units that the paper treats as scenes.","marker":"(Cohn, 2013)"},{"why":"Provides the pk segmentation metric used for both reliability assessment and benchmark evaluation.","marker":"(Glavaš et al., 2016)"},{"why":"Establishes scene detection in fiction as a segmentation task that the comic version extends.","marker":"(Zehe et al., 2021)"},{"why":"Supplies the comparative overview of comic datasets that positions ComicScene154's new scene-segmentation task.","marker":"(Vivoli et al., 2024)"}],"fun_headline_variants":["Scene-level comic dataset for narrative AI","ComicScene154: scenes not panels for story AI","154 pages of comics with scene boundaries for AI","Dataset flips comic analysis from pages to scenes"],"cache_read_input_tokens":9472,"weakest_assumption_plain":"The authors' own manual scene annotations, used as ground truth for all evaluations, are reliable enough to serve as a benchmark even though independent annotators agree with them only mildly (average pk of 0.17 on one-third of the dataset).","fun_headline_variants_meta":{"raw":{"variants":["Scene-level comic dataset for narrative AI","ComicScene154: scenes not panels for story AI","154 pages of comics with scene boundaries for AI","Dataset flips comic analysis from pages to scenes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000111,"raw_usage":{"total_tokens":833,"prompt_tokens":623,"completion_tokens":210,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":367,"completion_tokens_details":{"reasoning_tokens":163}},"tokens_in":367,"tokens_out":210,"duration_ms":2868,"temperature":1.0,"reasoning_tokens":163,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:26:53.669821+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recruit independent annotators to label all 154 pages and compute pairwise pk among them, then compare model-to-label pk against this human-human agreement. If human-human pk on the full dataset approaches the random baseline of 0.4 rather than the reported 0.17, the scene boundaries are not a stable ground truth; if model scores match human-human agreement, the benchmark is already at the ceiling set by annotation subjectivity.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the scene definition and the movie-scene analogy that the dataset adopts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines narrative arcs in comics, the structural units that the paper treats as scenes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes scene detection in fiction as a segmentation task that the comic version extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the comparative overview of comic datasets that positions ComicScene154's new scene-segmentation task."}],"review_version":1}