{"id":"2c48efa5-d005-4298-9ad7-0b9402e440ba","arxiv_id":"2508.14941","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"Semantic normalization of hierarchical manga knowledge graphs reduces label inconsistency and improves reasoning tasks, according to abstract-level claims.","lead":"This paper proposes a normalization step that merges similar action and event labels in comic story graphs, using word similarity and vector embeddings. It aims to make symbolic reasoning over visual narratives more consistent and robust while keeping graphs human-readable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed reasoning improvement is unverifiable as supplied: the full text is garbled and the abstract provides no metrics, baselines, or significance tests.","rationale":"The reader's verdict of UNVERDICTED is appropriate: with the full text unreadable, the central empirical claim cannot be assessed. I identify the same bottom line but frame the load-bearing issue as missing experimental evidence rather than the semantic-equivalence assumption. The semantic-equivalence failure mode is real, but it is not the most direct blocker; even a perfect clustering would not establish the headline unless the downstream evaluations are shown to be meaningful and compared against a proper raw-graph baseline. The concrete test targets exactly that gap by checking for metrics, statistics, and ablations in the clean PDF. No conclusion about author intent is drawn; the text simply does not support verification as supplied.","tokens_in":7984,"tokens_out":3109,"duration_ms":42474,"concrete_test":"Retrieve the clean PDF from arXiv:2508.14941 and inspect the experiments section. Verify that (1) exact task metrics are reported for normalized versus raw graphs on the same annotation graphs; (2) significance tests or confidence intervals are provided; and (3) an ablation separates lexical similarity from embedding clustering while controlling for reduced label count. If any of these are absent, the headline improvement remains unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires showing that semantically normalized hierarchical graphs outperform raw graphs on action retrieval, character grounding, and event summarization. The only inspectable evidence is the abstract, which calls the evaluations 'preliminary' and asserts improvements without reporting numbers, baseline definitions, error bars, or statistical tests. Because the supplied full text is mojibake, I cannot check whether normalization is compared with the raw graph under matched conditions, whether the reported gains exceed annotation noise, or whether 'maintaining symbolic transparency' is measured rather than asserted. The reader's semantic-equivalence concern is a plausible failure mode, but it is secondary: even a semantically valid clustering could artificially improve aggregate metrics simply by reducing label cardinality, and the supplied text gives no way to rule that out. The load-bearing vulnerability is therefore not a specific algorithmic step but the absence of inspectable quantitative support for the empirical headline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a semantic normalization framework for hierarchical narrative knowledge graphs derived from manga annotations. The method consolidates semantically related action/event labels using lexical similarity and embedding-based clustering, applied to panel-level, event-level, and story-level graphs built from Manga109. The abstract claims that preliminary evaluations on action retrieval, character grounding, and event summarization show improved coherence and robustness while maintaining symbolic transparency. However, the supplied full text is corrupted and unreadable (mojibake), so the methodology, experimental setup, and quantitative results cannot be inspected. The reviewable portion is therefore limited to the abstract.","tokens_in":8233,"tokens_out":3744,"duration_ms":46984,"significance":"If the claimed improvements were demonstrated with rigorous, reproducible experiments, the proposal would be a useful preprocessing step for symbolic multimodal narrative reasoning, and the application to Manga109 is a relevant test-bed. The promise of combining symbolic transparency with normalization is appealing. That said, the manuscript as supplied provides no numeric evidence, baselines, error analysis, or statistical tests; the central empirical claim is currently unsupported. No machine-checked proofs, reproducible code, or falsifiable quantitative predictions are visible. The potential significance is real but entirely conditional on evidence that the current submission does not make available for verification.","major_comments":[{"comment":"The headline claim that semantic normalization 'improves coherence and robustness' across the three tasks is asserted without reporting metric values, baseline definitions, error bars, or significance tests. Because normalization reduces label cardinality, aggregate metric improvements could arise mechanically from merging labels rather than from better semantic representation. The evaluation must compare the normalized graph against the raw graph under identical splits, task prompts, and annotation conditions, and must report per-task metrics with variance. This is load-bearing for the paper's central claim.","section":"Abstract"},{"comment":"The supplied body of the manuscript is unreadable due to character-encoding corruption; every equation, figure, table, algorithm description, and experimental result is inaccessible. This prevents reviewing the clustering algorithm (thresholds, embedding model, hierarchical construction), the exact definitions of panel/event/story graphs, and the evaluation protocols. A readable manuscript is a prerequisite for any further assessment and must be provided before the claims can be checked.","section":"Full text (all sections)"},{"comment":"The normalization step assumes that annotation noise is mostly surface variance and that semantically equivalent actions are detectable through lexical and embedding similarity. This assumption is not supported by data on Manga109 label distributions. If lexically close labels denote genuinely different narrative actions, or if synonymous labels are expressed in lexically distant surface forms, merging will corrupt the graph and degrade reasoning tasks. The authors should provide label-distribution evidence, a manual precision/recall evaluation of merged clusters, and an ablation that quantifies harm from incorrect merges.","section":"Abstract (normalization premise)"},{"comment":"The abstract states that evaluation is performed on the same graph representation that is normalized, but does not indicate whether the evaluation tasks use independently annotated gold labels or the normalized labels themselves. If the latter, the reported improvements are circular, because normalization changes the label space that the metrics measure. The authors must clarify whether the evaluation labels are independent of the normalization process and, if so, describe how the independence was enforced.","section":"Abstract (evaluation design)"}],"minor_comments":[{"comment":"The text encoding must be fixed in the authors' resubmission; as rendered, all sections, captions, and references are illegible. Please verify that the PDF conversion is not corrupt before resubmitting.","section":"Full text"},{"comment":"The phrase 'symbolic transparency' is asserted as a property of the normalized graphs but is never operationally defined in the abstract. Please specify what measurement, if any, was used to verify that transparency is maintained (e.g., human interpretability of merged labels, or equivalence of graph paths).","section":"Abstract"},{"comment":"Please state the exact clustering method and the source of the threshold parameters (e.g., similarity threshold for merging labels). At present the abstract mentions 'lexical similarity and embedding-based clustering' with no indication of how sensitive the results are to the chosen thresholds.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript as supplied is not reviewable in its current form because the body is corrupted. I recommend asking the authors to resubmit a readable version with full experimental details and a matched comparison against the raw-graph baseline. I would not reject on the basis of the abstract alone, but the absence of quantitative support and the unaddressed circularity risk are significant enough that a revision is required rather than a minor edit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper up front: the abstract describes a reasonable idea, but as supplied the paper has no inspectable evidence. The full text is mojibake, and the abstract itself says the evaluations are \"preliminary\" without reporting a single number. So there is nothing to verify.\n\nWhat's genuinely new: the application of semantic normalization—consolidating labels via lexical similarity and embedding clustering—to hierarchical narrative graphs at panel, event, and story levels, using Manga109. That's a legitimate extension of existing graph-construction work, and the framing around cognitively grounded narrative models is sensible. The paper also is honest that the results are preliminary, which counts for something.\n\nThe soft spots are in the evidence, not the idea. The abstract's claim that normalization \"improves coherence and robustness\" is unsupported by any quantitative comparison or baseline. The stress-test note is right that merging labels can trivially improve aggregate metrics by reducing label cardinality; unless the evaluation is matched for granularity or shows task-level gains beyond that, the headline could be an artifact. There is also no indication whether the evaluation tasks use independent annotations or the normalized labels themselves, which leaves a circularity worry. The reader's concern about lexical/embedding similarity failing on genuinely ambiguous semantics is real but secondary—the primary issue is the missing numbers. The citation pattern can't be checked either, since the references are in the garbled portion.\n\nWho should read this: people working on multimodal story understanding or manga annotation might find the framework worth considering as a pre-processing step. But no one should rely on the empirical claims as they stand.\n\nRecommendation: as supplied, this should not go to peer review. The body is unreadable and the abstract lacks evidence. The right move is to ask the authors for a clean, complete version with actual evaluation numbers, baselines, and a description of how the normalized labels are used in the tasks. If that version exists, it might earn a light referee; right now there's nothing to referee.","headline":"A plausible but unverifiable incremental idea: the abstract is coherent, but the supplied text has no numbers and the body is garbled, so the empirical claims are empty as they stand.","tokens_in":8586,"tokens_out":4772,"would_cite":false,"duration_ms":46548,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Semantic normalization of hierarchical manga knowledge graphs reduces annotation noise and improves narrative reasoning.","keywords":["semantic normalization","hierarchical knowledge graphs","visual narratives","manga understanding","symbolic reasoning","action retrieval","character grounding","event summarization"],"falsifier":"Have human annotators judge which Manga109 action labels are semantically equivalent; if the clustering merges labels humans keep separate or splits labels humans equate on a substantial share of labels, and corrected graphs do not outperform the unnormalized baseline on retrieval, grounding, and summarization, the claim is falsified.","tokens_in":7959,"feed_emoji":"📖","tokens_out":3708,"duration_ms":42542,"temperature":0.7,"pith_summary":"The paper proposes that messy, redundant labels for actions and events in hierarchical manga knowledge graphs can be cleaned by semantic normalization, and that this cleaning is what lets symbolic reasoning over visual narratives work reliably. It groups semantically related labels using lexical similarity and embedding-based clustering across panel-, event-, and story-level graphs. Preliminary evaluations show gains in action retrieval, character grounding, and event summarization. The payoff is a graph that remains human-readable while becoming more consistent and robust. If correct, normalization is a key preprocessing step for scalable multimodal narrative understanding.","feed_headline":"Semantic label cleanup sharpens manga story reasoning","feed_subtitle":"Consolidating similar action labels across panel, event, and story levels improves retrieval, grounding, and summarization.","key_machinery":"The hierarchical narrative knowledge graph (panel-, event-, and story-level) combined with semantic normalization, which consolidates semantically related action and event labels using lexical similarity and embedding-based clustering. This lets noisy labels be merged without converting the representation into opaque vectors, keeping the graph symbolically transparent.","core_discovery":"The paper's central claim is that semantic normalization—consolidating actions and events that mean the same thing but are labeled differently—makes hierarchical symbolic graphs of visual narratives more coherent and robust. Applied to manga stories, where graphs at panel, event, and story levels carry noisy and redundant labels, normalization reduces annotation noise, aligns symbolic categories across narrative levels, and preserves interpretability. After normalization, action retrieval, character grounding, and event summarization improve, positioning normalization as a necessary step before scalable reasoning over multimodal narratives.","pith_inferences":["The same normalization idea could transfer to other domains with noisy symbolic annotations, such as storyboards, video event graphs, or illustrated instruction sequences, without retraining the graph structure.","The paper's evaluation is preliminary; a stronger test would compare normalized graphs against a human-constructed ontology and measure downstream gains on held-out stories.","Embedding-based clustering could be complemented by a human-in-the-loop review of merge decisions, since the paper does not specify thresholds or quality checks for merges.","Normalization errors, if they occur, likely concentrate in rare action labels, where embedding similarity is less reliable."],"forward_implications":["Action retrieval improves because equivalent action labels collapse into shared symbolic categories, so matching no longer fails on surface wording differences.","Character grounding becomes more reliable because redundant or conflicting action labels no longer fragment the evidence about what a character is doing.","Event summarization becomes more coherent because story-level categories align with panel- and event-level labels.","The method preserves symbolic transparency, so each merged label remains a human-readable symbol rather than a learned vector.","Semantic normalization can serve as a reusable preprocessing step before any symbolic reasoner operates over multimodal narrative graphs."],"supporting_citations":[],"fun_headline_variants":["Label cleanup boosts manga narrative reasoning","Normalizing labels improves comic story understanding","Semantic normalization sharpens visual narrative graphs","Cleaning action labels enhances manga reasoning tasks","Unified labels strengthen symbolic manga narratives"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The method assumes that semantically equivalent actions and events can be recognized from the lexical and embedding similarity of their labels, so the noise it removes is surface-level wording variance rather than genuine category ambiguity.","fun_headline_variants_meta":{"raw":{"variants":["Label cleanup boosts manga narrative reasoning","Normalizing labels improves comic story understanding","Semantic normalization sharpens visual narrative graphs","Cleaning action labels enhances manga reasoning tasks","Unified labels strengthen symbolic manga narratives"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000126,"raw_usage":{"total_tokens":911,"prompt_tokens":670,"completion_tokens":241,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":414,"completion_tokens_details":{"reasoning_tokens":179}},"tokens_in":414,"tokens_out":241,"duration_ms":3075,"temperature":1.0,"reasoning_tokens":179,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:33:34.356238+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human annotators judge which Manga109 action labels are semantically equivalent; if the clustering merges labels humans keep separate or splits labels humans equate on a substantial share of labels, and corrected graphs do not outperform the unnormalized baseline on retrieval, grounding, and summarization, the claim is falsified.","supporting_citations":[],"review_version":1}