{"id":"84c80aa2-645c-4d7d-b877-1d4d3ab2f619","arxiv_id":"2606.09855","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Object presence alone fails to predict minhwa genres while image-text fusion succeeds, revealing faithful but insufficient object grounding where genre hinges on symbol arrangement rather than inventory.","lead":"The paper finds that listing symbolic objects in Korean folk paintings predicts genre poorly, while fusing image and curatorial text works better, and object-grounded representations hurt accuracy; the visual evidence is localized but insufficient because genre depends on arrangement. A smart generalist might read it to understand limits of object detection in symbolic cultural datasets and the value of multimodal fusion for heritage AI.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Expert object crops and eight-field captions may incompletely capture symbolic content, so the performance gap cannot be cleanly attributed to arrangement rather than missing symbols or annotation bias.","rationale":"The reader's weakest_assumption matches the load-bearing point exactly. Full text does not remove the need for this check; the argument's validity turns on whether the symbol lists are verifiably complete. The proposed test is a direct, low-cost validation that would either confirm or undermine the dissociation claim without requiring new models.","tokens_in":1814,"tokens_out":324,"duration_ms":22360,"concrete_test":"For the full corpus, extract all distinct symbols mentioned across the eight caption fields; compute recall of these against the expert crop annotations per painting. If mean recall < 0.85, augment the symbol-list model with the missing caption symbols and re-evaluate genre accuracy gap versus the image+text model; a >10-point reduction in the gap falsifies clean attribution to arrangement.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that the symbol inventory (derived from expert crops) is exhaustive and unbiased relative to the curatorial text and actual painting content. If crops systematically omit symbols referenced in captions or present in images, the symbol-list baseline is handicapped by incomplete input rather than by the claimed insufficiency of inventory alone. The image+text fusion model and the object-grounded penalty would then reflect data quality differences, not the arrangement-vs-inventory distinction. The leakage-safe evidence map addresses spatial faithfulness but does not test completeness of the symbol set itself.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents MinhwaNet for Korean folk painting (minhwa) analysis. It claims that genre prediction from a symbol inventory (derived from expert object crops) performs substantially worse than image+curatorial-text fusion, that forcing object-grounded representations hurts accuracy, and that this constitutes a 'faithful-but-insufficient dissociation': the part-level detector yields spatially faithful, leakage-safe evidence maps aligned with curator-isolated objects and gradient saliency, yet genre depends on symbol arrangement rather than presence. The work further shows that genre (content) transfers across held-out institutions while era (style) does not, with analogous behavior on two additional labels, and releases the multimodal system, a worked-example evidence-map reading, and heritage-collection evaluation cautions.","tokens_in":1937,"tokens_out":507,"duration_ms":31510,"significance":"If the dissociation result holds, the paper offers a concrete demonstration that object-centric grounding can be faithful yet insufficient for symbolic art domains where arrangement carries the semantic load. This has implications for explainable multimodal models in digital humanities. Explicit strengths include the use of separate expert crops, held-out-institution transfer tests, and the public release of code, worked examples, and recurring evaluation cautions for long-tailed heritage data.","major_comments":[{"comment":"Abstract and Methods: The central claim that the performance gap between symbol-list and image+text models is attributable to arrangement (rather than inventory completeness) rests on the assumption that expert object crops plus eight-field captions exhaustively capture all symbolically relevant content. No completeness verification (e.g., symbols mentioned in captions but absent from crops, or visible in full images but omitted) is described; if systematic omissions exist, the gap could reflect input quality differences rather than the claimed faithful-but-insufficient dissociation.","section":"Abstract/Methods"},{"comment":"Experiments section: The abstract reports clear performance differences and transfer results, yet provides no details on data splits, statistical significance tests, exact model architectures, or hyper-parameters. Without these, the reported superiority of fusion over inventory-only and the genre-vs-era transfer contrast cannot be independently verified.","section":"Experiments"}],"minor_comments":[{"comment":"The term 'faithful-but-insufficient dissociation' is introduced without a formal definition or comparison to related concepts in explainable AI or multimodal grounding literature.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. The comments highlight important points on assumption verification and experimental transparency. We respond to each major comment below and will incorporate revisions accordingly.","responses":[{"response":"We agree this is a valid concern: without explicit completeness checks, the observed gap between inventory-only and fusion models could partly reflect annotation differences rather than arrangement alone. The expert crops target curator-identified symbolic objects and captions provide detailed descriptions, but we performed no systematic audit for omissions. In revision we will add a Methods subsection discussing this assumption, its potential impact on the dissociation claim, and a qualitative sample-based check comparing captions, crops, and full images for discrepancies. This will qualify the interpretation without altering the core empirical results.","revision_made":"yes","referee_comment":"[Abstract/Methods] Abstract and Methods: The central claim that the performance gap between symbol-list and image+text models is attributable to arrangement (rather than inventory completeness) rests on the assumption that expert object crops plus eight-field captions exhaustively capture all symbolically relevant content. No completeness verification (e.g., symbols mentioned in captions but absent from crops, or visible in full images but omitted) is described; if systematic omissions exist, the gap could reflect input quality differences rather than the claimed faithful-but-insufficient dissociation."},{"response":"We acknowledge the need for full reproducibility details. The current manuscript version omits explicit reporting of these elements in the Experiments section. In the revision we will expand that section to include: institution-based data splits for the transfer experiments, statistical significance testing procedures and results, precise model architectures (including fusion components), and all hyper-parameters with training protocols. These will be presented in dedicated paragraphs and a summary table.","revision_made":"yes","referee_comment":"[Experiments] Experiments section: The abstract reports clear performance differences and transfer results, yet provides no details on data splits, statistical significance tests, exact model architectures, or hyper-parameters. Without these, the reported superiority of fusion over inventory-only and the genre-vs-era transfer contrast cannot be independently verified."}],"tokens_in":1519,"tokens_out":449,"duration_ms":28619,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that listing the symbols in these Korean folk paintings predicts genre far worse than fusing the image with the curatorial text, and pushing the model to stay object-grounded actually lowers accuracy. The predictions remain spatially faithful via a leakage-safe map from the part detector, which lines up with both the crops and gradient saliency.\n\nThe paper takes standard multimodal and explanation techniques and runs them on a public minhwa corpus with eight-field captions and separate expert crops. This yields the dissociation they name: the part-level view is honest about its inputs, yet genre turns on arrangement rather than inventory. They also show content labels like genre transfer across held-out institutions while style labels like era do not, and they repeat the check on two other labels. Releasing the system, a worked example, and some cautions for long-tailed heritage data is concrete and usable.\n\nThe soft spot is the completeness of the symbol inventory from the crops. The central claim attributes the gap to arrangement, but if the crops miss symbols that the captions or images contain, the list baseline is handicapped by missing input rather than by the claimed dissociation. The paper keeps institutions separate and uses the map for location, yet without reported coverage numbers or agreement checks on the symbol set, that assumption stays untested. The rest of the evidence looks solid enough.\n\nThis is for people working on computer vision for cultural heritage or on multimodal models with symbolic data. A reader who needs practical examples of faithful explanations on art collections would get value from it. It deserves a serious referee because the empirical result is clear and the methods are grounded, even if one data assumption needs more scrutiny in review.\n\nI would send it to peer review.","headline":"Object lists underperform image-text fusion for minhwa genre prediction because arrangement matters, with a faithful evidence map, but the crops may not fully capture the symbol set.","tokens_in":2404,"tokens_out":422,"would_cite":false,"duration_ms":35394,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"In Korean folk paintings, a list of which auspicious symbols appear predicts genre far worse than image-plus-text fusion, and forcing object grounding hurts accuracy.","keywords":["Korean folk painting","minhwa","object grounding","genre classification","multimodal fusion","cultural heritage","symbolic objects","evidence map"],"falsifier":"Retraining the symbol-list model after adding explicit spatial-relation features between detected objects and observing genre accuracy rise to match the image-text fusion model would falsify the insufficiency claim.","tokens_in":2706,"feed_emoji":"🖼️","tokens_out":715,"duration_ms":31772,"temperature":0.7,"pith_summary":"The paper establishes that genre classification in minhwa cannot be reduced to detecting the presence of a fixed set of symbolic objects such as tigers or peonies. Models relying solely on symbol inventories underperform those that combine the full painting with bilingual curatorial captions, and explicitly requiring the representation to be grounded in detected objects further lowers accuracy. Nevertheless the decisions remain spatially localizable through a leakage-safe evidence map that aligns with expert object crops and gradient saliency. A sympathetic reader would care because the result separates honest part-level explanation from the compositional information actually needed for the target label. The same separation holds for content versus style attributes across institutions.","feed_headline":"Symbol lists fail to predict minhwa genres","feed_subtitle":"Genre accuracy is higher with image-text fusion than with object inventories alone because arrangement, not presence, drives the label.","key_machinery":"Leakage-safe object evidence map projected from a part-level detector, which localizes the genre decision while remaining faithful to expert crops and model saliency.","core_discovery":"A model given only a list of which symbols a painting contains predicts the genre far worse than a model that fuses the image with the curatorial text, and forcing the genre representation to be object-grounded actively hurts accuracy. The visual evidence on which the genre prediction rests is nonetheless localized and inspectable via a leakage-safe object evidence map projected from a part-level detector that is spatially faithful to where curators isolated symbolic objects and to patch-based gradient saliency. This configuration is termed a faithful-but-insufficient dissociation: the part-level model is honest about what it sees, yet the genre target depends on how symbols are arranged rat","pith_inferences":["The same faithful-but-insufficient pattern is likely to appear in other symbolic art traditions that rely on recurring motifs whose meaning depends on placement.","Adding explicit composition encoders (spatial graphs or relation modules) between detected symbols could close the performance gap without losing inspectability.","Heritage datasets with long-tailed labels will repeatedly require separate checks for whether performance gaps trace to missing arrangement cues or to incomplete expert annotations."],"forward_implications":["Genre labels transfer successfully to held-out source institutions while style labels such as era do not.","Object-grounded representations actively reduce genre accuracy compared with ungrounded image-text fusion.","Symbol presence alone is insufficient; genre prediction requires modeling arrangement of the same symbols.","Content-based labels generalize across collections while style-based labels remain collection-specific."],"fun_headline_variants":["Symbol lists fail minhwa genre prediction","Minhwa genres require symbol arrangement not lists","Image fusion surpasses object-only minhwa models","Object evidence faithful yet insufficient for genre"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The expert object crops and eight-field bilingual curatorial captions accurately and exhaustively capture the symbolic content without annotation leakage or systematic bias.","fun_headline_variants_meta":{"raw":{"variants":["Symbol lists fail minhwa genre prediction","Minhwa genres require symbol arrangement not lists","Image fusion surpasses object-only minhwa models","Object evidence faithful yet insufficient for genre"]},"model":"grok-4.3","cost_usd":0.005384,"raw_usage":{"total_tokens":2654,"prompt_tokens":786,"num_sources_used":0,"completion_tokens":52,"cost_in_usd_ticks":53837000,"prompt_tokens_details":{"text_tokens":786,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1816,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":786,"tokens_out":52,"duration_ms":20406,"temperature":1.0,"reasoning_tokens":1816,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T14:52:19.072444+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Retraining the symbol-list model after adding explicit spatial-relation features between detected objects and observing genre accuracy rise to match the image-text fusion model would falsify the insufficiency claim.","supporting_citations":[],"review_version":1}