{"id":"eadea6e2-a251-4966-a647-13f707593c08","arxiv_id":"2607.05841","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"VisTCP pairs a TCP-trained structured-representation model with joint-embedding uncertainty views so historians can iteratively build knowledge-graph annotations of Traditional Chinese Paintings.","lead":"VisTCP is a human-in-the-loop visualization system that helps art historians build knowledge-graph representations of Traditional Chinese Paintings by pairing a domain-trained model with expert correction. Generalists may care because it targets culturally specific images that ordinary vision models misread, a recurring digital-heritage problem.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Trustworthiness of VisTCP representations rests on a three-expert taxonomy plus qualitative case/usage/interview evidence, without quantitative accuracy, IAA, or baseline tests of generalization.","rationale":"The reader’s weakest_assumption correctly isolates the load-bearing soft spot: generalization and evidentiary support for “trustworthy.” The method itself (taxonomy → TCP-oriented structured model → joint-embedding uncertainty view → iterative refine) is a coherent HITL design with no internal circularity or equation-level flaw. The concern is purely that the evaluation package described—case study, usage scenario, and expert interviews—does not underwrite the strength of the central claim when the ontology and feedback both originate from three experts and no quantitative anchors are supplied. Because the abstract (and the evaluation design it summarizes) leaves those anchors unreported, the appropriate stance remains UNVERDICTED pending inspection of the full results. No stronger technical inconsistency was identified; the issue is support for the claim as stated. Verdict and concern therefore stay aligned with the reader.","tokens_in":2138,"tokens_out":551,"duration_ms":32691,"concrete_test":"Extract from the full evaluation/results sections: (1) number of paintings and annotated objects/relations for train/test; (2) any reported model metrics or HITL ablation; (3) inter-expert agreement on taxonomy or labels; (4) whether interview participants are independent of the three taxonomy experts. If quantitative held-out metrics and independent IAA are absent or weak, the trustworthiness claim remains unsubstantiated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that a model trained on annotations from a three-expert pilot taxonomy, combined with joint-embedding visualization of expert–model differences and iterative expert refinement, yields structured KG representations that are trustworthy for art-history use on a broader TCP corpus. The paper’s evaluation design (case study, usage scenario, expert interviews on a real dataset) is the sole support offered for that claim. No quantitative extraction metrics (object/relation precision–recall or F1 on held-out paintings), inter-annotator agreement on the taxonomy or labels, dataset scale, or controlled comparison (model-only vs. HITL vs. pure expert) appear in the abstract’s description of the work. With the same small expert pool defining the ontology and supplying the refinement/evaluation feedback, the risk is that the loop mainly reconfirms a narrow consensus rather than demonstrating external correctness or stability. The joint-embedding view usefully surfaces disagreements but does not itself validate that refined outputs are accurate outside the study participants.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes VisTCP, a human-in-the-loop visualization framework for constructing knowledge-graph structured representations of Traditional Chinese Paintings (TCPs). Motivated by the claim that generic image structured-representation methods fail on TCPs (domain shift from natural images; hard even for experts), the authors first run a pilot study with three domain experts to build a TCP-oriented semantic taxonomy, then train a structured-representation model on expert annotations. A joint-embedding visualization view surfaces differences between expert labels and model predictions so that experts can iteratively refine outputs and, in principle, improve the model. Effectiveness is argued via a case study, a usage scenario, and expert interviews on a real TCP dataset, with the central claim that VisTCP yields trustworthy KG-style representations usable for archaeology and art-history research.","tokens_in":2306,"tokens_out":1246,"duration_ms":36836,"significance":"If the central claim holds, VisTCP would be a useful contribution at the intersection of visualization, digital humanities, and cultural-heritage computing: a domain-specific taxonomy plus a disagreement-aware HITL loop for a corpus that resists off-the-shelf scene-graph methods. The joint-embedding view of expert-versus-model differences is a concrete design idea that could transfer to other expert-scarce annotation settings. The work is primarily a systems/design contribution rather than a new learning algorithm; its significance therefore rests on whether the evaluation actually establishes trustworthiness and generalizability of the resulting representations beyond the study participants.","major_comments":[{"comment":"The abstract and evaluation design (case study, usage scenario, expert interviews) do not report quantitative extraction quality for the TCP-oriented structured-representation model—e.g., object/relation precision, recall, or F1 on held-out paintings, nor any measure of representation completeness against expert gold. The central claim of “trustworthy” KG representations is load-bearing and is not established by qualitative evidence alone; without held-out accuracy numbers it is unclear whether the model-plus-HITL loop improves correctness or mainly re-expresses the same experts’ judgments.","section":null},{"comment":"Taxonomy construction, training labels, iterative refinement, and interview-based effectiveness evidence all appear to draw on a three-expert pool (pilot study). This creates a methodological circularity risk for the trustworthiness claim: the same narrow consensus that defines the ontology and labels also supplies the evidence that the refined outputs are good. The manuscript needs either (a) inter-annotator agreement on taxonomy and labels from a larger or held-out expert set, or (b) external validation (additional experts or art-historical ground truth not involved in taxonomy design) before the representations can be treated as stable across a broader TCP corpus.","section":null},{"comment":"No controlled baseline or ablation is described comparing, for example, generic scene-graph / open-vocabulary methods, model-only extraction, pure expert annotation, and the full VisTCP HITL loop on the same paintings. Without such comparisons it is hard to attribute gains to the TCP-oriented taxonomy, the joint-embedding view, or iterative refinement, and hard to quantify how much the framework improves over existing structured-representation pipelines that the introduction claims “perform poorly on TCPs.”","section":null},{"comment":"Dataset scale and split protocol are not stated in the evaluation summary (number of paintings, objects/relations annotated, train/val/test or leave-painting-out design). For a claim that the trained model plus joint-embedding refinement generalizes as trustworthy structured representation, the manuscript must report corpus size, annotation volume, and how generalization beyond the annotated set was tested; otherwise the case study and usage scenario remain anecdotal relative to the stated research goal.","section":null}],"minor_comments":[{"comment":"Clarify early what “structured representation” concretely means in the KG (node/edge types, event vs. object relations, multi-instance handling) so that later claims about semantic understanding are checkable against a fixed schema.","section":null},{"comment":"The phrase “trustworthy structured representations” is strong; either define operational criteria (e.g., expert-verified precision thresholds, stability under re-annotation) or soften the claim language to match the qualitative evidence actually provided.","section":null},{"comment":"Describe the joint-embedding view’s construction more precisely (what is embedded—objects, relations, full graphs; which similarity space; how uncertainty is encoded) so that other systems can reproduce or compare the design.","section":null},{"comment":"State whether iterative expert refinements are used only for one-off correction of the current painting or are fed back to retrain/fine-tune the model, and under what schedule; the abstract implies both “refine the structured representation” and “iterative optimization of the model.”","section":null},{"comment":"Position related work more sharply against prior cultural-heritage KG / painting annotation tools and against recent open-vocabulary detection and scene-graph models, so the TCP-specific gap is measurable rather than asserted.","section":null}],"recommendation":"major_revision","confidential_remarks":"Evaluation is the main gate. For a cs.HC / VIS venue this qualitative package is not automatically disqualifying, but the paper’s own framing (“trustworthy,” archaeology/art-history use, model trained on expert data) invites ML-style scrutiny that the current evidence does not meet. I would accept a revision that adds held-out extraction metrics, IAA or external expert validation, dataset statistics, and at least one baseline comparison—even if the design contribution remains primary. If the authors cannot expand beyond the three-expert loop, the claim language and contribution framing should be narrowed substantially. Fit for a serious journal in the field is conditional on that revision; novelty of the joint-embedding HITL idea is real but secondary to the evaluation gap."},"author_rebuttal":{"model":"grok-4.5","summary":"We thank the referee for a careful and constructive review. The four major comments correctly identify that our central claim of trustworthy TCP structured representations currently rests too heavily on qualitative evidence (case study, usage scenario, expert interviews) and on a small expert pool. We agree that the manuscript must report quantitative extraction quality, clarify dataset scale and splits, reduce circularity risk around the three-expert consensus, and provide controlled comparisons against generic structured-representation pipelines. We will revise the evaluation section accordingly and temper claims where evidence remains qualitative. Below we respond point by point and indicate the concrete revisions planned for the next version.","responses":[{"response":"We agree that the trustworthiness claim is load-bearing and that qualitative evidence alone is insufficient. The current manuscript emphasizes the HITL workflow and expert feedback; it does not report held-out object/relation precision, recall, or F1, nor a completeness measure against expert gold. In revision we will (1) add a quantitative evaluation of the TCP-oriented structured-representation model on a held-out painting set (object detection and relation extraction P/R/F1, plus a simple completeness metric relative to expert gold graphs); (2) report before/after metrics for representations refined in the joint-embedding HITL loop so that gains from iterative refinement are separated from model-only output; and (3) revise the abstract and claims so that “trustworthy” is scoped to what the numbers and expert validation jointly support, rather than implied by qualitative evidence alone. Where sample size limits statistical strength, we will state that limitation explicitly.","revision_made":"yes","referee_comment":"The abstract and evaluation design (case study, usage scenario, expert interviews) do not report quantitative extraction quality for the TCP-oriented structured-representation model—e.g., object/relation precision, recall, or F1 on held-out paintings, nor any measure of representation completeness against expert gold. The central claim of “trustworthy” KG representations is load-bearing and is not established by qualitative evidence alone; without held-out accuracy numbers it is unclear whether the model-plus-HITL loop improves correctness or mainly re-expresses the same experts’ judgments."},{"response":"The referee is right that relying on the same three-expert pool for taxonomy design, labeling, refinement, and interview-based effectiveness creates a circularity risk. We will address this in two ways. First, we will report inter-annotator agreement (e.g., pairwise agreement / Cohen’s or Fleiss’ kappa where applicable) on taxonomy categories and on object/relation labels among the participating experts, and we will document how disagreements were resolved. Second, we will add external validation: at least one additional domain expert (or a small held-out expert set) who did not design the taxonomy will review a sample of model and HITL-refined graphs, and we will report their agreement with the refined representations and any systematic disagreements. We will also expand the Limitations section to state that the taxonomy reflects a small expert consensus and that broader multi-institution validation remains future work. We cannot fully eliminate the small-pool constraint within this revision cycle, but the IAA numbers plus held-out expert review will make the stability claim more honest and testable.","revision_made":"yes","referee_comment":"Taxonomy construction, training labels, iterative refinement, and interview-based effectiveness evidence all appear to draw on a three-expert pool (pilot study). This creates a methodological circularity risk for the trustworthiness claim: the same narrow consensus that defines the ontology and labels also supplies the evidence that the refined outputs are good. The manuscript needs either (a) inter-annotator agreement on taxonomy and labels from a larger or held-out expert set, or (b) external validation (additional experts or art-historical ground truth not involved in taxonomy design) before the representations can be treated as stable across a broader TCP corpus."},{"response":"We agree that the introduction’s claim that generic structured-representation methods perform poorly on TCPs is not backed by controlled comparisons in the current evaluation, and that attribution of gains to the taxonomy, joint-embedding view, and iterative refinement is therefore weak. In revision we will add a controlled comparison on a fixed set of paintings covering: (i) at least one generic scene-graph or open-vocabulary structured-representation baseline (as used or cited in the related work); (ii) our TCP-oriented model alone (no HITL); (iii) pure expert annotation time/quality where feasible; and (iv) the full VisTCP HITL loop. We will report the same object/relation quality metrics (and, where appropriate, expert preference or correction effort) across these conditions so that improvements can be attributed more clearly. If pure expert annotation of the full set is too costly, we will use a stratified subset and report effort (time, number of corrections) for model-only vs. HITL. Claims in the introduction and discussion will be rewritten to match what these comparisons actually show.","revision_made":"yes","referee_comment":"No controlled baseline or ablation is described comparing, for example, generic scene-graph / open-vocabulary methods, model-only extraction, pure expert annotation, and the full VisTCP HITL loop on the same paintings. Without such comparisons it is hard to attribute gains to the TCP-oriented taxonomy, the joint-embedding view, or iterative refinement, and hard to quantify how much the framework improves over existing structured-representation pipelines that the introduction claims “perform poorly on TCPs.”"},{"response":"This is a clear omission and we will fix it. The revised manuscript will report: total number of TCP images in the corpus; number of paintings annotated for training and evaluation; counts of annotated objects, relations, and taxonomy categories; and the train/validation/test (or leave-painting-out) split protocol used for the structured-representation model. We will also state how held-out paintings used in the quantitative evaluation and in the case study / usage scenario relate to the training set, so that generalization beyond the annotated set is explicit rather than implied. If the corpus is modest, we will say so and frame the contribution as a systems/design study with preliminary generalization evidence rather than as large-scale corpus-wide validation.","revision_made":"yes","referee_comment":"Dataset scale and split protocol are not stated in the evaluation summary (number of paintings, objects/relations annotated, train/val/test or leave-painting-out design). For a claim that the trained model plus joint-embedding refinement generalizes as trustworthy structured representation, the manuscript must report corpus size, annotation volume, and how generalization beyond the annotated set was tested; otherwise the case study and usage scenario remain anecdotal relative to the stated research goal."}],"tokens_in":2001,"tokens_out":1455,"duration_ms":26545,"standing_objections":[]},"desk_editor":{"model":"grok-4.5","letter":"Punchline: VisTCP is a competent cultural-heritage HCI system—TCP taxonomy, a domain-tuned structured model, and a joint-embedding view that surfaces expert–model disagreement for iterative fix-up. That package is useful for art historians who need searchable semantic graphs of paintings. It is not a general CV advance, and the central “trustworthy” claim rests on thin evidence.\n\nWhat is actually new and done well: modern scene-graph and open-vocab methods fail on Traditional Chinese Paintings because objects, styles, and events diverge hard from natural images. The authors run a three-expert pilot to build a TCP semantic taxonomy, train a structured extractor on those labels, then expose residual uncertainty in a joint embedding of expert annotations versus model predictions so historians can correct and retrain. The joint-embedding view is the cleanest piece: it makes disagreement visible instead of hiding it behind a single score. Case study, usage scenario, and expert interviews on a real corpus are the right first-order evaluation for a systems paper of this type.\n\nSoft spots, in proportion: the stress-test lands. The abstract (and the evaluation design it describes) gives no object/relation precision–recall, no inter-annotator agreement on the taxonomy or labels, no dataset scale, and no controlled comparison of model-only vs HITL vs pure expert. With the same small expert pool defining the ontology and supplying the interview feedback, you mainly learn that the loop is usable inside that consensus, not that the graphs generalize as trustworthy across a broader TCP corpus. That is a real but fixable gap for a VIS/HCI paper, not a load-bearing mathematical flaw. Mild design–evaluate overlap risk; no equation circularity.\n\nWho it is for: people building digital-art-history tools, cultural-heritage search, or domain HITL visualization. A serious referee at IEEE VIS, CHI, or a digital-humanities venue should see it. I would send it to review and expect requests for quantitative extraction metrics, IAA, dataset details, and clearer separation of taxonomy builders from evaluators. Worth engaging; not a desk reject.","headline":"Solid domain HITL system for TCP knowledge graphs; the joint-embedding refinement loop is the real contribution, but “trustworthy” is overclaimed on qualitative evidence alone.","tokens_in":3012,"tokens_out":528,"would_cite":false,"duration_ms":25693,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"VisTCP lets art historians build trustworthy knowledge-graph representations of Traditional Chinese Paintings through human-in-the-loop model refinement.","keywords":["Traditional Chinese Painting","knowledge graph","structured representation","visualization framework","human-in-the-loop","semantic taxonomy","joint embedding","art history"],"falsifier":"On a held-out set of Traditional Chinese Paintings, measure whether VisTCP graphs improve object-and-relation coverage, inter-expert agreement, or research utility relative to unaided expert annotation and to standard vision models; a null or negative result would falsify the claim.","tokens_in":2987,"feed_emoji":"🎨","tokens_out":793,"duration_ms":22332,"temperature":0.7,"pith_summary":"Standard image understanding tools fail on Traditional Chinese Paintings because the objects, scenes, and events look nothing like modern photographs, and even specialists often cannot name every ancient motif with certainty. VisTCP is a visualization framework that closes that gap by first building a TCP-specific semantic taxonomy with domain experts, training a structured-representation model on their annotations, and then showing experts where the model and their own labels diverge in a joint embedding view. Experts correct the graph; the corrections retrain the model. The loop is meant to produce knowledge-graph style descriptions of paintings that art historians can trust for archaeology and art-history work. The paper demonstrates the pipeline through a case study, a usage scenario, and expert interviews on a real painting corpus.","feed_headline":"VisTCP builds knowledge graphs of Chinese paintings with experts in the loop","feed_subtitle":"A visualization loop trains on expert labels, surfaces model uncertainty, and iterates toward trustworthy TCP semantics.","key_machinery":"The joint-embedding visualization view that places expert annotations and model predictions in a shared space so users can see uncertainty, correct the structured representation, and iteratively improve the TCP-oriented model.","core_discovery":"A human-in-the-loop visualization framework can produce trustworthy knowledge-graph structured representations of Traditional Chinese Paintings by combining a TCP-oriented extraction model trained on expert labels, a joint-embedding view that surfaces expert-versus-model differences, and iterative expert refinement of the resulting graph.","pith_inferences":["The same taxonomy-plus-joint-embedding loop could transfer to other heritage image domains whose visual vocabulary diverges from natural-image benchmarks.","If the taxonomy is published, it becomes a community resource for labeling larger TCP datasets even outside this tool.","Without reported quantitative accuracy or baseline comparisons, adoption will hinge more on expert qualitative trust than on measured gains.","The joint-embedding view itself is a reusable pattern for any human-in-the-loop structured-representation task where model and expert disagree on rare classes."],"forward_implications":["Art historians can obtain consistent, machine-readable knowledge graphs of TCP objects and relationships instead of relying only on free-text notes.","Archaeology and art-history studies gain a reusable semantic layer for comparing motifs, events, and compositions across paintings.","Model uncertainty becomes visible rather than hidden, so experts know where their domain knowledge must intervene.","Each round of expert correction can retrain the extractor, gradually reducing the effort needed for new paintings."],"fun_headline_variants":["VisTCP: Expert-AI loop builds knowledge graphs of Chinese paintings","VisTCP surfaces model-expert gaps for TCP knowledge graphs","Human-in-the-loop VisTCP extracts structured TCP semantics","Experts refine VisTCP models into trustworthy painting graphs","VisTCP joint embeddings guide TCP knowledge-graph construction"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"A taxonomy and training labels produced with three domain experts are complete and stable enough that a model trained on them, plus iterative human correction shown in joint embeddings, will yield representations that generalize as trustworthy across a wider corpus of Traditional Chinese Paintings.","fun_headline_variants_meta":{"raw":{"variants":["VisTCP: Expert-AI loop builds knowledge graphs of Chinese paintings","VisTCP surfaces model-expert gaps for TCP knowledge graphs","Human-in-the-loop VisTCP extracts structured TCP semantics","Experts refine VisTCP models into trustworthy painting graphs","VisTCP joint embeddings guide TCP knowledge-graph construction"]},"model":"grok-4.5","cost_usd":0.008392,"raw_usage":{"total_tokens":2000,"prompt_tokens":806,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":83920000,"prompt_tokens_details":{"text_tokens":806,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1128,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":806,"tokens_out":66,"duration_ms":14825,"temperature":1.0,"reasoning_tokens":1128,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T19:40:34.512993+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a held-out set of Traditional Chinese Paintings, measure whether VisTCP graphs improve object-and-relation coverage, inter-expert agreement, or research utility relative to unaided expert annotation and to standard vision models; a null or negative result would falsify the claim.","supporting_citations":[],"review_version":1}