{"id":"c4a5939d-5b2b-420b-a52d-80f6cdc2905e","arxiv_id":"2607.28509","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Mixed-data SFT plus Hierarchical Coverage-Discounted GRPO yields open-source SOTA multi-reference image-grounded video captions on MRVBench without hurting general captioning.","lead":"RefCaptioner teaches video models to write captions that attach multiple reference-image tags to the right local phrases, while rejecting distractors. The method and MRVBench matter for multi-reference video generation and editing pipelines that need source-faithful, phrase-level grounding.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Judge–baseline coupling is real but partly mitigated; the more load-bearing gap is that human preference/binding evidence is too small and reconstruction confounds caption quality with reference selection.","rationale":"I agree with the reader’s CONDITIONAL verdict and with judge dependence as a material caveat: Gemini judging a leaderboard that includes Gemini is a real bias risk for Table 1 margins, and training rewards also use LLM judges (B.3). I only partially agree that this is the single most load-bearing soft spot for the stated strongest claim. The paper already provides partial insulation—open-source SOTA vs many non-Gemini models, CaptionRefine controls, VDC/VCapsBench gains, reward ablations, and some human numbers—so judge coupling alone does not collapse the contribution. The claim as written bundles automatic SOTA with human preference and reconstruction fidelity. That bundle’s weakest load-bearing piece is the small-N, expert-only human study plus reconstruction that confounds text quality with which references are passed to the generator. Strengthening or failing that test decides whether CONDITIONAL should stay (method + benchmark still useful) or harden toward stronger skepticism on the human/downstream rhetoric. No change off CONDITIONAL: novelty and engineering are real; unreproducible train data and judge-heavy eval remain appropriate conditions, with human/downstream scale as the sharper check.","tokens_in":23585,"tokens_out":709,"duration_ms":19918,"concrete_test":"Re-run the §6.7 / Fig. 6 reconstruction protocol on ≥100 held-out MRVBench videos with a fixed gold positive-reference set for every captioner (strip model-selected refs; feed identical refs + each model’s tag-free caption to Wan/Seedance). Separately collect blinded preference ranks from ≥10 annotators on the same 100 captions. If RefCaptioner’s GSB Good rate vs Gemini/GPT falls to ~chance or mean rank is no longer best, the human/downstream half of the strongest claim does not hold.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The reader correctly flags Gemini-3.1-Pro as both a top proprietary baseline (Table 1) and the primary automatic judge for KP-Cov, VQA, Ref-Bind, and Subj metrics (§5.1, A.4–A.5). That coupling can inflate automatic margins. However, the paper’s central claim does not rest only on those scores: it also cites human preference and caption-conditioned reconstruction (§6.7, Fig. 6, F.1–F.2). Those supports are thinner than they appear. Human ranking/binding uses only 40 videos and three graduate annotators (F.1); reconstruction GSB uses the same small expert pool and conditions the generator on both the caption text and the subset of references the captioner chose (E.1), so better reconstructions can come from better reference selection alone rather than better phrase-level grounded captions. Ablations (Table 5) and CaptionRefine (Table 2) support the method, but the headline “preferred by annotators and more source-faithful reconstruction” overclaims relative to the scale and confounds of the human evidence. If human preference does not replicate at larger N, or if reconstruction wins vanish when all methods receive the same gold reference set, the claim’s human/downstream pillar weakens even if automatic MRVScore remains high.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces multi-reference image-grounded video captioning: given a video and candidate reference images (including distractors), a model must produce a factual caption with phrase-level <Image_i> tags, correct selection, local binding, distractor rejection, and same-entity grouping. RefCaptioner post-trains Qwen3-VL-8B via mixed multi-reference/general-caption SFT, then Hierarchical Coverage-Discounted GRPO with a dual-branch reward (factual coverage discounted by VQA errors; grounding coverage discounted by binding errors, DAES, and CRSC). The authors build a 20k-video training corpus and MRVBench (462 held-out real/AIGC videos, 3,831 references, 2,172 QA pairs). On MRVBench, RefCaptioner leads open-source models (MRVScore 0.882) and is competitive with proprietary systems; it also improves or holds on VDC and VCapsBench. Ablations, a CaptionRefine control, reference-count robustness, human preference (n=40), and caption-conditioned reconstruction GSB are reported as supporting evidence.","tokens_in":23963,"tokens_out":1630,"duration_ms":34611,"significance":"If the results hold under stronger human validation, this is a clear and timely contribution: it defines a practically motivated task at the intersection of video captioning and multi-reference generation/editing, ships a dedicated benchmark with explicit distractors and multi-view entity groups, and shows that joint post-training (not post-hoc tag insertion) is needed for phrase-level grounding. Strengths include a held-out test set, broad open-source and proprietary baselines, CaptionRefine and reward ablations (Tables 2 and 5), retention of general captioning ability (Tables 3–4), and public release of MRVBench evaluation data. The work is empirical systems research rather than theory; its value is the task formulation, training recipe, and evidence that multi-reference grounding can be added without collapsing standard caption quality.","major_comments":[{"comment":"§5.1, Appendix A.4–A.5, and Table 1: Gemini-3.1-Pro is both a top proprietary baseline and the primary automatic judge for KP-Cov, VQA/VQA-Cov, Ref-Bind, and Subj-R/F1. This coupling can systematically favor styles the judge prefers and inflate margins relative to Gemini (and possibly other models). The paper should (i) quantify judge–human agreement on a stratified subset for binding and subject consistency, (ii) report at least one alternate-judge or cross-model re-score for the main metrics, and (iii) state clearly which claims rest only on automatic scores versus human checks. Without this, the SOTA comparison to Gemini in Table 1 is not fully trustworthy even if open-source rankings remain informative.","section":"§5.1, Table 1, Appendix A.4–A.5"},{"comment":"§6.7, Fig. 6, and Appendix E.1/F.1–F.2: The abstract and conclusion claim captions are “preferred by annotators” and enable “more source-faithful video reconstruction.” Human ranking/binding uses only 40 videos and three graduate annotators (F.1); reconstruction GSB uses the same small expert pool and conditions the generator on both caption text and the cited reference subset (E.1). Wins can therefore come from better reference selection alone rather than better phrase-level grounded language. Please either (a) run a controlled reconstruction where all methods receive the same gold positive references (text-only difference), and/or (b) substantially enlarge and statistically report the human study, and temper abstract/conclusion language to match the actual support.","section":"§6.7, Fig. 6, Appendix E.1, F.1–F.2"},{"comment":"§4.3 and Appendix B.3: HCD-GRPO’s load-bearing signals (keypoint scores, VQA error discounts, binding accuracy, CRSC) are themselves LLM-judge outputs with fixed ordinal scales and hand-set λb, λd, λc, λqa and 0.5/0.5 branch weights. Table 5 shows component necessity but not sensitivity to judge choice or λ’s. A short sensitivity sweep (e.g., ±λ or an alternate binding judge) and an explicit limitation that reward and eval judges share the same family of proxies would strengthen the claim that gains are from the hierarchical coverage-discount design rather than judge-specific overfitting.","section":"§4.3, Table 5, Appendix B.3"}],"minor_comments":[{"comment":"Appendix A.2 notes the 20k training corpus will not be released. Please state this prominently in the main text (contributions or experiments) and clarify what artifacts will be released (code, MRVBench, prompts, reward configs) so reproducibility expectations are clear.","section":"§5, Appendix A.2"},{"comment":"Table 1 header says “best result in each column is highlighted in green,” but the provided text does not convey highlighting; ensure camera-ready formatting makes open-source bests unambiguous, and that proprietary rows remain clearly excluded from “best.”","section":"Table 1"},{"comment":"Eqs. (1)–(4): define more explicitly how item-level binding decisions enter C_ref (whether a tag must pass the same judge as A_bind), and whether malformed-tag capping (τ_struct=0.2) was ever triggered on validation—useful for interpreting reward scale.","section":"§4.3.1–4.3.2"},{"comment":"Figure 4’s robustness score is a product of recall × binding × distractor rejection; briefly justify multiplicativity versus a weighted sum, and report per-component curves in the appendix to avoid hiding trade-offs.","section":"§6.5, Figure 4"},{"comment":"Minor polish: “V ATEX” spacing (Related Work); consistent naming of Qwen3.6 vs Qwen3-VL variants; ensure all baselines listed in §6.1 (e.g., VideoLLaMA3, LLaVA-Video) appear in the relevant tables or are marked as appendix-only.","section":"§2, §6.1"},{"comment":"Human study (F.1) reports mean ranks and binding accuracy but no inter-annotator agreement (e.g., Kendall’s W or pairwise agreement). Adding IAA would help readers weight Fig. S5.","section":"Appendix F.1"}],"recommendation":"minor_revision","confidential_remarks":"Solid systems paper for a CV/ML venue: new task + benchmark + competent post-training recipe. I recommend minor_revision rather than major because the automatic open-source SOTA, CaptionRefine control, general-caption retention, and ablations already support the core method claim; the main risk is over-claiming human/downstream fidelity and under-discussing Gemini-as-judge coupling. If the authors only add boilerplate limitations without any alternate-judge check or reconstruction control, the editor may want a second look at the abstract wording. Training-data non-release is acceptable if MRVBench and code/prompts ship as promised."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing: this is not another generic video-caption fine-tune. They define a concrete task—factual captions with phrase-level multi-image tags, distractor rejection, and same-entity grouping—and back it with MRVBench plus a two-stage recipe (mixed SFT + HCD-GRPO with DAES/CRSC) that actually moves the needle on grounding while holding general captioning on VDC/VCapsBench.\n\nWhat is new is the joint formulation and evaluation, not GRPO itself. Related work covers detailed captioning or multi-ref generation, but not whether an MLLM can select, bind, and group references inside the caption. The CaptionRefine control is the right one: post-hoc tag insertion preserves coverage but fails binding and subject consistency, so learning grounding with generation matters. Ablations line up with the claimed roles of the factual, DAES, and CRSC terms. An 8B model leading open-source on MRVScore and staying competitive with Gemini/GPT on several grounding metrics is a real result for people building multi-ref gen/edit pipelines.\n\nSoft spots, in proportion. Training data stays private—annoying for reproduction, normal for licensed video, and they do release the test set. Heavier issue: Gemini-3.1-Pro is both a top baseline and the main automatic judge for KP/VQA/binding/subject metrics, so automatic margins can be partly style-coupled. That does not sink the paper; human preference and reconstruction are independent pillars, but those pillars are thin—40 videos, three graduate annotators, and reconstruction conditions on both caption text and the cited reference subset, so wins can come from better reference selection alone. The abstract’s “preferred by annotators / more source-faithful reconstruction” language is a bit ahead of that N and that confound. Free reward weights and ordinal judge scores are ordinary systems knobs, not hidden circularity.\n\nMath is standard GRPO plus coverage-discounted rewards; citations look appropriate. Who it’s for: multi-ref video generation, editing, and evaluation people. Not a rewrite of video understanding. I’d send it to peer review, read the method and MRVBench sections carefully, and cite the task/benchmark if I work in this lane. Engage; don’t treat the human/downstream claims as settled until someone scales them or fixes the reference set in reconstruction.","headline":"Solid systems paper that cleanly defines multi-ref phrase-grounded video captioning, ships a usable benchmark and post-training recipe, and earns its open-source lead—with judge coupling and thin human/downstream evidence as real but bounded caveats.","tokens_in":24659,"tokens_out":603,"would_cite":true,"duration_ms":17422,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Video captions can name which reference images match which local phrases, not just describe the scene.","keywords":["video captioning","multi-reference grounding","phrase-level binding","GRPO","distractor rejection","MRVBench","multimodal LLMs","video reconstruction"],"falsifier":"If independent human raters (not the automatic judge) reverse the ranking on phrase-level binding, distractor rejection, and reconstruction fidelity—especially against the same proprietary models used as judges—or if post-hoc tag insertion matched one-pass RefCaptioner under blind human scoring, the central claim would fail.","tokens_in":24441,"feed_emoji":"🎬","tokens_out":864,"duration_ms":17572,"temperature":0.7,"pith_summary":"Standard video captioners describe what happens on screen but never say which of several reference photos those details come from. This paper defines multi-reference image-grounded video captioning: write a factual caption and attach each usable image tag right after the phrase it supports, while ignoring distractors and grouping multiple views of the same subject. RefCaptioner is a two-stage post-training recipe—mixed supervised fine-tuning plus Hierarchical Coverage-Discounted GRPO—that jointly trains reference selection, phrase-level binding, distractor rejection, and cross-reference consistency without giving up ordinary caption quality. The authors release MRVBench (real and AI-generated videos) and show the open 8B model leads open-source systems on that benchmark, stays competitive on standard caption benchmarks, and produces captions humans prefer and that reconstruct source videos more faithfully when fed to video generators.","feed_headline":"Captions that pin each detail to the right reference image","feed_subtitle":"An open 8B model learns selection, local binding, and distractor rejection without losing ordinary video description.","key_machinery":"Hierarchical Coverage-Discounted GRPO (HCD-GRPO): a dual-branch group-relative policy reward in which coverage of correct facts or correctly bound references is the positive signal and measurable errors (QA mistakes, bad bindings, distractor use, split same-entity tags) discount that signal.","core_discovery":"Multi-reference grounding can be learned together with factual video captioning: after mixed-data SFT, Hierarchical Coverage-Discounted GRPO—with a factuality branch and a grounding branch that rewards correct bindings, suppresses distractors (DAES), and enforces same-entity tag grouping (CRSC)—yields an open model that leads open-source systems on MRVBench while remaining competitive on general caption benchmarks and improving human preference and caption-conditioned reconstruction.","pith_inferences":["The same coverage-discounted split between ‘say what is true’ and ‘cite only supported evidence’ could transfer to multi-image document or long-video grounding.","Heavy dependence on one family of judge models for both rewards and leaderboard scores suggests a follow-up with held-out human-only labels to stress-test the reported margins.","Caption-conditioned reconstruction is an early signal that grounded captions may act as a portable interface layer between understanding models and generators."],"forward_implications":["Reference-conditioned video generation and editing can take phrase-tagged captions as explicit subject–attribute–scene wiring instead of free text alone.","Open captioners can approach proprietary multi-reference grounding without sacrificing ordinary detailed video description.","Post-hoc rewriting of plain captions is a weak substitute for learning selection and binding jointly with generation.","MRVBench becomes a shared test for whether models pick the right images, bind them locally, reject distractors, and group multi-view subjects."],"fun_headline_variants":["RefCaptioner binds video phrases to multiple reference images","Multi-reference grounding meets factual video captioning","Open model learns reference selection and distractor rejection","Hierarchical GRPO yields phrase-level reference-grounded captions","Mixed SFT plus GRPO leads open models on multi-reference bench"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The training rewards and benchmark scores both lean on large language-model judges as stand-ins for human judgments of factual coverage and whether a phrase truly matches a reference image.","fun_headline_variants_meta":{"raw":{"variants":["RefCaptioner binds video phrases to multiple reference images","Multi-reference grounding meets factual video captioning","Open model learns reference selection and distractor rejection","Hierarchical GRPO yields phrase-level reference-grounded captions","Mixed SFT plus GRPO leads open models on multi-reference bench"]},"model":"grok-4.5","effort":"low","cost_usd":0.002824,"raw_usage":{"total_tokens":1019,"prompt_tokens":768,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":28244000,"prompt_tokens_details":{"text_tokens":768,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":188,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":768,"tokens_out":63,"duration_ms":5473,"temperature":1.0,"reasoning_tokens":188,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T05:06:14.149803+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If independent human raters (not the automatic judge) reverse the ranking on phrase-level binding, distractor rejection, and reconstruction fidelity—especially against the same proprietary models used as judges—or if post-hoc tag insertion matched one-pass RefCaptioner under blind human scoring, the central claim would fail.","supporting_citations":[],"review_version":1}