{"id":"1fdd056f-393c-45f6-a705-ee586c3478d3","arxiv_id":"2506.11737","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Standard fine-tuned LLaVA-NeXT-Interleave outscores its DCI-enhanced variant on overall validation accuracy, though the DCI variant wins on coherence and slide-question tasks.","lead":"This challenge report fine-tunes LLaVA-NeXT-Interleave with and without an added dense vision connector and scores both on 22 multi-image, document, and dialogue benchmarks. It is a practical test of whether plug-and-play connector modules improve an already strong multimodal model, relevant to anyone assembling multimodal systems on limited GPUs.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table I contradicts the abstract and Section IV-B: DCI is reported lower than FT on the two highlighted 'DCI-strength' datasets (PropertyCoherence and SlideVQA) and higher on VISION, which is attributed to FT, so the headline qualitative claim is not supported by the displayed numbers.","rationale":"The reader correctly identified aggregation and post hoc labeling as weaknesses, but the more load-bearing issue is an internal contradiction between Table I and the paper's central claim. The abstract and Section IV-B name specific datasets as evidence for the FT-versus-DCI trade-off, yet those exact rows in Table I point the opposite way. This is not a matter of statistical methodology or interpretation; it is a factual inconsistency that any reader can verify from the printed table. The arithmetic error in the Original model's Document Understanding average reinforces that the numerical results have not been carefully checked. The core ranking (FT beats DCI on average in every task group) may survive, but the qualitative story about task-dependent strengths is unsupported as currently written. Therefore the paper should not be accepted without mandatory correction: the authors must either fix the table, fix the text, or report that the highlighted examples were chosen by mistake. A conditional verdict is appropriate rather than outright rejection, because the error is addressable and the overall comparison setup is reasonable for a challenge report. The proposed concrete test of reproducing three key rows from official evaluation would settle which part of the manuscript is wrong and allow the authors to issue a corrected version.","tokens_in":7106,"tokens_out":7791,"duration_ms":68707,"concrete_test":"Reproduce the validation scores for VISION, MIT-States PropertyCoherence, and SlideVQA using the released code and checkpoints on the official validation split, and compare each score against Table I. Specifically, determine whether the DCI model scores are 88.90, 91.50, and 51.50 (as printed in the first column) or 88.55, 94.75, and 65.25 (as printed in the second column). Also recompute the Document Understanding average from the six Original-model values; if 288.30/6 does not equal 35.30, the table needs a numerical correction. If the table is reproduced, the abstract's attribution of VISION to FT and PropertyCoherence/SlideVQA to DCI is false; if the table is not reproduced, the reported numbers are unreliable.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim in the abstract and Section IV-B is that the standard FT model excels in vision-heavy tasks like VISION, NLVR2, and Fashion200K, while the DCI-enhanced version shows particular strength on MIT-States PropertyCoherence and SlideVQA. However, Table I, with columns labeled 'FT w/ DCI, FT, Original', gives: VISION = 88.90 (DCI), 88.55 (FT), 86.40 (Original); PropertyCoherence = 91.50 (DCI), 94.75 (FT), 93.50 (Original); SlideVQA = 51.50 (DCI), 65.25 (FT), 63.50 (Original). Under the printed labels, DCI beats FT on VISION but loses on PropertyCoherence and SlideVQA, directly contradicting the text. If the column labels were swapped, DCI would beat FT on PropertyCoherence and SlideVQA, but then DCI would also beat FT on Fashion200K and NLVR2, contradicting the claim that FT excels there. No consistent labeling makes all the named examples match the reported qualitative conclusions. Additionally, the Document Understanding average for the Original model is arithmetically wrong: the six listed values sum to 288.30, giving 48.05, not the reported 35.30. This indicates the table's numerical entries or aggregation have not been carefully validated. Because the headline conclusion is supported by the examples that are internally inconsistent with Table I, the central qualitative claim is not trustworthy as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes the authors' submission to the ICME25 INOVA Challenge Track A. The authors take LLaVA-NeXT-Interleave, fine-tune it on 22 challenge datasets, and compare three configurations: the original model, standard fine-tuning (FT), and fine-tuning with the Dense Channel Integration (DCI) connector. The reported validation results in Table I are used to argue that standard FT gives the best overall accuracy, excelling on VISION, NLVR2, and Fashion200K, while DCI helps on datasets requiring semantic coherence such as MIT-States PropertyCoherence and SlideVQA. The authors also report their second-place test-set score of 74.9 and discuss training-loss behavior.","tokens_in":7384,"tokens_out":4548,"duration_ms":40959,"significance":"If the comparisons were reliable, this would be a useful ablation of a plug-and-play connector in the interleaved multi-image setting, with practical implications for choosing between standard fine-tuning and DCI-enhanced fine-tuning depending on task family. The paper is also valuable as a reproducible challenge report: the code is released, the setup is simple (one A100 GPU, one epoch of fine-tuning), and the test-set result confirms competitiveness. However, the qualitative conclusions as stated are not supported by the printed table, and the single-run nature and unweighted aggregation limit the strength of the quantitative claims; the manuscript therefore requires correction before its central message can be accepted.","major_comments":[{"comment":"The text's central qualitative claims are contradicted by the numbers in Table I. The abstract and Section IV-B state that FT excels on VISION, NLVR2, and Fashion200K, while DCI is strong on MIT-States PropertyCoherence and SlideVQA. Under the column labels \"FT w/ DCI, FT, Original\", the table shows VISION 88.90 (DCI) vs. 88.55 (FT), PropertyCoherence 91.50 (DCI) vs. 94.75 (FT), and SlideVQA 51.50 (DCI) vs. 65.25 (FT). Thus DCI beats FT on VISION, which the text attributes to FT, and loses on the two datasets cited as DCI strengths. Swapping the first two columns would make the two DCI examples consistent but would make DCI beat FT on Fashion200K and NLVR2 as well, contradicting the other half of the claim. No consistent labeling of the columns makes all named examples match the text; the headline comparison must be corrected or the table entries must be verified.","section":"Section IV-B, Table I"},{"comment":"The average reported for the Original model in the Document and Knowledge-Based Understanding block is arithmetically wrong. The six row entries (63.50, 90.50, 26.38, 78.50, 17.90, 11.52) sum to 288.30, so the average is 48.05, not the printed 35.30. The corresponding DCI and FT averages are consistent (49.17 and 56.10), which suggests a localized error, but it undermines confidence in the table's other entries and in the \"highest overall accuracy\" comparison, since those averages are the basis for the claim.","section":"Table I, Document and Knowledge-Based Understanding average"},{"comment":"The conclusion that FT \"achieves the highest overall accuracy\" rests on unweighted averages over 22 datasets from a single validation split, with no error bars or repeated runs. The paper uses a 9:1 split of the provided training data and reports one run per configuration, so it is not possible to assess whether the observed margins (e.g., 1-2 accuracy points in several blocks) are stable. The authors should add confidence intervals or at least report the number of examples per dataset, and should temper the claim accordingly if those are unavailable.","section":"Section IV-B and Section III-C"}],"minor_comments":[{"comment":"Please fix the typos \"A VERAGE\" in Table I and Figure 4, \"ROGUE-L\" for ROUGE-L, and the line-break artifacts in \"LLaV A\" throughout the text.","section":"Throughout"},{"comment":"References [4] and [8] are the same CLIP paper; the reference numbering should be corrected, since [3] is GPT-3 and [4] is currently listed as CLIP.","section":"References"},{"comment":"The dataset names are inconsistent: Section III-B writes \"MMQA\", Table I uses \"MultiModalQA\" and \"ManyModalQA\", and Figure 4 uses only \"MultiModalQA\"; please align the names between text and tables.","section":"Section III-B and Table I"},{"comment":"Figure 4 reports the test-set result but does not state which of the three validation configurations (Original, FT, or FT with DCI) produced that submitted run; please clarify.","section":"Section IV-B and Figure 4"},{"comment":"The test-set scores in Figure 4 differ substantially from the validation scores for several datasets (e.g., SlideVQA 81.40 test vs. 51.50/65.25 validation), but the paper does not explain whether different splits or test-time procedures were used; please provide the protocol.","section":"Section IV-B"},{"comment":"The phrase \"we just pre-train the model\" is inaccurate for the described procedure; the paper performs fine-tuning, not pre-training, so the wording should be corrected.","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"The paper is essentially a challenge report with limited novelty beyond the application of an existing connector to an interleaved model. The main issue is the table-text inconsistency, which is serious enough to require a major revision; if the authors correct the table or the claims and add basic uncertainty information, the manuscript could be acceptable for a workshop-style or challenge-track venue. The self-citation of [14] is not directly relevant to the method, but it is not a problem by itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper is a clean challenge-style empirical comparison of standard fine-tuning vs DCI-enhanced fine-tuning of LLaVA-NeXT-Interleave on 22 datasets, but its headline qualitative conclusion is contradicted by the numbers in its own Table I, and the table has an arithmetic error. So treat the central claim with suspicion, though the raw data may still be useful.\n\nWhat's new: the specific comparison itself — DCI plugged into LLaVA-NeXT-Interleave, evaluated on the INOVA Track A datasets — appears not to be in the prior literature. The paper is honest about the setup: one A100, one epoch, 9:1 split, and it includes training loss curves and test set numbers for their second-place challenge entry. That transparency is worth credit.\n\nThe soft spots, in order. First, the abstract and Section IV-B claim the standard FT model excels on VISION, NLVR2, Fashion200K, while DCI is stronger on MIT-States PropertyCoherence and SlideVQA. Table I shows the opposite: DCI beats FT on VISION (88.90 vs 88.55), but loses on PropertyCoherence (91.50 vs 94.75) and SlideVQA (51.50 vs 65.25). No column swap fixes this — if you swap labels, DCI would beat FT on NLVR2 and Fashion200K too. So the example-based claim doesn't match any consistent labeling. That's a load-bearing inconsistency, not a typo. Second, the Document Understanding average for the Original model is arithmetically wrong: the six values sum to 288.30, average 48.05, but the table reports 35.30. That suggests the numerical entries weren't carefully checked. Third, the comparison relies on single-run validation scores without error bars, unweighted per-dataset averages, and post hoc grouping of \"semantic coherence\" datasets, which weakens the overall ranking.\n\nNone of this kills the paper's value as a challenge report. The data table, once corrected, tells a plausible story: standard FT is better on average, DCI helps some tasks and hurts others. But as written, the central qualitative claim is not supported by the displayed evidence.\n\nWho's this for? People building interleaved multimodal systems and INOVA challenge participants. It deserves a serious referee — the empirical comparison is useful and the flaws are fixable — but the authors should be asked to reconcile the table and text, add error bars or multiple seeds, and justify the dataset grouping. I wouldn't cite it in my own work until those corrections are made.","headline":"A useful challenge-report numbers table that contradicts its own abstract and Section IV-B, so the headline claim is not trustworthy as written.","tokens_in":7916,"tokens_out":3045,"would_cite":false,"duration_ms":27497,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Standard fine-tuning achieves the highest overall accuracy, while the DCI-enhanced variant wins on semantic-coherence and structured-change datasets.","keywords":["interleave","LLaVA","comprehension","dense connector","Dense Channel Integration","multi-image reasoning","plug-and-play technique","vision-language model"],"falsifier":"Recompute the comparison with per-example weighting or repeated seeds across the 22 datasets: if the DCI-enhanced model's average reaches or exceeds the standard model's under any plausible reweighting, the claim that standard fine-tuning is best overall fails.","tokens_in":6904,"feed_emoji":"🧩","tokens_out":4208,"duration_ms":36098,"temperature":0.7,"pith_summary":"This paper asks whether a plug-and-play connector, Dense Channel Integration (DCI), improves an interleaved multi-image model, LLaVA-NeXT-Interleave, across 22 datasets spanning multi-image reasoning, document and knowledge understanding, and interactive multimodal dialogue. Fine-tuning the base model gives the best overall accuracy, while the DCI-enhanced version matches or beats it on tasks that require semantic coherence and structured change, such as MIT-States PropertyCoherence and SlideVQA. The paper reports that both fine-tuned variants far outperform the original model, and identifies a trade-off between general balance and coherence-specific strength.","feed_headline":"Standard fine-tuning beats dense-connector model overall","feed_subtitle":"The DCI variant wins on coherence-heavy benchmarks like MIT-States and SlideVQA.","key_machinery":"The key object is the Dense Channel Integration (DCI) connector, which replaces the vision encoder's final-only output with a hierarchical fusion of all layer features. The connector partitions the $L$ vision-encoder layers into $G$ groups of $M = L/G$ adjacent layers, averages each group to produce fused representations $G^L_i = \\frac{1}{M}\\sum_{k=(i-1)M+1}^{iM} V_k$, and concatenates these with the final layer feature $V_L$ to form the vision embedding $E_V = \\mathrm{Concatenate}([G^L_1, \\dots, G^L_G, V_L])$. This adds fewer than 2% trainable parameters and needs only single-stage instruction tuning; the paper's comparison isolates the effect of this connector on top of the LLaVA-NeXT-Interleave base model.","core_discovery":"The central claim is that standard fine-tuning of LLaVA-NeXT-Interleave achieves the highest overall accuracy, while adding the DCI connector changes the strength profile rather than uniformly improving it. On vision-heavy benchmarks (VISION, NLVR2, Fashion200K) the standard model excels; on datasets demanding semantic coherence or structured change understanding (MIT-States PropertyCoherence, SlideVQA) the DCI version is stronger. The paper also reports training-loss dynamics in which DCI converges faster early but shows higher variance later, and it reports a second-place test result with strong accuracy on majority datasets but low ROUGE-L on free-answer generation datasets.","pith_inferences":["Because the paper's overall ranking depends on unweighted averages, a task-weighted or example-weighted aggregation might shift the standard-vs-DCI comparison; this is an editorial caveat, not a paper claim.","DCI's faster early convergence, if it transfers to low-data regimes, could make it the better choice for few-shot adaptation even though it loses on the full benchmark average.","The task-family split (vision-heavy vs coherence-heavy) hints at a simple selection rule: use standard fine-tuning for generic visual QA, DCI for tasks whose questions hinge on property consistency or structured change.","Combining the two models in an ensemble or a two-stage router could recover the best of both profiles; this is a testable extension the paper does not run."],"forward_implications":["Standard fine-tuning of LLaVA-NeXT-Interleave yields the best overall accuracy across the three challenge tasks, so a single model choice is adequate for general use.","The DCI connector gives a measurable advantage on semantic-coherence and structured-change benchmarks such as MIT-States PropertyCoherence and SlideVQA.","Both fine-tuned variants exceed the original model by wide margins, so task-specific fine-tuning matters more than connector choice for interleaved multi-image models.","The low ROUGE-L scores on free-answer datasets (Spot-the-Diff, IEdit, Birds-to-Words) indicate that open-ended generation remains a bottleneck for this training strategy.","A hybrid fine-tuning strategy combining standard and DCI approaches is a natural next step suggested by the observed trade-off."],"supporting_citations":[{"why":"Defines the LLaVA-NeXT-Interleave base model that is fine-tuned and modified with the DCI connector.","marker":"[10]"},{"why":"Supplies the Dense Connector method, including the DCI computation that the paper integrates into the vision encoder.","marker":"[16]"},{"why":"Specifies the SigLIP vision backbone used for training and evaluation of both model variants.","marker":"[19]"},{"why":"Introduces LLaVA-1.5, the improved baseline architecture that LLaVA-NeXT-Interleave builds on.","marker":"[9]"},{"why":"Motivates the contrastive vision-language pretraining that the vision encoder relies on, and is cited as a compatible encoder for the DCI connector.","marker":"[8]"}],"fun_headline_variants":["Standard model beats DCI overall in 22-dataset test","DCI connector shifts strengths, not overall win","Standard fine-tuning tops DCI in comprehensive benchmark","No overall win for DCI connector despite niche gains","Standard LLaVA beats DCI on overall accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The overall comparison rests on unweighted averages of one validation-run accuracy or ROUGE-L score per dataset; if datasets were weighted by size or difficulty, or runs repeated with different seeds, the ranking between standard and DCI fine-tuning could change.","fun_headline_variants_meta":{"raw":{"variants":["Standard model beats DCI overall in 22-dataset test","DCI connector shifts strengths, not overall win","Standard fine-tuning tops DCI in comprehensive benchmark","No overall win for DCI connector despite niche gains","Standard LLaVA beats DCI on overall accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000475,"raw_usage":{"total_tokens":2307,"prompt_tokens":847,"completion_tokens":1460,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":1384}},"tokens_in":463,"tokens_out":1460,"duration_ms":10321,"temperature":1.0,"reasoning_tokens":1384,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:03:47.986800+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the comparison with per-example weighting or repeated seeds across the 22 datasets: if the DCI-enhanced model's average reaches or exceeds the standard model's under any plausible reweighting, the claim that standard fine-tuning is best overall fails.","supporting_citations":[{"cited_title":"Llava-next: Improved reasoning, ocr, and world knowledge,","cited_arxiv_id":null,"evidence_quote":"Defines the LLaVA-NeXT-Interleave base model that is fine-tuned and modified with the DCI connector."},{"cited_title":"Dense connector for mllms,","cited_arxiv_id":null,"evidence_quote":"Supplies the Dense Connector method, including the DCI computation that the paper integrates into the vision encoder."},{"cited_title":"Sigmoid loss for language image pre-training,","cited_arxiv_id":null,"evidence_quote":"Specifies the SigLIP vision backbone used for training and evaluation of both model variants."},{"cited_title":"Learning transferable visual models from natural language supervision,","cited_arxiv_id":null,"evidence_quote":"Motivates the contrastive vision-language pretraining that the vision encoder relies on, and is cited as a compatible encoder for the DCI connector."}],"review_version":1}