{"id":"053e1092-574d-4bea-b6c1-16bd7220cc97","arxiv_id":"2504.20199","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A focus-centric reasoning format plus a 150K synthetic dataset improves VLM accuracy across seven multi-image benchmarks.","lead":"This paper trains vision-language models to answer multi-image questions by decomposing them into smaller steps, each focusing on a subset of images. It introduces a 150,000-example synthetic dataset built with open-source models and reports average gains of about three and two percentage points on two model families.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported gains are not attributable to the Focus-Centric Visual Chain format because the experimental design includes no equivalently sized control dataset without that format.","rationale":"The reader's weakest assumption is correct and is the main load-bearing gap. The paper shows that LoRA fine-tuning on VISC-150K improves scores on seven multi-image benchmarks and does not degrade four general single-image benchmarks (Table 2), which is genuine evidence that the dataset is useful for multi-image instruction tuning. However, the paper's headline contribution is the Focus-Centric Visual Chain paradigm, and Section 4.4 attributes the gains to the chain format ('These improvements can be attributed to three key characteristics...'). No experiment varies the reasoning format while holding data constant. RQ1 varies dataset size, RQ2/RQ3 analyze sub-tasks and image counts, RQ4 checks general ability, and RQ5 is a human quality audit; none isolates the format. The causal claim is therefore unsupported as presented. The proposed control experiment is feasible and standard for a paradigm-introduction paper. The SOTA statement is also overstated as written: Table 1 shows GPT-4V/GPT-4o ahead of the VISC models on MMIU (55.70 vs 52.76), so 'state-of-the-art on four out of the seven' seems to rely on an unstated open-source-only comparison. This is secondary but should be corrected. Given the missing control, the verdict should remain CONDITIONAL: the dataset may well be valuable, but acceptance of the paradigm attribution requires the control experiment.","tokens_in":17036,"tokens_out":5528,"duration_ms":56433,"concrete_test":"Construct two control datasets matched to VISC-150K in size, image source, and final question/answer content: (A) flat QA containing only the final question and final answer, and (B) the same final QA preceded by a generic 'think step by step' instruction. Fine-tune LLaVA-OneVision-7B and Qwen2-VL-7B on each control and on VISC-150K using identical LoRA hyperparameters (rank 16, learning rate 1e-5, one epoch, batch size 8), then evaluate on MMIU, MuirBench, MIRB, and BLINK. If the VISC-trained models do not beat both controls by a clear margin on most benchmarks, the reported gains cannot be attributed to Focus-Centric Visual Chains rather than to additional multi-image instruction tuning.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that fine-tuning on VISC-150K improves multi-image reasoning because the data follows the Focus-Centric Visual Chain paradigm. The evidence in Section 4.4 and RQ1-RQ5 supports only the weaker conclusion that fine-tuning on VISC-150K improves benchmark scores; it does not support the causal attribution to the chain format. There is no control condition in which the same 150K images, final questions, and final answers are presented without the focus-centric sub-question/focus decomposition (e.g., as flat QA or with a generic CoT prompt), and no comparison to an equally sized alternative multi-image instruction dataset. RQ1 varies dataset size, RQ2/RQ3 analyze sub-tasks and image counts, RQ4 checks general ability, and RQ5 is a human quality audit; none isolates the reasoning format. A secondary but related issue is that the statement 'achieving new state-of-the-art on four out of the seven' is not supported by Table 1 as written if closed-source models are included: GPT-4V/GPT-4o scores 55.70 on MMIU, above the best VISC model's 52.76, so the four SOTA claims appear to rely on an unstated open-source-only comparison. The missing control is the more fundamental concern: without it, a generic benefit from additional multi-image instruction data cannot be distinguished from the benefit of the proposed paradigm.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Focus-Centric Visual Chain (FCVC), a multi-step reasoning format for multi-image vision-language tasks, and Focus-Centric Data Synthesis (FCDS), a bottom-up pipeline that builds the VISC-150K dataset of 150K reasoning instances using open-source models. The authors fine-tune LLaVA-OneVision-7B and Qwen2-VL-7B via LoRA on VISC-150K and report average accuracy gains of 3.16% and 2.24% on seven multi-image benchmarks, along with analyses of data scale, sub-task breakdown, input-image count, general capability retention, and a human quality audit. The paper claims the gains validate the FCVC paradigm and that the method achieves state-of-the-art on four of the seven benchmarks.","tokens_in":17300,"tokens_out":5588,"duration_ms":46849,"significance":"The VISC-150K dataset and the FCDS pipeline are potentially useful community resources: they are built entirely with open-source models, are accompanied by a human-evaluation audit (97.5% validity on 200 samples), and the paper demonstrates consistent benchmark improvements across two different model architectures. If the causal attribution to the focus-centric format were established, the work would offer a scalable recipe for improving multi-image reasoning in VLMs. However, the current experimental design does not isolate the proposed paradigm from the generic effect of adding multi-image instruction data, and the state-of-the-art claim is not supported as written when closed-source models are included.","major_comments":[{"comment":"The central claim that the performance gains are due to the Focus-Centric Visual Chain format is not supported by the experimental design. There is no control condition in which the same 150K instances are fine-tuned without the focus-centric sub-question decomposition (e.g., as flat QA or with a generic chain-of-thought prompt), nor a comparison against an equally sized alternative multi-image instruction dataset. RQ1–RQ5 vary data scale, sub-tasks, image counts, general ability, and data quality, but none of them manipulates the reasoning format while holding the underlying images, questions, and answers fixed. Without such an ablation, the observed improvements could be a generic effect of additional multi-image instruction data rather than a benefit of the proposed paradigm.","section":"§4.4, §4.5 (RQ1–RQ5)"},{"comment":"The statement that the method 'achieves new state-of-the-art on four out of the seven' benchmarks is only true if the comparison is restricted to open-source models. Table 1 lists GPT-4V/GPT-4o scores of 55.70 on MMIU and 68.00 on MuirBench, both higher than the best VISC-fine-tuned model's 52.76 and 60.16, respectively. The text should explicitly qualify that the state-of-the-art claim is among open-source models, or otherwise revise the claim to reflect the full baseline set.","section":"§4.4, Table 1, Abstract"},{"comment":"The paper states that VISC-150K 'consistently brings performance improvements across all challenging benchmarks,' but Table 1 reports a 0.48-point drop for Qwen2-VL-VISC on Mantis-Eval relative to the Qwen2-VL baseline. The reported average gain of 2.24% includes this negative value, so the wording 'consistently' and 'across all' is inaccurate. This should be reworded to describe the actual pattern (e.g., improvements on six of seven benchmarks).","section":"Abstract, §4.4, Table 1"},{"comment":"The FCDS pipeline uses LLaVA-OneVision-7B as both the Extractor and the Annotator, and LLaVA-OneVision-7B is also one of the two models fine-tuned on the resulting dataset. This creates a potential self-distillation effect that could inflate the reported gains on the LLaVA-OneVision architecture, because the training data is generated by the same model family that is later evaluated. The paper does not discuss or control for this distributional overlap. The gains on Qwen2-VL suggest this is not the sole explanation, but the missing control condition in the first major comment also prevents this confound from being resolved.","section":"§3.2, §4.1"}],"minor_comments":[{"comment":"The abstract ends with ': /gtbVISC', which appears to be a leftover placeholder or malformed link and should be removed or replaced with the actual dataset URL.","section":"Abstract"},{"comment":"The text says 'The statics of each benchmark' and should read 'statistics'.","section":"§4.3"},{"comment":"The phrase 'close-source VLMs' should be 'closed-source VLMs'.","section":"§4.2"},{"comment":"In the formalization, 'the model focuses on a visual evidence subset G⊆G' uses the same symbol for the subset and the full collection; this should be written as G_i ⊆ G (or an equivalent distinct notation) to avoid ambiguity.","section":"§3.1"},{"comment":"The appendix ends with '...Fleiss’κ = 0.637). s', which contains a stray 's' at the end; this is likely a typographical artifact.","section":"Appendix F"},{"comment":"The MVBench reference (Li et al., 2024b) contains a long sequence of identical-looking author names that appears to be a placeholder or formatting error; the authors should verify the citation.","section":"References"},{"comment":"The caption contains the typo 'Foucs' and should read 'Focus'.","section":"Figure 7 caption"}],"recommendation":"major_revision","confidential_remarks":"The dataset and synthesis pipeline are valuable contributions, and the benchmark gains are promising. The missing control ablation is, however, a standard requirement for the paper's central causal claim, and the state-of-the-art overclaim relative to closed-source models should be corrected. I would accept a revised version that adds an equivalently sized control without the focus-centric format and that qualifies the SOTA statement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things worth knowing about this paper. First, it is a legitimate empirical contribution: a 150K multi-image reasoning dataset (VISC-150K) plus a bottom-up synthesis pipeline (FCDS) that produces focus-centric chain-of-thought style supervision. Fine-tuning LLaVA-OneVision-7B and Qwen2-VL-7B on this data gives consistent gains across seven multi-image benchmarks, with a modest human quality audit and a check on general VLM abilities. That is useful, reproducible-in-principle work, and the dataset release, if it happens, would likely be used by people training multi-image VLMs.\n\nThe second thing is that the paper's central claim—that the gains come from the Focus-Centric Visual Chain format—is not supported by the experiments. There is no control condition. The paper does not compare against fine-tuning on the same 150K questions and answers formatted as flat QA, or against an equally sized alternative multi-image instruction set. RQ1 through RQ5 cover data scale, sub-tasks, image counts, general ability, and data quality; none of them isolate the format. So the evidence supports 'VISC-150K improves multi-image benchmarks,' not 'focus-centric chaining is the cause.' This is a load-bearing gap, not a nit.\n\nTwo smaller issues. The state-of-the-art claim on four of seven benchmarks does not match Table 1 as printed: GPT-4V beats the best fine-tuned model on MMIU (55.70 vs. 52.76) and MuirBench (68.00 vs. 49.62), so that statement apparently depends on an open-source-only comparison. The dataset and code are described as released, but the anonymous link in the abstract is garbled, so independent replication is not currently possible. The human evaluation is only 200 samples with Fleiss' kappa of 0.637, moderate agreement; the 97.5% validity number should be read in that context. These are fixable but need attention.\n\nWho this is for: people building multi-image instruction data and evaluating VLMs on multi-image reasoning. It deserves a serious referee. I would ask the authors to add a control dataset and to rewrite the SOTA claim. Without the control, the contribution is a dataset and recipe—still useful and publishable, but the title-level claim about the paradigm needs to be dialed back.","headline":"A genuinely useful dataset and recipe for multi-image VLM training, but the headline claim that the chain format causes the gains is not isolated by the experiments.","tokens_in":17807,"tokens_out":3077,"would_cite":true,"duration_ms":29871,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Teaching vision-language models to reason stepwise across images improves multi-image benchmark accuracy by 3.16 percentage points in LLaVA-OneVision and 2.24 in Qwen2-VL.","keywords":["multi-image understanding","vision-language models","reasoning chain","data synthesis","fine-tuning","VISC-150K","Focus-Centric Visual Chain","LoRA"],"falsifier":"Train the same two base models with LoRA on a matched 150K-sample dataset that contains the same questions, images, and final answers but replaces the multi-step focus-chain traces with a single flat answer (or with an unstructured chain-of-thought that does not explicitly select image subsets). If MMIU and MuirBench gains shrink to near zero, the focus-centric format is the active ingredient; if they persist, the paper's attribution to the chain approach is unsupported.","tokens_in":16868,"feed_emoji":"🧩","tokens_out":12769,"duration_ms":99551,"temperature":0.7,"pith_summary":"This paper claims that vision-language models can be made substantially better at tasks involving several images at once by teaching them to reason in a stepwise 'focus-centric' manner: at each step the model names a sub-question, points to the specific image or images it is looking at, answers, and then moves on. To make this trainable, the authors build a 150K-sample dataset, VISC-150K, whose entries consist of such visual reasoning chains, synthesized bottom-up with open-source models. Fine-tuning two strong base VLMs on this dataset improves their average accuracy on seven multi-image benchmarks by 3.16 and 2.24 percentage points respectively, and reaches a new best score on four of them, while general single-image capabilities stay roughly flat. The paper's core assertion is that the chain structure, not just additional multi-image data, is what drives the gains.","feed_headline":"Chained visual focus lifts multi-image reasoning accuracy 3.16%","feed_subtitle":"A 150K-sample dataset of chained reasoning steps drives the gain; four benchmarks set new best scores.","key_machinery":"The central object is the Focus-Centric Visual Chain, an ordered reasoning trace $R = [(q_i, G_i, a_i, z_i)]_{i=1}^N$ in which each step $i$ generates a sub-question $q_i$, selects a visual focus $G_i$ (a subset of the input image collection), answers $a_i$, and emits a stopping signal $z_i$. The chain is what the synthesized dataset encodes and what the fine-tuned model learns to reproduce. Its load-bearing property is that every step is anchored to a minimal image subset, so the model practises selective attention across images rather than attending to all inputs at once. The data-side machinery is the four-stage bottom-up synthesis loop (feature extraction, pair connection, relevance annotation, question generation) that turns raw image collections into chained question–answer paths without relying on expensive closed-source generation.","core_discovery":"The discovery is a training approach and a data-generation pipeline. In the Focus-Centric Visual Chain approach, a VLM confronted with a set of images and a question iteratively produces a sub-question, selects the minimal subset of images needed to answer it, derives an answer, and accumulates these until it can synthesize a final answer. The accompanying Focus-Centric Data Synthesis framework builds training examples from the bottom up: an extractor writes detailed textual profiles of each image, a connector links related image pairs, an annotator labels each link as temporal, spatial, or semantic, and a questioner produces chained sub-questions plus a composite question. All components use open-source models, so the pipeline is reproducible and cheap. Fine-tuning LLaVA-OneVision-7B and Qwen2-VL-7B with LoRA on the resulting VISC-150K yields average gains of 3.16% and 2.24% across seven multi-image benchmarks, and the authors report that the model learns to apply the chained reasoning format at test time, including on video frames and GUI navigation screenshots.","pith_inferences":["The most direct test the paper does not run is a matched control: fine-tuning on the same 150K question–answer pairs with the chain structure flattened or removed. If the gains persist, the improvement may come from more multi-image instruction data rather than from the focus-centric format itself; if they vanish, the format is the active ingredient.","Because the synthesis pipeline uses only open-source models and structured prompts, it can be pointed at new image domains (diagrams, charts, code screenshots) that the authors note are currently untested, effectively turning the framework into a general multi-image reasoning data generator.","The dynamic focus selection mechanism resembles an attention-routing procedure; a natural extension is to make the focus subset explicit at test time (e.g., allowing the model to crop or zoom into chosen images), which could further reduce interference from irrelevant images in large sets.","The reported 97.5% human-validated accuracy of the synthesized chains, with inter-annotator agreement kappa = 0.637, suggests the data pipeline is reliable enough to serve as a cheaper substitute for closed-source distillation in other multimodal reasoning settings."],"forward_implications":["Fine-tuning on VISC-150K improves both LLaVA-OneVision-7B and Qwen2-VL-7B on every one of the seven benchmarks except one (Mantis-Eval for Qwen2-VL, which drops 0.48 points), with average gains of 3.16 and 2.24 percentage points.","The approach transfers to video understanding: MVBench improves by 1.53 and 1.01 points, consistent with treating video as a temporally ordered multi-image input.","General capabilities are not sacrificed: on HallusionBench, MMStar, MMMU, and MathVista, the fine-tuned Qwen2-VL stays within ±0.3 points on three benchmarks and improves by 1.5 on HallusionBench.","Performance scales with data size: increasing the fine-tuning subset from 25K to 125K yields rapid gains, with diminishing but continued improvements toward 150K, suggesting the synthetic pipeline can be scaled further.","The model generalizes to sub-tasks and image types not present in VISC-150K, e.g., geographic understanding in MuirBench, indicating the chains teach a transferable skill rather than memorized question formats."],"supporting_citations":[{"why":"base model for one fine-tuning branch and for the extractor/annotator components of the synthesis pipeline","marker":"(Li et al., 2024a)"},{"why":"second base model; the paper fine-tunes Qwen2-VL-7B on VISC-150K and evaluates the gains","marker":"(Wang et al., 2024b)"},{"why":"LoRA is the fine-tuning method used to adapt both base models","marker":"(Hu et al., 2022)"},{"why":"MMIU benchmark, one of the seven evaluation suites and a source of the largest reported gains","marker":"(Meng et al., 2024)"},{"why":"MuirBench benchmark, used to analyse sub-task gains and cross-domain generalization","marker":"(Wang et al., 2024a)"},{"why":"MIRB benchmark, one of the seven evaluation suites","marker":"(Zhao et al., 2024)"},{"why":"provides the Mantis-Idefics2 baseline and the Mantis-Eval benchmark, both used in the evaluation table","marker":"(Jiang et al., 2024)"},{"why":"MVBench video benchmark, used to show the approach transfers to temporally ordered frames","marker":"(Li et al., 2024b)"},{"why":"chain-of-thought reasoning, the conceptual basis for decomposing the task into sub-questions","marker":"(Wei et al., 2022)"}],"fun_headline_variants":["Focus-chain gains 3.16% on multi-image VLM","Visual chain lifts multi-image VLM by 3.16%","Chained focus adds 3.16% to seven multi-image tests","Focus-chain method hikes VLM multi-image scores 3.16%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the performance gains come from the focus-centric chain format itself, but it never compares against fine-tuning on an equally large set of multi-image instruction data without that chain structure, so extra data alone could explain part or all of the improvement.","fun_headline_variants_meta":{"raw":{"variants":["Focus-chain gains 3.16% on multi-image VLM","Visual chain lifts multi-image VLM by 3.16%","Chained focus adds 3.16% to seven multi-image tests","Focus-chain method hikes VLM multi-image scores 3.16%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001346,"raw_usage":{"total_tokens":5468,"prompt_tokens":948,"completion_tokens":4520,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":4443}},"tokens_in":564,"tokens_out":4520,"duration_ms":29063,"temperature":1.0,"reasoning_tokens":4443,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:34:45.205242+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same two base models with LoRA on a matched 150K-sample dataset that contains the same questions, images, and final answers but replaces the multi-step focus-chain traces with a single flat answer (or with an unstructured chain-of-thought that does not explicitly select image subsets). If MMIU and MuirBench gains shrink to near zero, the focus-centric format is the active ingredient; if they persist, the paper's attribution to the chain approach is unsupported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"chain-of-thought reasoning, the conceptual basis for decomposing the task into sub-questions"}],"review_version":1}