{"id":"fcc0ef05-c84e-44f4-b428-f888515d58ee","arxiv_id":"2603.12478","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Goal-driven selection of 1× multimodal instruction subsets reaches a 512k Uni-10x baseline after ~27–35k samples and improves accuracy by up to +3.08 pp under a fixed Qwen3-VL recipe.","lead":"GDO picks smaller, goal-tuned image-video instruction subsets using six sample descriptors and beats a 512k uniform baseline with roughly 15–19× fewer samples under a fixed Qwen3-VL training recipe. It matters because multimodal post-training is expensive and most mixed pools waste budget on redundant short-video and image QA.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Peak-match gains may be inflated by comparing optimized 1× subsets only against a 10× uniform baseline, not against strong 1× selection methods under the same budget.","rationale":"The reader correctly flags the hand-fixed scorer and single-backbone probe as risks, and the fixed-recipe design is a real strength. The more load-bearing gap for the title claim, however, is the comparison contract itself: all headline Peak Match / Reduction numbers (Table 1) and the abstract’s “far fewer samples than Uni-10x” are measured against a 10× uniform baseline, not against competitive 1× selectors at matched budget. That leaves open whether GDO’s six descriptors and feasibility presets are doing the work, or whether almost any aggressive 1× filter would produce similar early crossings. Profile spectrum and ablations give partial internal evidence but do not replace the missing same-budget baselines. Verdict stays CONDITIONAL (same direction as the reader) until those controls or multi-seed/backbone checks land; agreement is partial because the reader’s weakest assumption is real but secondary to the baseline confound for the central “less data, faster convergence” claim.","tokens_in":17730,"tokens_out":811,"duration_ms":8154,"concrete_test":"Under the identical one-epoch Qwen3-VL-8B recipe and checkpoints, train three additional 1× controls at the Temp+ budget (53.3k): (a) uniform random 1×, (b) top-k by quality_score alone, (c) top-k by VDS alone (same video-ratio floors as Temp+). Recompute peak-match vs Uni-10x and final Acc on MVBench/VideoMME/MLVU. If random or quality-only already reaches Uni-10x by ~30k and matches Temp+ within ~0.5 pp, the load-bearing claim weakens to budget compression; if Temp+ retains ≥1 pp and earlier crossing, the goal-driven scorer is supported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (less data + faster convergence + higher accuracy under a fixed train/eval contract) rests on GDO 1× subsets beating a fixed 512k Uni-10x uniform baseline at peak-match points of 26.6k–35.4k samples (Table 1, Fig. 1/3). That comparison confounds two effects: (i) discarding ~90% of the pool and (ii) goal-driven scoring/feasibility. Because Uni-10x is a uniform 10× control rather than a strong same-budget 1× selector (e.g., random 1×, quality-only, VDS-only, or prior instruction-data selectors such as LESS/AlpaGasus-style ranking), the reported 14–19× reductions and +0.84–+3.08 pp gains may largely reflect “any non-uniform 1× curation beats 10× uniform” rather than the specific six-descriptor shared scorer (Eqs. 9–11) and goal presets. The paper’s own MinLoss/Diverse/Temp spectrum and Temp+ ablations (Table 5) show internal structure, but they do not close the missing same-budget external baseline. Without that, the title claim is only partially secured.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes Goal-Driven Data Optimization (GDO) for multimodal instruction tuning. It extracts six probe-derived sample descriptors (flow, VDS, temporal necessity, self-consistency, PPL-like difficulty, coverage), ranks candidates with a shared hand-weighted scorer (Eqs. 9–11), and builds 1× subsets under goal-specific feasibility presets (MinLoss, Diverse, Temp, Temp+). Under a fixed one-epoch Qwen3-VL-8B-Instruct SFT recipe, the optimized subsets are compared to a fixed 512k Uni-10x uniform baseline on MVBench, VideoMME, MLVU, and LVBench. The headline result is that GDO reaches the Uni-10x reference after 26.6k–35.4k samples while improving Accuracy by +0.84 to +3.08 pp, with stronger temporal presets improving long-video behavior. Subtask heatmaps, trajectory plots, and Temp+ ablations support a structured, capability-dependent effect rather than a uniform score lift.","tokens_in":18238,"tokens_out":1541,"duration_ms":16761,"significance":"If the result holds under stronger same-budget controls, the paper makes a useful systems contribution: it treats multimodal SFT data allocation as a first-class design lever under a deliberately strict train/eval contract (fixed model, optimizer, checkpoints, benchmarks; benchmark-blind construction). The released code, four-goal spectrum, peak-match framing, and honest LVBench mismatch discussion are concrete strengths. The work is timely for video-centric post-training, where mixed image–video pools are large and unevenly useful. Significance is currently limited by the missing same-budget selection baselines; with those, the paper would more cleanly separate “less data helps” from “this goal-driven scorer/feasibility design helps.”","major_comments":[{"comment":"Table 1 / Fig. 1 / §3.2: The central less-data and peak-match claims compare optimized 1× subsets only to a fixed 512k Uni-10x uniform baseline. This confounds (i) discarding most of the pool with (ii) the six-descriptor scorer and goal presets. Without same-budget 1× controls—at minimum random 1× at each Ng, and preferably quality-only, VDS-only, or prior instruction-selection baselines (e.g., LESS/AlpaGasus-style ranking)—the reported 14–19× reductions and +0.84–+3.08 pp gains may largely reflect any non-uniform curation beating 10× uniform rather than GDO’s specific design. This is load-bearing for the title claim and should be added under the same fixed train/eval contract.","section":"Table 1, Fig. 1, §3.2"},{"comment":"§2.3, Eqs. (9)–(11) and Supp. Table 6: The shared scorer uses fixed hand-chosen mixture weights (e.g., 0.35/0.95/0.35 for video; 0.90/0.15 for image; bvid coefficients 0.85/0.9/0.55/0.15), and the four goals mainly change budgets and feasibility floors rather than the scorer. Table 5 ablates score components for Temp+ but does not test sensitivity of the mixture weights or compare against a pure feasibility-only / pure score-only builder at matched Ng. A short sensitivity or alternative-weight experiment is needed to support the claim that the shared scorer plus goal presets—not an accidental coefficient choice—drive the frontier shifts.","section":"§2.3, Eqs. (9)–(11)"},{"comment":"§2.1–2.2 and §3.1: Descriptors (VDS, temporal necessity, self-consistency, PPL) are obtained from a frozen Qwen3-VL-8B-Instruct probe that is the same model family as the SFT target. While construction is benchmark-blind and not circular by construction, the paper should quantify how much the result depends on this same-family oracle—e.g., by recomputing a subset of descriptors with a different frozen probe, or by reporting correlation of VDS/PPL ranks across probes. Without that, generalization of GDO beyond “self-probe the model you will fine-tune” remains an open assumption behind the utility oracle.","section":"§2.2, Eqs. (4)–(7)"}],"minor_comments":[{"comment":"§2.1 Eq. (1) defines a per-goal 10× uniform control U_g = Uniform(D; 10|S_g|), but all reported tables use a single fixed 512k Uni-10x. Clarify whether per-goal 10× controls were run and, if not, why the fixed 512k anchor is preferred.","section":"§2.1, Eq. (1)"},{"comment":"Table 2: MinLoss and Diverse show negative LVBench deltas while Temp/Temp+ are positive. A one-sentence interpretation in the table caption would help readers see the intended goal spectrum without flipping to the discussion.","section":"Table 2"},{"comment":"Fig. 4 heatmap and Table 4: adverse subtasks (e.g., MLVU Ego under MinLoss −15.1 pp) are important; consider marking statistically or practically large regressions more explicitly so redistribution is as visible as gains.","section":"Fig. 4, Table 4"},{"comment":"Notation: m_tnc is defined as T(q) from a probe judgment, but the prompt template and mapping to [0,1] are only sketched. A short appendix box with the exact probe prompt would improve reproducibility beyond the code link.","section":"§2.2, Eq. (5)"},{"comment":"Related Work §4 is comprehensive but dense; a short paragraph explicitly contrasting GDO’s fixed-contract comparison with curriculum/reweighting/active-selection pipelines would sharpen the positioning.","section":"§4"},{"comment":"Minor typos/consistency: “ML VU” / “MLVU” and “L VBench” / “LVBench” spacing varies across abstract, tables, and body; unify.","section":"Abstract / Tables"}],"recommendation":"major_revision","confidential_remarks":"The experimental contract and honesty about LVBench mismatch are above average for this area. The skeptic’s same-budget baseline concern is the main reason I chose major_revision rather than minor_revision; if the authors already have random/quality 1× runs in the code release or logs, this could convert quickly. Scope fit for a solid systems/data-centric multimodal venue is good; novelty is incremental relative to instruction-data selection literature but the fixed multimodal video contract is a real contribution if baselines are completed."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is the comparison contract. They freeze Qwen3-VL-8B, one-epoch SFT, checkpoints, and eval, and only change subset construction. Under that contract, Temp+ reaches the 512k Uni-10x reference at 26.6k–35.4k samples with +0.84 to +3.08 pp, and MinLoss→Temp+ moves long-video behavior in the expected direction. That is a real empirical outcome, not a recipe confounded with a new backbone.\n\nWhat is new is the packaging, not the ingredients. Flow, VDS loss gap, temporal-necessity proxy, self-consistency, PPL-like difficulty, and coverage are known cues. The useful piece is one shared scorer plus goal feasibility presets (budget, video ratio, VDS+ floors, source floors) that let them reallocate capability without retuning the model. Subset construction is benchmark-blind. Trajectories, subtask heatmaps, and Temp+ ablations are coherent: gains cluster on order/motion/temporal perception; LVBench is smaller and they say so because the pool is short-video/image-heavy. Code is promised.\n\nThe soft spot that matters is the baseline the stress-test flags. Uni-10x is a 10× uniform control, not a strong same-budget 1× selector (random 1×, quality-only, VDS-only, LESS-style, etc.). So “less data, faster convergence” partly confounds discarding ~90% of the pool with the specific six-descriptor scorer. Internal profile structure and ablations help, but they do not fully close that gap. Secondary soft spots: hand-fixed mixture weights, single backbone/family as both probe and trainee, no multi-seed error bars. None of those make the fixed-contract result fake; they bound how far you can generalize the title.\n\nThis is for people who actually run multimodal SFT budgets and care about allocation under a fixed recipe. Not a foundational theory paper. I would bring it to reading group for the experimental design and the honest LVBench discussion. It deserves peer review; I would not desk-reject it. Engage if you care about data allocation for video VLMs; treat the 14–19× reduction numbers as upper bounds until same-budget 1× baselines appear.","headline":"Clean fixed-recipe result: goal-driven 1× subsets hit a 512k Uni-10x bar much earlier with small accuracy gains, but the title claim is only half-secured without same-budget 1× selectors.","tokens_in":18818,"tokens_out":624,"would_cite":true,"duration_ms":5628,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Under a fixed multimodal training recipe, goal-driven 1× data subsets reach a 512k uniform baseline with far fewer samples and higher accuracy.","keywords":["Data Optimization","Instruction Tuning","Vision Language Models","multimodal learning","video understanding","sample selection","efficient fine-tuning"],"falsifier":"Train the identical fixed recipe on random or purely difficulty-based 1× subsets of the same sizes as the GDO profiles; if those controls match or exceed GDO’s peak-match sample counts and final accuracy deltas on MVBench, VideoMME, MLVU, and LVBench, the claim that the descriptors and goal presets drive the gains fails.","tokens_in":18645,"feed_emoji":"⚡","tokens_out":1019,"duration_ms":20176,"temperature":0.7,"pith_summary":"Multimodal instruction tuning often wastes compute by spreading a fixed budget evenly over large mixed image–video pools whose samples have very uneven value. This paper introduces Goal-Driven Data Optimization (GDO): every candidate is scored with six probe-derived descriptors, then small optimized training subsets are built for explicit goals such as minimum loss, diversity, or stronger temporal video emphasis. Holding the model, one-epoch recipe, checkpoints, and evaluation fixed, these 1× subsets reach the performance of a fixed 512k-sample uniform baseline after roughly 27k–35k samples on four video-understanding benchmarks, while also finishing higher in accuracy. Gains are largest on subtask-focused temporal benchmarks and grow as the allocation goal stresses temporally informative video; ultra-long-video benchmarks improve more modestly because the training pool is short-video and image dominant. The result matters because it treats data allocation itself as a first-class lever for efficiency and capability once the backbone and recipe are locked.","feed_headline":"30k curated samples beat 512k uniform multimodal training","feed_subtitle":"Goal-driven 1× subsets hit a fixed baseline earlier and finish higher on four video benchmarks.","key_machinery":"Goal-Driven Data Optimization (GDO): six sample descriptors (flow magnitude, video-dependence score, temporal necessity, self-consistency, PPL-like difficulty, coverage) feed one shared fixed scorer; goal-specific feasibility presets then control budget, video ratio, temporal-positive coverage, source floors, and oversampling to build 1× subsets. Only subset construction changes; model, optimizer, and evaluation stay fixed.","core_discovery":"Under one fixed one-epoch Qwen3-VL-8B-Instruct train/eval contract, goal-driven optimized 1× subsets reach the fixed 512k-sample Uni-10x reference after 35.4k samples on MVBench, 26.6k on VideoMME, 27.3k on MLVU, and 34.7k on LVBench, while improving Accuracy by +1.38, +1.67, +3.08, and +0.84 percentage points respectively. Across MinLoss, Diverse, Temp, and Temp+, stronger temporal emphasis yields steadily better long-video understanding behavior.","pith_inferences":["The same descriptor-plus-feasibility pattern could be reused as a pre-filter for other uneven multimodal corpora, including long-context or audio-visual instruction pools.","Adding native ultra-long-video coverage to the pool would likely let Temp-style presets close more of the LVBench gap without changing the scorer.","That one fixed scorer works across goals suggests admissibility constraints, not preference ranking, are the main dial for goal specialization.","If the frozen probe were a weaker or different model family, the utility ranking might degrade, so probe–backbone alignment itself is a testable design variable."],"forward_implications":["Multimodal SFT budgets can be cut by more than 10× relative to uniform 10× sampling while matching or exceeding the same reference accuracy.","Different explicit allocation goals produce distinct efficiency–capability trade-offs under an identical train/eval contract.","Stronger temporal data pressure systematically improves motion, order, and temporal-perception subtasks.","Gains stay smaller when the target regime (ultra-long video) is poorly matched by a short-video/image training pool.","Once backbone and recipe are fixed, data allocation becomes a controllable design lever rather than a fixed background choice."],"fun_headline_variants":["Goal-driven 1× subsets beat 512k Uni-10x after 26–35k samples","GDO reaches Uni-10x reference with ~30k samples and higher accuracy","Six descriptors yield 1× data that converges earlier than 512k baseline","Temporal GDO subsets lift MVBench/MLVU with far fewer training samples","Fixed-protocol GDO: fewer samples, earlier hit, +0.8–3.1 pp on video benches"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The six descriptors from one frozen probe model, mixed by a single hand-fixed shared scoring formula, rank true training utility well enough that only budget and composition presets need to change across goals.","fun_headline_variants_meta":{"raw":{"variants":["Goal-driven 1× subsets beat 512k Uni-10x after 26–35k samples","GDO reaches Uni-10x reference with ~30k samples and higher accuracy","Six descriptors yield 1× data that converges earlier than 512k baseline","Temporal GDO subsets lift MVBench/MLVU with far fewer training samples","Fixed-protocol GDO: fewer samples, earlier hit, +0.8–3.1 pp on video benches"]},"model":"grok-4.5","effort":"low","cost_usd":0.005618,"raw_usage":{"total_tokens":1601,"prompt_tokens":901,"num_sources_used":0,"completion_tokens":120,"cost_in_usd_ticks":56180000,"prompt_tokens_details":{"text_tokens":901,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":580,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":901,"tokens_out":120,"duration_ms":6982,"temperature":1.0,"reasoning_tokens":580,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T22:17:43.715071+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train the identical fixed recipe on random or purely difficulty-based 1× subsets of the same sizes as the GDO profiles; if those controls match or exceed GDO’s peak-match sample counts and final accuracy deltas on MVBench, VideoMME, MLVU, and LVBench, the claim that the descriptors and goal presets drive the gains fails.","supporting_citations":[],"review_version":1}