{"id":"c88cef3c-b3bc-4bf4-9409-9a7423135cad","arxiv_id":"2601.10922","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"For multimodal reasoning fine-tuning under a fixed protocol, difficulty-filtered small datasets on an aligned source outperform larger or more diverse alternatives, with diversity and synthetic mixtures adding no gains.","lead":"This paper identifies which data-curation choices actually matter when fine-tuning a vision-language model for reasoning under a fixed training recipe. It finds that difficulty-filtered small datasets on a well-aligned source outperform larger or more diverse datasets, and that common diversity and synthetic-data heuristics add little.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 3's difficulty bins are overlapping/inconsistent: moderate (0≤k≤8) is a superset of super-difficult (k≤3), and the asserted random-subsample comparison is not shown, so the central difficulty-filtering claim is not cleanly supported.","rationale":"I read the paper as a scoped empirical study of data curation under a fixed protocol. The reader's weakest assumption — that the fixed training recipe may not be representative for diverse/synthetic data — is a legitimate external-validity caveat, but the paper explicitly restricts its conclusions to the DCVLR regime and acknowledges this limitation in Section 8; it does not undermine the internal claim. The more pressing issue is internal: the evidence for the main positive result is not cleanly displayed. Table 3's difficulty ranges overlap under the most natural parsing, and the claimed random-subsample superiority is asserted without a shown baseline. That makes the central 'difficulty filtering dominates' claim less secure than the surrounding text suggests. A disjoint-bin rerun with an explicit random baseline would settle whether moderate difficulty is genuinely the driver. I do not see grounds to reject the paper; the appropriate verdict remains conditional on this clarification. The reader's fixed-recipe concern and my Table 3 concern are different, hence disagreement.","tokens_in":12011,"tokens_out":7601,"duration_ms":75183,"concrete_test":"Rerun the Table 3 ablation with disjoint bins k∈[0,2], [3,5], [6,8], [9,14], [15,16] (correct of 16), three seeds each, 1k examples, plus a 1k random Walton subsample as an explicit baseline. Report means and per-seed values. If the middle disjoint bin does not significantly exceed the random subsample and adjacent bins, the 'moderate difficulty is dominant' conclusion fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim — that moderate-difficulty examples drive gains and that difficulty-filtered Walton beats random subsampling — rests on Table 3 and the prose in §6.2. Table 3 defines super-difficult as k≤3 and moderately difficult as 0≤k≤8 (correct out of 16, higher k = easier). The moderate bucket therefore contains the super-difficult bucket; it is not a disjoint 'fails frequently but not uniformly' interval. The comparison cannot attribute the 1k-example gain to moderate difficulty as distinct from easy/super-difficult; the k=4–8 slice is never isolated. Moreover, §6.2 states that difficulty-filtered subsets outperform randomly subsampled Walton of comparable size, but no random 1k baseline appears in Table 3 or elsewhere in the body. The only matched-size comparisons are among difficulty thresholds. If the 'k≤3' reading is wrong and 'k≤30' is literal, that is impossible for k out of 16, so either way the threshold definitions are not reproducible as written. This is more load-bearing than the fixed-recipe concern because it targets the positive evidence, not just the scope of negative results.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports the first-place solution to the NeurIPS 2025 DCVLR challenge and uses post-competition ablations under the fixed official training recipe to argue that (i) model-relative difficulty filtering of an aligned base corpus (Walton) is the dominant curation lever; (ii) increasing data size beyond roughly 1k examples does not improve mean accuracy but mainly reduces run-to-run variance; and (iii) diversity heuristics and synthetic CoSyn mixtures do not improve over difficulty-filtered Walton. The authors repeatedly state that the conclusions are scoped to the DCVLR fixed-recipe, saturation regime.","tokens_in":12308,"tokens_out":6271,"duration_ms":69980,"significance":"If the findings are supported, the paper would be a useful controlled empirical contribution: it reframes a competition result, uses the official training and evaluation pipeline, and includes three-seed replication for the difficulty-threshold ablation. The explicit negative results for diversity and synthetic augmentation are also valuable, even though they are regime-specific. However, the central positive evidence has definitional problems and lacks the random-baseline comparison needed for the paper's main claim, while the abstract contains claims not present in the body. The paper is potentially publishable after targeted major revisions.","major_comments":[{"comment":"The difficulty bins are not disjoint and the asserted random baseline is missing. Table 3 defines 'Super-difficult' as k≤3 and 'Moderately difficult' as 0≤k≤8, with k = number of correct answers out of 16 stochastic decoding passes (higher k = easier). The moderate set is therefore a superset of the super-difficult set, so the comparison cannot isolate examples that 'fail frequently but not uniformly.' The k=4–8 slice is never reported. In addition, §6.2 and §7.1 claim that difficulty-filtered 1k Walton subsets outperform randomly subsampled 1k Walton subsets, but Table 3 contains no random 1k row and §6.3/Figure 4 do not provide this matched comparison. Please report disjoint intervals (e.g., k≤3, 4≤k≤8, 9≤k≤14, k≥15) and add a random 1k baseline with seed-level results.","section":"Table 3; §6.2"},{"comment":"The abstract makes two claims that do not appear in the body: a per-benchmark decomposition attributing much of the improvement over random sampling to OlympiadBench, and transfer of Qwen-derived difficulty scores to other model families. The body has no random-baseline decomposition table for per-benchmark gains and no experiments with other model families. These claims must either be supported by new experiments or removed from the abstract. As written, the abstract overstates the paper's evidence.","section":"Abstract vs. body"},{"comment":"Table 2 reports 'Ours 1k' as Overall (weighted) 46.0, while Table 3 reports the moderately difficult threshold as 0.491. If both numbers refer to the same aggregate evaluation, they are inconsistent. If Table 3 uses a different metric (e.g., unweighted accuracy, a different score, or a different sample), the caption should say so explicitly. Without clarification, the reader cannot tell which number represents the final submission or the difficulty-filtered ablation.","section":"Table 2 vs. Table 3"},{"comment":"The negative results for diversity and CoSyn mixing appear to be single training runs without error bars. Section 5.3 states that only selected ablations, particularly dataset size, are repeated with multiple seeds. Given that the three-seed ranges in Table 3 span roughly 0.02–0.05 in accuracy, single-run point comparisons in Figures 5 and 6 are insufficient to support claims that diversity heuristics 'do not improve' or that CoSyn mixtures 'consistently degrade' performance. Please provide repeated-seed results for these variants, or substantially weaken the wording.","section":"§5.3, §6.4, Figures 5–6"}],"minor_comments":[{"comment":"The caption says 'higher k = easier' but the row labels 'Super-difficult k≤3' and 'Moderately difficult 0≤k≤8' remain ambiguous because the lower bound 0 does not disambiguate a correct-count score from a difficulty score. Use a single unambiguous notation and explicitly state that k is the number of correct rollouts.","section":"Table 3 caption"},{"comment":"The symbol k is used both for the number of rollouts (k=16) and for the number of correct answers in Table 3. Rename one of these to avoid confusion.","section":"§4.2"},{"comment":"The figures are referenced but their axes, error bars, and seed counts are not described in the text. The captions should state whether error bars are standard deviations across seeds, and for which variants seeds were run.","section":"Figures 4–6"},{"comment":"The limitation about the fixed training recipe is acknowledged and is appropriate. However, because this limitation is load-bearing for the negative diversity/synthetic-data conclusions, it should also be flagged in the abstract or introduction rather than only in the Limitations section.","section":"§8"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the competition result and controlled setup give the paper a real empirical contribution, but the body currently does not support several abstract claims and the main difficulty-filtering comparison is not cleanly specified. The needed changes are local: add disjoint difficulty bins and a random baseline, reconcile the aggregate numbers, and either add cross-model/per-benchmark experiments or remove those claims. I do not see a reason to question the integrity of the work; the issues are presentation and missing evidence that can be addressed in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the DCVLR paper. Bottom line: it's a legitimate, scoped empirical study, but the central positive claim is messier than it looks, and the abstract goes beyond what the body supports.\n\nWhat's genuinely new: the controlled comparison of difficulty filtering against diversity heuristics and CoSyn synthetic mixing under the fixed DCVLR training recipe. The negative results—diversity and synthetic data don't help, and larger aligned datasets mainly reduce variance—are useful if you take them as regime-specific. The three-seed replication for the main difficulty ablation is solid.\n\nThe soft spots are real. First, the difficulty buckets in Table 3 overlap: 'moderately difficult' is 0≤k≤8 and 'super-difficult' is k≤3 out of 16, so the moderate bin contains the super-difficult bin. That means the analysis doesn't isolate the 'challenging but learnable' (k=4–8) examples that the prose claims drive the gains. You'd need a disjoint interval to support that. Second, the paper repeatedly claims difficulty-filtered subsets beat random subsampling at matched size, but the random baseline lives only in Figure 4, not in Table 3, so the matched-size comparison is asserted rather than directly shown. Third, the abstract promises a per-benchmark decomposition and transfer to other model families that never appear in the body. Fourth, the diversity and CoSyn negative results appear to be single-run, no error bars, so they're worthwhile hints, not demonstrations.\n\nThe fixed-recipe concern—that negative results might be an artifact of the organizers' hyperparameter sweep—is legitimate and the paper acknowledges it. That's the main limit on generalizing the negative claims. The circularity from using base-model rollouts for difficulty is mild, since held-out benchmarks are included.\n\nWho's this for? People doing data curation for VLM reasoning fine-tuning under tight compute budgets. It's not a field-shifter, but it's a useful data point. With a cleaned-up Table 3, an explicit random-matched comparison, and an abstract that matches the body, it would be a solid contribution. As is, it deserves peer review but needs revision first.\n\nRecommendation: engage with it, but referee it carefully.","headline":"Useful scoped negative results, but the headline claim about moderate difficulty is undercut by overlapping bins and an abstract that outruns the body.","tokens_in":12830,"tokens_out":3953,"would_cite":true,"duration_ms":38727,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Difficulty-filtered example selection, not data scale, drives multimodal reasoning gains under a fixed training protocol.","keywords":["data curation","multimodal reasoning","difficulty filtering","fine-tuning","vision-language models","dataset scaling","diversity heuristics","saturation regime"],"falsifier":"Retrain the same curated subsets under per-composition hyperparameter sweeps (e.g., a longer schedule or different learning rate for the diverse or synthetic mixtures), then check whether any diversity or synthetic mixture beats the difficulty-filtered aligned set at equal evaluation cost; if one does, the paper's negative claims about those heuristics are falsified.","tokens_in":11904,"feed_emoji":"🎯","tokens_out":5651,"duration_ms":62821,"temperature":0.7,"pith_summary":"The paper argues that when the base model, optimizer, schedule, and evaluation are all fixed, the main lever in multimodal reasoning fine-tuning is which examples you choose—not how many. It reports that selecting examples the base model frequently fails on but can sometimes solve (moderate difficulty) from an already aligned corpus produces the largest accuracy gains, and that a 1,000-example filtered set performs comparably to a 10,000-example random baseline while beating random subsets of the same size. It also finds that adding cluster-based diversity or mixing in rewritten synthetic traces does not improve on difficulty filtering, and can hurt. The practical target is a scoped recipe for data-constrained reasoning fine-tuning: start aligned, filter by model-relative difficulty, keep the set small.","feed_headline":"Pick middle-difficulty examples, not more data, for reasoning gains","feed_subtitle":"In a fixed-recipe study, filtered 1k sets beat random same-size sets; size mostly just steadies runs.","key_machinery":"The central object is a per-example difficulty score: the number of correct answers the base model gives on a question over 16 temperature-0.7 stochastic decoding passes, where a higher score means the example is easier. The paper uses this score to cut an aligned starting corpus into easy, moderate, and super-difficult bands, then samples a fixed budget from the moderate band. The mechanism works because the score is model-relative and target-aligned: it identifies examples the fixed base model is on the verge of mastering within the evaluation distribution, rather than examples that are hard in an absolute sense.","core_discovery":"On the paper's own terms, the central discovery is that difficulty-based filtering is the dominant curation signal in a regime of diminishing returns from added data. For each candidate example, the paper scores difficulty by how often the frozen base model answers correctly across 16 stochastic decoding runs; examples that land in the middle band—frequently wrong but not uniformly wrong—yield the strongest downstream accuracy. Filtered 1k subsets beat random 1k subsamples, and the gain is not accounted for by the most heavily weighted aligned benchmark alone: a per-benchmark decomposition shows that much of the improvement over random sampling comes from the largest non-aligned math benchma","pith_inferences":["The negative results for diversity and synthetic data may be tied to the single fixed hyperparameter recipe; under schedules tuned per data composition, diverse or synthetic data could behave differently. This is my inference, not the paper's claim.","The per-benchmark decomposition suggests the headline effect is not purely in-distribution overfitting: if the largest non-aligned math benchmark drives much of the gain, difficulty filtering may be selecting transferable reasoning behaviors—a hypothesis the paper leaves partially open.","A testable extension is to use the same difficulty-filtering procedure across different model families; the paper's own transfer results hint that difficulty scores from one model help some other models more than others, so score transfer could serve as a cheap probe before fine-tuning.","Benchmark designers can exploit this result: aggregate scores in a fixed-recipe challenge mostly measure example selection, so reporting per-benchmark decompositions is necessary to avoid mistaking specialization for general reasoning ability."],"forward_implications":["Under a fixed training recipe, a small (~1k) difficulty-filtered set can match a 10k random set, so data-constrained teams should spend their budget on selection, not volume.","Moderate-difficulty examples—those the model sometimes gets right but often misses—carry most of the learning signal; filtering out both easy and extremely hard examples improves accuracy.","Scaling an aligned dataset beyond roughly 1k examples mostly reduces run-to-run variance; it does not reliably improve mean accuracy and can mildly hurt less-aligned benchmarks.","Diversity heuristics such as clustering-based balancing and category-level balancing do not add to difficulty filtering, and stacking them on top can suppress learning signal.","Synthetic rewritten data mixed at low ratios is neutral, and at higher ratios consistently degrades performance in this fixed-recipe regime."],"fun_headline_variants":["Middle-difficulty picks beat more data for multimodal reasoning","Filter by difficulty, not dataset size, for reasoning FT gains","Reasoning gains from mid-difficulty curation, not added data","Why mid-difficulty examples outpace larger datasets in FT","Curation: choose hard-but-passable examples over bulk data"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The negative conclusions about diversity and synthetic data rest on the assumption that the single fixed training recipe—settings chosen by the organizers on a baseline and applied uniformly to all datasets—interacts with every data type in the same way; if diverse or synthetic data prefer a different schedule, those conclusions could be artifacts of the recipe.","fun_headline_variants_meta":{"raw":{"variants":["Middle-difficulty picks beat more data for multimodal reasoning","Filter by difficulty, not dataset size, for reasoning FT gains","Reasoning gains from mid-difficulty curation, not added data","Why mid-difficulty examples outpace larger datasets in FT","Curation: choose hard-but-passable examples over bulk data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000173,"raw_usage":{"total_tokens":1120,"prompt_tokens":754,"completion_tokens":366,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":280}},"tokens_in":498,"tokens_out":366,"duration_ms":4557,"temperature":1.0,"reasoning_tokens":280,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T10:08:56.659136+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the same curated subsets under per-composition hyperparameter sweeps (e.g., a longer schedule or different learning rate for the diverse or synthetic mixtures), then check whether any diversity or synthetic mixture beats the difficulty-filtered aligned set at equal evaluation cost; if one does, the paper's negative claims about those heuristics are falsified.","supporting_citations":[],"review_version":1}