{"id":"5036ac58-8fbf-4685-b8cb-162291425dd0","arxiv_id":"2606.02830","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Proposes memorization-guided two-stage scoring to select debiased training subsets, enabling ERM models to achieve better performance than SOTA debiasing techniques using only 10% of data.","lead":"The paper introduces a two-stage sample scoring function based on memorization dynamics to select a small subset of training data that reduces reliance on spurious correlations without needing group labels. A basic model trained on this 10% subset reportedly outperforms existing debiasing methods.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Two-stage scoring function's claimed disentanglement of core vs. spurious dynamics rests on unverified separation of learning trajectories.","rationale":"Reader's weakest_assumption directly identifies the same unverified separation step that the performance claim depends on; full-text access does not alter this because the abstract already flags the dependence issue and no independent verification (e.g., theoretical bound or controlled ablation) is described.","tokens_in":1700,"tokens_out":305,"duration_ms":12052,"concrete_test":"On ColoredMNIST or Waterbirds with known group labels, compute the two-stage scores; measure correlation between stage-1 and stage-2 rankings versus the spurious attribute; if rank correlation exceeds 0.4 or if the selected subset's spurious-group balance deviates >15% from uniform, the disentanglement claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the proposed two-stage metric (first stage on early training dynamics, second on later) isolates difficulty attributable to core features rather than spurious ones. The abstract asserts existing scores \"largely depend on spurious features,\" yet the argument provides no formal criterion, identifiability condition, or ablation showing that the two stages actually achieve separation (e.g., no proof that stage-1 scores are invariant to spurious correlation strength). If the stages remain coupled through shared optimization trajectory, the selected 10% subset could still be biased toward majority spurious patterns, undermining the ERM superiority result.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that existing sample scoring functions in coreset/invariant subset selection largely depend on spurious features. It proposes a two-stage memorization-guided scoring function that disentangles core-feature and spurious-feature learning dynamics, uses this to select an informative 10% subset (with and without spurious correlations), and shows that standard ERM trained on this subset outperforms state-of-the-art debiasing methods.","tokens_in":1804,"tokens_out":506,"duration_ms":12863,"significance":"If the two-stage metric reliably isolates core-feature difficulty, the result would be significant: it offers a label-free route to data-efficient debiasing that reduces training data to 10% while beating specialized debiasing algorithms. The approach also supplies a concrete, falsifiable test of whether early vs. late training dynamics can be used to separate core and spurious signals.","major_comments":[{"comment":"§3 (two-stage scoring function): the central claim that the first stage (early dynamics) and second stage (later dynamics) disentangle core vs. spurious difficulty lacks an identifiability argument or ablation. No formal criterion is given showing that stage-1 scores remain invariant when the strength of the spurious correlation is varied while core features are held fixed; without this, the selected 10% subset could still be dominated by majority spurious patterns.","section":"§3"},{"comment":"§4–5 (experimental validation): the superiority of ERM on the selected subset over SOTA debiasing baselines is reported, but the manuscript provides no controlled experiment that varies only the spurious-correlation strength while measuring whether the two-stage scores correctly up-weight minority core samples. The 10% data-sufficiency claim therefore rests on the unverified separation assumption.","section":"§4–5"}],"minor_comments":[{"comment":"Notation for the two-stage metric (Eq. (3) or equivalent) should explicitly define the early-training window and the memorization threshold used in each stage.","section":"§3"},{"comment":"The abstract states that existing scoring functions 'largely depend on spurious features'; this should be supported by a quantitative comparison (e.g., correlation of each baseline score with spurious vs. core labels) rather than left as a qualitative assertion.","section":"Abstract / §2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their insightful comments on our manuscript. We address the major comments point by point below, indicating where we will make revisions to strengthen the paper.","responses":[{"response":"We acknowledge that our manuscript does not include a formal identifiability argument or a specific ablation varying spurious correlation strength while holding core features fixed. The two-stage approach is motivated by established observations in the memorization literature that models learn spurious features faster than core features in the presence of strong correlations. In the revised manuscript, we will add an ablation study on synthetic datasets where we systematically vary the spurious correlation strength and demonstrate that the stage-1 scores prioritize samples based on core feature difficulty, leading to subsets that improve minority group performance.","revision_made":"yes","referee_comment":"[§3] §3 (two-stage scoring function): the central claim that the first stage (early dynamics) and second stage (later dynamics) disentangle core vs. spurious difficulty lacks an identifiability argument or ablation. No formal criterion is given showing that stage-1 scores remain invariant when the strength of the spurious correlation is varied while core features are held fixed; without this, the selected 10% subset could still be dominated by majority spurious patterns."},{"response":"The experiments in the manuscript are performed on several benchmark datasets that exhibit different levels of spurious correlations, and we consistently observe that ERM on the 10% subset outperforms debiasing methods. However, we agree that a more controlled experiment isolating the effect of spurious correlation strength would provide stronger validation. We will include such an experiment in the revision using a controlled synthetic dataset to explicitly measure how the two-stage scores up-weight minority core samples as spurious strength varies.","revision_made":"yes","referee_comment":"[§4–5] §4–5 (experimental validation): the superiority of ERM on the selected subset over SOTA debiasing baselines is reported, but the manuscript provides no controlled experiment that varies only the spurious-correlation strength while measuring whether the two-stage scores correctly up-weight minority core samples. The 10% data-sufficiency claim therefore rests on the unverified separation assumption."}],"tokens_in":1363,"tokens_out":466,"duration_ms":23119,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main move is a two-stage sample scoring function that uses memorization dynamics to pick a small subset of training data. The first stage looks at early training and the second at later stages, then the algorithm prioritizes samples both with and without spurious correlations. A standard ERM model on that subset is said to beat existing debiasing methods while using only 10% of the data and no group labels.\n\nWhat is actually new is the explicit attempt to make the scoring function ignore spurious features by splitting the difficulty assessment across training phases. Prior scoring methods are called out for latching onto the spurious signal instead, which is a fair observation in this literature.\n\nThe soft spot is that the abstract supplies no equations for the two-stage metric, no ablation showing the stages are independent of spurious correlation strength, and no experimental details at all. The stress-test concern holds: if the stages remain coupled through the shared optimization path, the selected 10% could still favor majority spurious patterns, and nothing here rules that out. The performance claim therefore sits on an unverified assumption.\n\nThis is aimed at people working on data selection and robustness when group annotations are missing. A reader already thinking about memorization effects or coreset methods might pick up a usable heuristic idea.\n\nSend it to peer review so the full experiments and any checks on the separation can be examined, but the current write-up does not yet make the central claim convincing.","headline":"The two-stage memorization scoring claims to isolate core-feature difficulty from spurious ones for cheap debiased subsets, but the abstract gives no evidence the stages actually separate them.","tokens_in":2274,"tokens_out":372,"would_cite":false,"duration_ms":17343,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A two-stage scoring function selects small data subsets that let standard models outperform specialized debiasing methods on spurious correlations.","keywords":["spurious correlations","dataset debiasing","sample selection","empirical risk minimization","memorization","core features","invariant learning","subset selection"],"falsifier":"Retraining a standard model on the selected 10% subset yields lower accuracy on minority-group test samples than full-data training or competing debiasing methods.","tokens_in":2592,"feed_emoji":"🔍","tokens_out":588,"duration_ms":14819,"temperature":0.7,"pith_summary":"The paper addresses the problem that real-world datasets often contain spurious correlations unrelated to the target label, causing models to misclassify minority samples that lack those patterns. Existing sample scoring methods used for subset selection tend to rely on the spurious features themselves and therefore fail to identify the truly important samples. The authors develop a two-stage scoring approach that tracks the learning dynamics of core features separately from spurious ones and uses the resulting metric to prioritize informative samples both with and without the spurious patterns. Experiments show that a standard empirical risk minimization model trained on the selected subset beats current state-of-the-art debiasing techniques while using as little as 10 percent of the original training data and without requiring group labels.","feed_headline":"10% data subset lets standard models beat debiasing techniques","feed_subtitle":"Two-stage scoring selects informative samples without group labels and yields higher accuracy than specialized methods on spurious-correlati","key_machinery":"two-stage sample scoring function that evaluates difficulty of core features separately from spurious features","core_discovery":"We propose a two-stage sample scoring function that disentangles the learning dynamics of core and spurious features and evaluates their difficulty separately. Based on our proposed metric, we introduce a new algorithm to find and prioritize informative samples both with and without spurious correlations. A standard ERM model trained on our selected samples achieves superior performance compared to state-of-the-art debiasing techniques, while requiring as little as 10% of the original training data.","pith_inferences":["The selection procedure could lower the cost of training on large real-world datasets that contain hidden biases.","The same scoring idea might extend to other label-free data pruning tasks beyond spurious-correlation mitigation.","Testing the method on datasets where spurious features are harder to isolate would reveal the limits of the two-stage separation.","Pairing the selected subset with lightweight regularization could produce further gains on minority samples."],"forward_implications":["A standard ERM model on the selected subset outperforms state-of-the-art debiasing techniques.","The required training set can be reduced to 10% of the original data.","Sample selection succeeds without access to group labels.","The method prioritizes informative samples both inside and outside the spurious-correlation majority group.","Existing scoring functions are shown to depend on spurious features and therefore mis-rank sample importance."],"fun_headline_variants":["Two-stage scoring selects informative samples with 10% data","Disentangling core and spurious features for targeted selection","New metric prioritizes samples both with and without correlations","10% data via separated feature difficulty for ERM models"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The two-stage scoring function can disentangle the learning dynamics of core features from those of spurious features in a way that existing scoring functions cannot.","fun_headline_variants_meta":{"raw":{"variants":["Two-stage scoring selects informative samples with 10% data","Disentangling core and spurious features for targeted selection","New metric prioritizes samples both with and without correlations","10% data via separated feature difficulty for ERM models"]},"model":"grok-4.3","cost_usd":0.005749,"raw_usage":{"total_tokens":2736,"prompt_tokens":658,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":57487000,"prompt_tokens_details":{"text_tokens":658,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2015,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":658,"tokens_out":63,"duration_ms":12945,"temperature":1.0,"reasoning_tokens":2015,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T15:45:06.893818+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Retraining a standard model on the selected 10% subset yields lower accuracy on minority-group test samples than full-data training or competing debiasing methods.","supporting_citations":[],"review_version":1}