{"id":"44eebb42-1467-47c9-9061-9d598c6013f1","arxiv_id":"2606.00628","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"DASD dynamically selects tokens in self-distillation to keep logical corrections while suppressing stylistic noise, improving robustness on math, code, and commonsense benchmarks.","lead":"The paper proposes Distribution-Aligned Self-Distillation (DASD) to dynamically filter high-perplexity tokens during self-distillation using base model confidence. A smart generalist might read it for insights into reducing stylistic biases when training AI models on reasoning tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader's weakest_assumption directly identifies the separability premise as the load-bearing point. With only the abstract available in the provided context, no additional technical flaw (e.g., normalization error, unstated assumption in an equation) can be located. Verdict therefore remains UNVERDICTED.","tokens_in":1678,"tokens_out":261,"duration_ms":10299,"concrete_test":"Re-run the main experiments (math/code/commonsense benchmarks) with the dynamic filter replaced by a random high-PPL token mask of identical cardinality; if the performance gap disappears, the separability claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the premise that high-PPL tokens can be partitioned into beneficial logical corrections versus harmful stylistic drift, with the answer-aware reference model plus base-model confidence serving as a reliable separator. The abstract describes the mechanism but supplies no internal inconsistency or hidden assumption that would falsify the construction on its own terms; the separation is presented as an empirical design choice rather than a formally derived necessity. Absent the full manuscript, no equation, ablation, or derivation is available to expose a circularity or unstated bound that would break the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces Distribution-Aligned Self-Distillation (DASD), which identifies high-perplexity tokens in self-distillation data as arising from either beneficial logical corrections or harmful stylistic drift. It proposes using an answer-aware reference model to generate candidates and dynamically filtering them via the base model's confidence to retain useful knowledge tokens while suppressing misaligned style noise. The central claim is that this yields consistent outperformance over baselines on math, code, and commonsense reasoning benchmarks, reduces high-PPL tokens, and improves robustness across task difficulties.","tokens_in":1749,"tokens_out":470,"duration_ms":19257,"significance":"If the empirical claims hold with proper controls and ablations, the work could offer a practical refinement to self-distillation pipelines by mitigating distribution shift from stylistic imitation. The distinction between sources of high-PPL tokens is a reasonable empirical observation, and the dynamic selection approach is a targeted intervention. No machine-checked proofs or parameter-free derivations are present; credit is due for framing the problem around token-level distribution alignment rather than global loss terms.","major_comments":[{"comment":"Abstract: The assertion that DASD 'consistently outperforms competitive baselines' and 'reduces high-PPL tokens' supplies no quantitative results, error bars, dataset names, baseline identities, or effect sizes. This absence is load-bearing for the central empirical claim and prevents evaluation of whether the token-selection mechanism delivers the stated gains.","section":"Abstract"},{"comment":"Method description (inferred from abstract and introduction): The separation of high-PPL tokens into 'beneficial knowledge-enhancing logical corrections' versus 'harmful stylistic drift' is presented as an empirical design choice without an explicit algorithm, threshold formula, or ablation showing that the answer-aware reference model plus base-model confidence reliably partitions the two sources. This directly affects the validity of the dynamic filtering step.","section":"Method"}],"minor_comments":[{"comment":"Abstract: Consider including one sentence with the specific benchmarks (e.g., GSM8K, HumanEval) and at least one numeric improvement to allow readers to gauge the scale of the reported gains.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address each major comment below and indicate where revisions will strengthen the manuscript.","responses":[{"response":"We agree that the abstract is too high-level. The revised version will incorporate specific quantitative highlights drawn from the experimental results, including approximate gains on named benchmarks (MATH, GSM8K, HumanEval, etc.), the primary baselines, and the observed reduction in high-PPL tokens, along with a brief note on robustness across difficulty levels.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The assertion that DASD 'consistently outperforms competitive baselines' and 'reduces high-PPL tokens' supplies no quantitative results, error bars, dataset names, baseline identities, or effect sizes. This absence is load-bearing for the central empirical claim and prevents evaluation of whether the token-selection mechanism delivers the stated gains."},{"response":"The full method section already specifies the answer-aware reference model for candidate generation and dynamic filtering via base-model token confidence. To make the partitioning criterion fully explicit, we will add a formal algorithmic description (including the exact selection rule) and an ablation isolating the effect of this confidence-based filter. This addresses the concern about clarity without altering the core approach.","revision_made":"yes","referee_comment":"[Method] Method description (inferred from abstract and introduction): The separation of high-PPL tokens into 'beneficial knowledge-enhancing logical corrections' versus 'harmful stylistic drift' is presented as an empirical design choice without an explicit algorithm, threshold formula, or ablation showing that the answer-aware reference model plus base-model confidence reliably partitions the two sources. This directly affects the validity of the dynamic filtering step."}],"tokens_in":1358,"tokens_out":379,"duration_ms":18845,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The central move is to treat high-perplexity tokens in rewritten reference answers as coming from two sources: useful logical fixes that the base model should absorb, and stylistic drift that pulls the model off its original distribution. DASD uses an answer-aware reference model to propose tokens and then keeps only those where the base model already has decent confidence.\n\nThis addresses a real issue in self-distillation for reasoning. Standard rewriting often makes the student copy surface patterns instead of internal logic, and the paper shows why that hurts more on harder problems. The dynamic selection is a direct response to that observation, and the claim that it reduces high-PPL tokens while improving robustness on math, code, and commonsense tasks follows from the setup.\n\nThe experiments are described only at the level of consistent outperformance and token reduction, with no numbers, baselines, or ablations visible here. That leaves the strength of the gains unclear. The separation between the two kinds of high-PPL tokens also rests on the reference model and confidence scores doing the job reliably; if that split is noisy in practice, the method could just be discarding useful signal along with the noise.\n\nThe work is aimed at people already running self-distillation loops on reasoning models and looking for small efficiency or robustness tweaks. A reader in that niche would see a concrete implementation choice worth testing.\n\nI would send it to review. The idea is simple enough to check quickly and the problem it targets is common enough that even modest gains would be useful if the results hold.","headline":"The paper gives a practical filter for high-PPL tokens in self-distillation, keeping logical corrections while dropping style noise.","tokens_in":2208,"tokens_out":378,"would_cite":false,"duration_ms":15782,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Dynamic token filtering in self-distillation preserves logical knowledge while suppressing stylistic noise to improve reasoning.","keywords":["self-distillation","dynamic token selection","reasoning benchmarks","high-perplexity tokens","distribution alignment","logical corrections","stylistic drift","robustness"],"falsifier":"A controlled experiment in which DASD-trained models show no improvement over baselines on difficult reasoning tasks or fail to reduce stylistic imitation in generated outputs would falsify the claim.","tokens_in":2570,"feed_emoji":"","tokens_out":612,"duration_ms":19774,"temperature":0.7,"pith_summary":"Self-distillation rewrites reference answers to better match the model's distribution but introduces stylistic biases that cause imitation of surface forms rather than reasoning patterns. High-perplexity tokens in the data come from two sources: beneficial logical corrections and harmful stylistic drift. DASD generates candidate tokens with an answer-aware reference model and dynamically filters them according to the base model's confidence scores. This keeps tokens that carry useful logical knowledge and discards distributionally misaligned style noise. The method yields consistent gains over baselines on math, code, and commonsense reasoning benchmarks while reducing disruptive high-PPL tokens.","feed_headline":"Dynamic filtering keeps logical tokens in self-distillation","feed_subtitle":"By retaining beneficial corrections and dropping stylistic noise, the method improves math, code and commonsense reasoning performance.","key_machinery":"Dynamic token selection in DASD that filters high-perplexity tokens by combining answer-aware reference generation with base-model confidence to align training data with the original distribution.","core_discovery":"Distribution-Aligned Self-Distillation (DASD) uses an answer-aware reference model to generate candidate tokens and applies dynamic selection based on the base model's confidence, thereby preserving tokens that encode useful logical knowledge while suppressing tokens that represent distributionally misaligned style noise.","pith_inferences":["The same separation of logical versus stylistic tokens could be tested in standard knowledge distillation beyond the self-distillation setting.","If the filtering reliably isolates logical content, it may reduce certain forms of output bias on tasks outside reasoning benchmarks.","Repeating the token-selection process on models of different sizes would test whether the two sources of high-perplexity tokens remain separable at scale."],"forward_implications":["Consistent outperformance on math, code, and commonsense reasoning benchmarks compared with competitive baselines.","Measurable reduction in high-perplexity tokens present in the rewritten training data.","Improved robustness on tasks of varying difficulty without disrupting the base model's original distribution.","Better preservation of useful reasoning patterns instead of surface-form imitation."],"fun_headline_variants":["Dynamic selection keeps logical tokens aligned","Confidence filtering preserves logic in DASD","Answer-aware model filters stylistic noise","Token selection aligns self-distillation distribution"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"High-perplexity tokens arise from two cleanly separable sources—beneficial logical corrections versus harmful stylistic drift—that an answer-aware reference model and base-model confidence can reliably distinguish.","fun_headline_variants_meta":{"raw":{"variants":["Dynamic selection keeps logical tokens aligned","Confidence filtering preserves logic in DASD","Answer-aware model filters stylistic noise","Token selection aligns self-distillation distribution"]},"model":"grok-4.3","cost_usd":0.003889,"raw_usage":{"total_tokens":1963,"prompt_tokens":600,"num_sources_used":0,"completion_tokens":47,"cost_in_usd_ticks":38887000,"prompt_tokens_details":{"text_tokens":600,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1316,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":600,"tokens_out":47,"duration_ms":9305,"temperature":1.0,"reasoning_tokens":1316,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T19:10:55.191283+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled experiment in which DASD-trained models show no improvement over baselines on difficult reasoning tasks or fail to reduce stylistic imitation in generated outputs would falsify the claim.","supporting_citations":[],"review_version":1}