{"id":"b1f841b6-0c05-44ac-9a10-7a0e1b25448c","arxiv_id":"2506.13681","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A comprehensive reanalysis finds that min-p sampling does not outperform top-p, top-k, or basic sampling once the original data are re-tested and hyperparameter budgets are equalized.","lead":"This paper re-examines the evidence behind min-p sampling and finds that the original claims of superior quality and diversity are not supported by the data. It matters because min-p is widely used and was presented as an ICLR Oral, so a careful check of whether the hype survived statistical and benchmark scrutiny is valuable.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own Table 1 reports a significant min-p quality win over basic at τ=3.0 after Bonferroni correction; dismissing it because absolute scores are lower in that regime contradicts the blanket 'does not improve' claim, since high-temperature creative generation is exactly the regime min-p…","rationale":"I read the paper as an evidence audit of Nguyen et al. (2024), and much of the audit is valuable and well-documented: the omitted basic-sampler scores, the pooled t-test error, the selective reporting of LLM-as-a-Judge win rates, and the retracted community-adoption numbers are all concrete findings supported by data and code. The paper also earns credit for releasing its sweeps, rerunning with standard GSM8K formatting, and reporting the original authors' new human study. The most load-bearing problem I see is not the benchmark-grid fairness concern that the reader emphasized, though that is also real; it is that the paper's own Table 1 contains a statistically significant result that contradicts the abstract's blanket negative claim. Table 1 shows min-p beating basic on quality at τ=3.0 with p=.001 after Bonferroni correction, and Sec. 2.4 dismisses this by pointing to lower absolute scores in that regime. That dismissal is a normative choice about which temperature regime matters. Since the original min-p paper's central promise is improved quality at high temperature, the significant τ=3.0 result is relevant evidence, and the conclusion must be qualified to survive. The Discussion already concedes a possible high-temperature benefit, so the abstract overstates the finding. This is an internal consistency problem, not a disagreement with community consensus. The reader's weakest_assumption correctly identified the benchmark fairness assumption, so my agreement is partial: I agree the benchmark analysis needs scrutiny, but I would prioritize the human-evaluation contradiction as the single most decisive issue. The appropriate resolution is to keep the paper's findings but require the authors to narrow the central claim, which matches the reader's CONDITIONAL verdict; I do not move the verdict.","tokens_in":17991,"tokens_out":9694,"duration_ms":115182,"concrete_test":"Using the publicly posted human evaluation data behind Table 1, construct the full quality-diversity scatter for every sampler at every temperature and diversity setting, and compute the Pareto frontier (maximal quality for each level of diversity, or vice versa). Then determine whether the min-p point that significantly beats basic at τ=3.0 is strictly dominated by some basic or top-p point at τ≤2.0. If every min-p winning point is strictly dominated in both quality and diversity by a baseline point at a lower temperature, the paper's 'no trade-off advantage' conclusion survives in narrowed form. If any min-p winning point is not dominated, the blanket negative claim fails and the abstract and Sec. 1 must be revised to acknowledge a conditional high-temperature benefit. This is a single computation using data already available in the paper's repository.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Sec. 1 and the abstract is that min-p does not improve quality, diversity, or the quality-diversity trade-off relative to basic, top-k, and top-p sampling. The most direct evidence against that claim is in the paper's own Table 1, not in any external benchmark. Of the 12 one-sided paired t-tests on the original human evaluation data, one remains significant after Bonferroni correction: min-p > basic on quality at temperature 3.0 (t=3.13, p=.001). The paper acknowledges this but then neutralizes it in Sec. 2.4 by arguing that whenever min-p wins, it does so under conditions with lower absolute scores, and therefore offers no apparent advantage. That is a value judgment about which operating points matter, not a statistical refutation. The original min-p paper's claimed benefit is specifically in the high-temperature, creative regime, so a relative quality gain at τ=3.0 is exactly the kind of event that would support min-p's claim, not undermine it. The paper's own Discussion concedes that min-p 'can sometimes provide a benefit at higher temperatures,' which directly undercuts the abstract's unconditional negative. Unless the conclusion is explicitly narrowed to 'no Pareto improvement over the best achievable quality-diversity frontier across all temperatures and hyperparameters,' the blanket 'does not improve' statement is internally inconsistent with the paper's own accepted significant result. A second, independent weakness is the benchmark section's reliance on equalizing the number of swept hyperparameters and its sparse, hand-picked grids on GSM8K only; but the Table 1 contradiction alone is sufficient to require a qualified conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript re-examines the evidence presented in Nguyen et al. (2024) for min-p sampling, arguing that the original paper's four lines of evidence—human evaluations, NLP benchmarks, LLM-as-a-Judge evaluations, and community adoption claims—do not support min-p's superiority. The authors re-analyze the original human evaluation data (finding omitted basic-sampler scores and conducting corrected hypothesis tests), add a new human study conducted by the original authors, perform large GSM8K hyperparameter sweeps with equalized hyperparameter counts, investigate the LLM-as-a-Judge methodology and reporting, and document that the GitHub adoption figures were retracted. The paper concludes that min-p does not improve quality, diversity, or the quality-diversity trade-off relative to basic, top-k, and top-p sampling.","tokens_in":18292,"tokens_out":3634,"duration_ms":39663,"significance":"If the conclusions are correct and properly scoped, this is a valuable corrective to a high-profile ICLR 2025 oral paper. The re-analysis of the original human evaluation data is concrete and reproducible: the discovery that one-third of the human scores were omitted, the corrected Bonferroni and IUT tests, and the documentation of the Table 3(b) reporting inconsistency are substantive contributions. The paper also ships public code and W&B sweeps, which strengthens the benchmark portion. The documented retraction of the community adoption numbers is an important scientific-record correction. However, the paper's central claim is overstated relative to its own evidence, so its significance depends on whether the conclusions are narrowed and the internal inconsistency in the treatment of the significant τ=3.0 quality result is resolved. As written, the blanket 'does not improve' claim in the abstract is not supported even by the paper's own Table 1.","major_comments":[{"comment":"The abstract and Section 1 claim, without qualification, that min-p \"does not improve quality or diversity or the trade-off between quality and diversity.\" Yet Table 1 reports a Bonferroni-corrected significant win for min-p over basic on quality at temperature 3.0 (t=3.13, p=.001). The paper's response in Section 2.4—that this advantage occurs only under conditions with lower absolute scores—is a value judgment about which operating points matter, not a statistical refutation of a relative improvement claim. Furthermore, Section 6 itself concedes that \"min-p sampling can sometimes provide a benefit at higher temperatures,\" which directly contradicts the unconditional wording of the abstract. The conclusion should be narrowed, e.g., to \"no Pareto improvement over the best achievable quality-diversity frontier across all temperatures\" or \"no consistent improvement across all settings,\" and the abstract and Section 1 revised accordingly. As written, the central claim is internally inconsistent with the paper's own accepted significant result.","section":"Sec. 2.2/Table 1 and Sec. 6"},{"comment":"The abstract states that \"comprehensively sweeping the original paper's NLP benchmarks reveals min-p does not surpass baselines when controlling for the number of hyperparameters,\" but the manuscript only sweeps GSM8K Chain-of-Thought. The original paper's NLP benchmark evaluations also included GPQA (5-shot), which is not evaluated here. The manuscript acknowledges this limitation (\"we only evaluated GSM8K CoT\") but still draws a benchmark-wide conclusion. The claim should be explicitly restricted to GSM8K CoT, or GPQA must be swept before the broad statement is made.","section":"Sec. 3.1"},{"comment":"The core benchmark argument depends on the fairness criterion of equalizing the number of hyperparameter settings per sampler. This criterion is asserted, not justified. The hyperparameter grids (six values per sampler, \"taken from the original paper; some were lightly edited\") are hand-picked and may not be representative of each sampler's practical tuning range. Moreover, for basic sampling, temperature is the only hyperparameter, so the \"Best-of-N\" subsampling with N ranging up to 100 is not well-defined unless sampling with replacement is explicitly specified; the manuscript does not specify this detail. Without a justification of the fairness criterion or a sensitivity analysis over alternative grids (e.g., controlling for compute or for best score per hyperparameter dimension), the benchmark-based rejection of min-p's superiority is not fully established.","section":"Sec. 3.1, Best-of-N analysis"}],"minor_comments":[{"comment":"The claim that the value 7.80 in the original Table 15 should be 5.80 is specific and testable; given that the paper presents this as a typo in the original authors' work, it would strengthen the report to include the relevant row from the publicly posted data in a supplementary table.","section":"Sec. 2.4"},{"comment":"The phrase \"in our opinion, were preposterous\" is subjective and not needed to make the point that the numbers were unverified; the subsequent analysis of GitHub repositories is sufficient. Consider removing such value-laden wording.","section":"Sec. 5.1"},{"comment":"The manual annotation of qualitative responses is a reasonable exploratory step, but the paper does not report inter-annotator agreement or whether the annotation procedure was pre-registered. A sentence acknowledging this subjective component would increase methodological transparency.","section":"Sec. 2.3"},{"comment":"The section \"What Went Wrong During the ICLR 2025 Review Process?\" goes beyond the paper's stated scientific scope and includes speculative statements about reviewer behavior. While the review-process critique is motivated by the evidence, the presentation would be more measured if the section focused on the documented failures rather than on evaluative comments about the reviewers.","section":"Sec. 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful critique with solid re-analysis of the human data and concrete documentation of the adoption-claim retraction. The main problem is the gap between the strong, unconditional conclusion and the paper's own evidence, particularly the significant τ=3.0 quality result. This can be fixed by carefully restating the claims and adding the missing GPQA sweep or qualifying the benchmark conclusion. The tone in a few places (e.g., \"preposterous,\" the review-process section) may be off-putting to some readers but is not a scientific flaw. If the authors narrow the conclusions, the paper would likely be a valuable contribution to the field."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is the paper to send anyone who still cites min-p as a validated sampler. What is actually new is the forensic work: documenting that a third of the human evaluation scores were omitted, re-running the paired tests correctly, catching the cherry-picked LLM-as-a-judge scores, and showing the GitHub adoption numbers were unsubstantiated. The reanalysis of the original human data is careful, and the matched-budget GSM8K sweep is a lot of honest work; they shipped code and the W&B sweeps. That reproducibility deserves credit.\n\nThe soft spots are real, too, and one is load-bearing for the wording of the conclusion. Their own Table 1 shows min-p beats basic on quality at tau=3.0 even after Bonferroni correction (t=3.13, p=.001). The paper acknowledges it in Section 2.4, then neutralizes it by arguing this happens in a regime where absolute scores are lower, so it offers no apparent advantage for practitioners. That is a value judgment about operating points, not a statistical refutation. The original min-p paper explicitly targets the high-temperature creative regime, so a significant relative quality win there actually supports the original claim. The Discussion concedes min-p 'can sometimes provide a benefit at higher temperatures,' which is hard to square with the abstract's blanket 'does not improve quality, diversity, or the trade-off.' The conclusion should be qualified: no Pareto improvement, no benefit at standard operating points, or no evidence of improvement that matters for typical use. As written, the headline overreaches.\n\nThe benchmark section has a smaller, second weakness: equalizing the number of swept hyperparameters is a defensible fairness criterion, but not the only one, and the grids are hand-picked and only on GSM8K, not GPQA. A different grid or a compute-per-sampler criterion could reorder the samplers. Still, the direction of the finding is probably robust enough for a claim of 'no consistent superiority.'\n\nThe statistical corrections and documentation of missing data are solid; the citation pattern is fine; the few self-citations in the gold-standard list are not load-bearing. The paper is probably right that the original min-p paper's evidence base is much weaker than its reception suggested. It just needs to say exactly that, not 'min-p does not improve anything.'\n\nWho is this for? Anyone using min-p as a default decoder, and anyone who teaches sampling methods. It deserves a serious referee; the core documentation is careful and the main weakness is a fixable overstatement.","headline":"A careful, mostly convincing re-analysis of min-p's evidence base, but the paper's own headline conclusion is stronger than its own significant high-temperature result allows.","tokens_in":18832,"tokens_out":2435,"would_cite":true,"duration_ms":25419,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A re-analysis finds that min-p sampling does not improve output quality or diversity over basic, top-k, or top-p sampling.","keywords":["min-p sampling","language model decoding","sampling methods comparison","human evaluation","hyperparameter sweep","LLM-as-a-judge","community adoption","replication study"],"falsifier":"Run a pre-registered, large-N human evaluation of basic, top-p, top-k, and min-p sampling across the original temperatures and diversity settings, report all scores, and test the 'min-p better in all settings' claim with an Intersection-Union Test; or run a benchmark sweep where min-p is given equal or fewer hyperparameter settings than the baselines on an out-of-distribution task and still beats them by a margin larger than the reported standard errors.","tokens_in":2855,"feed_emoji":"🔍","tokens_out":1423,"duration_ms":69508,"temperature":0.7,"pith_summary":"This paper re-examines the four lines of evidence behind the claim that min-p sampling makes language-model output both better and more diverse. It re-analyzes the original human evaluations, sweeps the original NLP benchmarks while giving every sampler the same number of tuned hyperparameters, inspects the LLM-as-a-judge comparisons, and checks the community-adoption numbers. Across all four, the paper finds the evidence does not support min-p's superiority: under matched tuning, min-p is largely indistinguishable from basic, top-k, and top-p sampling. If this conclusion holds, a widely adopted sampling method is being used on the strength of claims that closer analysis does not sustain.","feed_headline":"Min-p sampling shows no edge over standard samplers","feed_subtitle":"Re-checking human scores, benchmarks, judge evals, and adoption claims finds the claims don't hold up.","key_machinery":"The argument's technical load is carried by two re-analysis tools. First, a re-running of the original human-evaluation scores with one-sided paired t-tests, a Bonferroni correction, and an Intersection-Union Test, which converts the original 'consistently better' claim into a testable hypothesis about all twelve comparisons. Second, a Best-of-N hyperparameter sweep on GSM8K chain-of-thought: nine models, thirty-one temperatures, three seeds, and six hyperparameter values per sampler, compared by subsampling equal numbers of hyperparameters per sampler so that each method receives an identical tuning budget. The sweep answers the question 'best performance at equal tuning cost' rather than 'best performance when each paper chooses its own grid.'","core_discovery":"The central claim is that min-p sampling does not improve quality, diversity, or the quality-diversity trade-off relative to commonly used samplers. The paper reaches this by re-analyzing the original paper's own data: the human-evaluation results lose significance when the omitted basic-sampler scores are restored and the tests are run correctly (only 1 of 12 comparisons survives Bonferroni correction); the GSM8K benchmark sweep shows min-p matching or underperforming other samplers once each sampler receives the same number of hyperparameter settings; the LLM-as-a-judge results are under-specified, and the reported table chose the higher win rate for min-p and the lower for top-p; and the advertised adoption figures were retracted as unsubstantiated. The paper's conclusion is negative: the published evidence fails to show that min-p outperforms the alternatives.","pith_inferences":["The equal-hyperparameter-volume standard, if applied more broadly, could be a template for re-evaluating other decoding-method claims.","A reasonable next test is whether min-p's possible benefit at very high temperatures (where all samplers do worse) is worth the cost in absolute performance; the paper leaves that as a possible niche.","The pattern of reviewers being swayed by unverified adoption statistics suggests that venue-level practices for checking empirical claims deserve scrutiny.","One could extend the Best-of-N sweep to other tasks (e.g., GPQA, summarization, creative writing) to see whether the null result generalizes."],"forward_implications":["Practitioners should not assume min-p is a strict upgrade; at equal tuning budget, basic, top-k, and top-p samplers perform about as well.","Claims of community adoption should be verified at the repository level before being treated as evidence of a method's value.","Human-evaluation studies in this area need to report all participants' scores, pre-specify their statistical tests, and account for multiple comparisons.","LLM-as-a-judge comparisons should either compare samplers directly or justify why an indirect comparison is valid, given that judge preferences may not be transitive.","Sampling-method choice may be a smaller lever than model choice or prompt format for output quality."],"supporting_citations":[{"why":"The paper that introduced min-p sampling; its data, tables, and claims are the object of this re-analysis.","marker":"Nguyen et al. (2024)"},{"why":"Provides the GSM8K chain-of-thought benchmark used in the hyperparameter sweeps.","marker":"Cobbe et al. (2021)"},{"why":"The GPQA benchmark from the original paper's evaluation, mentioned for consistency.","marker":"Rein et al. (2023)"},{"why":"AlpacaEval, the LLM-as-a-judge evaluation suite whose win-rate reporting is criticized.","marker":"Dubois et al. (2023)"},{"why":"Introduced the LLM-as-a-judge methodology that the original paper's evaluations rely on.","marker":"Zheng et al. (2023)"},{"why":"Supports the concern that LLM-judge preferences may be non-transitive, undermining the indirect comparison.","marker":"Xu et al. (2025)"},{"why":"The LM Evaluation Harness used to run the GSM8K sweeps.","marker":"Gao et al. (2021)"},{"why":"Defines top-p sampling, one of the baseline samplers min-p is compared against.","marker":"Holtzman et al. (2020)"}],"fun_headline_variants":["Min-p sampling fails to beat simple baselines","Study finds min-p sampling offers no quality boost","Min-p sampling's claimed edge unravels under scrutiny","Reanalysis: min-p sampling not superior to standard samplers","Min-p sampling claims overstated, new analysis finds"],"cache_read_input_tokens":20864,"weakest_assumption_plain":"The benchmark comparison assumes that giving each sampler the same number of swept hyperparameter values is the fair way to compare them, and that the chosen grids (six values per sampler) represent each method fairly; a different fairness rule could change the ordering.","fun_headline_variants_meta":{"raw":{"variants":["Min-p sampling fails to beat simple baselines","Study finds min-p sampling offers no quality boost","Min-p sampling's claimed edge unravels under scrutiny","Reanalysis: min-p sampling not superior to standard samplers","Min-p sampling claims overstated, new analysis finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1358,"prompt_tokens":1046,"completion_tokens":312,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":662,"completion_tokens_details":{"reasoning_tokens":236}},"tokens_in":662,"tokens_out":312,"duration_ms":3388,"temperature":1.0,"reasoning_tokens":236,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:27:08.117424+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a pre-registered, large-N human evaluation of basic, top-p, top-k, and min-p sampling across the original temperatures and diversity settings, report all scores, and test the 'min-p better in all settings' claim with an Intersection-Union Test; or run a benchmark sweep where min-p is given equal or fewer hyperparameter settings than the baselines on an out-of-distribution task and still beats them by a margin larger than the reported standard errors.","supporting_citations":[],"review_version":1}