{"id":"8a162e83-da48-4e71-8a9a-80a09aed6503","arxiv_id":"2508.05131","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For destructive testing, ISO 2859-2 plans do not bound the Bayesian risk of an unsatisfactory remaining lot; new plans and a remaining-lot-size tabulation are provided.","lead":"This paper argues that standard acceptance-sampling plans miss the mark when testing destroys the sampled items, and it offers Bayesian plans that protect the quality of the lot left behind. It also proposes a new way to tabulate such plans, fixing the remaining lot size instead of the sample size.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposed plans and 44% risk bound rest on a uniform prior that is not the reference prior for the hypergeometric model; small prior shifts exceed the 10% limit for small lots.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the plans and the 44% risk bound depend on the uniform prior, and sensitivity analysis shows small changes (a from 1.0 to 1.1) can break the 10% limit for small lots. My reading agrees. The paper is valuable: the critique of ISO 2859-2 for destructive testing is real, the Bayesian framing is coherent, and the new [N−n, Ac] representation is a useful standardization idea. However, the prior is a subjective input, and the paper's claim that uniform is the reference prior is questionable; moreover, the plans' sensitivity for small lots means the advertised 10% risk is not robust. The reader's CONDITIONAL verdict already reflects this. I see no internal inconsistency that would justify rejection, nor a flaw that would demand unverified status. The conditional acceptance is appropriate. The one additional nuance I would stress is that the 'reference prior' label may be factually incorrect for the hypergeometric model, which strengthens the prior-dependence concern: the 44% peak might be an artifact of choosing uniform rather than a true reference prior. The concrete test above would settle this. If the reference prior indeed yields a lower peak, the paper should soften its abstract and qualify the strength of its critique of ISO plans. As it stands, the paper overstates the certainty of the 44% claim and the robustness of the proposed plans for the smallest lots.","tokens_in":11013,"tokens_out":17626,"duration_ms":184522,"concrete_test":"Recompute, for each plan in Table 1 (LQ=2% and LQ=20%), the specific consumer's risk (Eq. 5) over the full stated lot-size range under the beta-binomial prior with (a,b)=(1.1,1) and also (1.1,1.1), using the same computational method ('stats.betabinom.cdf' from SciPy). Determine the maximum risk for each plan. If any plan's maximum risk exceeds 10% for a lot size inside its advertised range, the claim that these plans limit risk to 10% is not robust to this prior misspecification. Additionally, derive the reference prior for the hypergeometric likelihood (3) and recompute the peak risk for ISO plan (50,0) at N=90; if it drops well below 44%, the 'reference prior' justification for the central critique is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central argument has two pillars: (i) ISO 2859-2 plans yield high Bayesian specific consumer's risk (peak 44%) and are thus ill-suited for destructive sampling, and (ii) the newly proposed plans (Table 1) limit this risk to 10%. Both pillars are computed under the uniform prior P(kwhole)=1/N (Section 3, Eq. 6), which the paper calls 'the reference prior' citing Bernardo [2003]. This prior choice is load-bearing. For the hypergeometric likelihood (Eq. 3), the reference prior for the underlying proportion is Beta(1/2,1/2), not uniform; a uniform prior over kwhole overweights intermediate values and underweights the boundary p=0, so the 44% peak and the tabulated plans are not calibrated to a reference prior. More directly, even if uniform is accepted as a defensible non-informative prior, the paper's own sensitivity analysis (Section 4, Fig. 5a-5c) shows that for the smallest lots in each range of Table 1, a slight shift toward nonconformance (e.g., a=1.0 -> 1.1) pushes the specific consumer's risk above the 10% limit. The paper concedes this and suggests adjusting ranges or plans, but the abstract and conclusions state without qualification that the plans 'limit the (Bayesian) specific consumer's risk'. Thus the central claim that the new plans provide 10% consumer protection is conditional on a specific prior, and for small lots it is not robust to small prior perturbations. Similarly, the headline that ISO plans fail to meet 10% protection is prior-dependent: under a prior with more mass near perfect quality (e.g., Beta(1/2,1/2)), the 44% figure would likely be substantially lower, undermining the dramatic claim. This is the softest spot in the argument because it affects both the diagnosis and the remedy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses acceptance sampling by attributes when the sampled items are destroyed, so that the lot remaining after sampling, rather than the original whole lot, is the relevant quantity. It argues that ISO 2859-2 plans, designed to limit the consumer's risk for the whole lot, do not control the risk for the remaining lot. Section 2 notes that the frequentist risk in Eq. (4), which conditions on the number of nonconforming items in the remaining lot, is not directly computable from the hypergeometric distribution because the conditioning event involves the sample count Y. Section 3 defines a Bayesian specific consumer's risk in Eq. (5) under a uniform prior for kwhole, shows that an ISO plan can have a risk of 44% at N=90, and Section 4 proposes tabulated zero-acceptance plans in Table 1 using a new representation [N-n, Ac] that fixes the remaining lot size. A prior sensitivity analysis with beta-binomial priors is included.","tokens_in":11388,"tokens_out":27127,"duration_ms":312916,"significance":"If correct, the paper provides a practical, tabulated alternative to ISO 2859-2 for destructive testing and a simple representation that is efficient for small lots. The 44% example is concrete and checkable and usefully demonstrates that the standard's plans do not control the quality of the remaining lot under a uniform prior. The paper is transparent about restricting attention to Ac=0, about the role of the prior, and about sensitivity to prior perturbations, which strengthens reproducibility. The main weaknesses are framing rather than mathematical: the abstract overstates the frequentist impossibility, and the reference-prior label is supported only by a single citation.","major_comments":[],"minor_comments":[{"comment":"The uniform prior is written as P(kwhole)=1/N, but kwhole takes N+1 values (0,...,N), so the normalized uniform prior is 1/(N+1). The beta-binomial with a=b=1 indeed gives 1/(N+1). Please correct the formula and the corresponding sentence. Also, the notation \"P(kwhole - y | y,N,n)\" is informal; it should be written as P(kwhole = krem + y | y,N,n).","section":"Section 3, Eq. (6)"},{"comment":"The statement that the hypergeometric distribution cannot describe the frequentist consumer's risk is too broad. What is shown correctly is that Eq. (4), which conditions on the random variable krem=kwhole-Y, is not a frequentist probability. For a fixed kwhole, however, the probability of accepting a remaining lot with krem >= m is a well-defined hypergeometric tail: F(min(Ac,kwhole-m); N,kwhole,n). The paper should state this distinction and restrict the claim to Eq. (4), or explain why a worst-case frequentist criterion over kwhole is rejected.","section":"Abstract and Section 2"},{"comment":"The claim that the uniform prior is the reference prior for the hypergeometric sampling distribution is load-bearing for the numerical 44% figure and for Table 1, but it is supported only by a page citation. Please provide a short derivation or exact quotation, or rephrase as \"an assumed noninformative prior.\" It would also help to note explicitly that the hypergeometric parameter here is discrete K, so the continuous binomial reference prior Beta(1/2,1/2) is not directly applicable.","section":"Section 3, prior discussion"},{"comment":"The entries \"no plan\" for N<50 (LQ=2%) and N<17 (LQ=20%) are not explained. For example, [5,0] at N=49 gives a risk of about 10% and [6,0] at N=16 gives about 11% under the uniform prior. Please state the selection criterion that excludes these cases, so that the table is self-contained.","section":"Table 1 and Section 4"},{"comment":"The phrase \"the risks ... increase only at multiples of 1/LQ\" is imprecise. The risk jumps when the threshold ceil(LQ*(N-n)) changes, i.e., when LQ*(N-n) crosses an integer. Please rephrase to avoid confusion.","section":"Section 4, Figure 3"},{"comment":"The table caption should state explicitly that all plans are zero-acceptance (Ac=0) and that the risk limit is one-sided (P(krem >= ceil(LQ*(N-n)) | y,N,n) <= 10%). This is clear in the body text but should be in the caption for standalone use.","section":"Table 1 caption"}],"recommendation":"minor_revision","confidential_remarks":"The stress-test concern about circularity does not land: the plans are decision-theoretic designs under a stated prior, not fitted quantities. The concern about the uniform prior not being the reference prior is also not obviously correct, because the hypergeometric parameter is discrete K rather than a continuous proportion; the paper's citation to Bernardo may well be right. Still, because the 44% headline and Table 1 both depend on this prior, I would ask the authors to provide a short derivation or exact quote so that the claim can be verified. The overstatement about the frequentist risk in the abstract should be corrected, but the central applied contribution is sound and within the scope of the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you care about acceptance sampling in destructive-testing settings. The paper does two useful things: it shows, with a concrete 44% peak, that ISO 2859-2 plans do not control the risk of accepting an unsatisfactory remaining lot, and it proposes a practical tabulation format — fixing the remaining lot size instead of the sample size — that makes Bayesian sampling plans efficient and suitable for standardization. That representation is the real novelty; the Bayesian specific consumer's risk is not new (Uhlig et al., Wilson and Farrow), and Herman and Robbins already noted the hypergeometric overstates confidence in small residual lots. But the systematic design of destructive-sampling plans for remaining lots, and the table format, are new and well executed. The 44% calculation checks out, and the sensitivity analysis is a credit: they explore how prior changes affect the applicable lot-size ranges and honestly report that small prior shifts can push the risk over 10% for the smallest lots.\n\nThe main soft spot is the claim that the frequentist risk 'cannot be described' by the hypergeometric distribution. That is overstated; it can be expressed as a sum of hypergeometric terms, but the conditioning event involves the sample outcome, so it is not identifiable without a prior. The paper's real point is that the standard's risk concept does not address the remaining lot, not that the math is impossible. Second, the plans' 10% guarantee is conditional on the uniform prior, which the paper calls 'the reference prior' citing Bernardo. That label is questionable — the reference prior for the underlying proportion is typically Beta(1/2,1/2) — and the sensitivity analysis shows the guarantee is not robust for small lots. The paper does acknowledge this in Section 5, but the abstract and conclusions are unqualified, which is a bit much. The lack of shipped code is minor; the computations are standard beta-binomial.\n\nWho is this for? Statisticians and metrologists working on standards for acceptance sampling, seed testing, utility metering. It deserves a serious referee. I'd send it to peer review: the central idea is sound, the contribution is practical, and the weaknesses are caveats rather than fatal flaws. A referee should push them to soften the 'cannot be described' claim and to condition the abstract's statement on the prior. The [N-n, Ac] representation is a genuinely useful idea for standardization, and the paper is a fair and honest piece of applied work.","headline":"A genuinely useful applied paper on destructive acceptance sampling, with a real practical contribution in the [N-n, Ac] representation, but it overstates what the hypergeometric distribution cannot do and the 10% guarantee is prior-dependent, especially for small lots.","tokens_in":11917,"tokens_out":5397,"would_cite":false,"duration_ms":58288,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P30","62F15","62C10"],"pacs":[],"model":"deepseek-v4-flash","headline":"ISO 2859-2 sampling plans do not protect the lot that remains after destructive testing: the probability of accepting an unsatisfactory remaining lot can reach 44%, and the paper tabulates replacement plans that hold it to 10%.","keywords":["acceptance sampling by attributes","destructive testing","specific consumer's risk","Bayesian statistics","hypergeometric distribution","remaining lot quality","ISO 2859-2","reference prior"],"falsifier":"Recompute the Bayesian specific consumer's risk (5) for the ISO 2859-2 plan $n=50$, $Ac=0$, $LQ=2\\%$, $N=90$ with the uniform prior: the paper predicts 44%. If an independent calculation, or a simulation of repeated destructive-sampling campaigns following ISO 2859-2, puts that risk at or below 10%, the central claim fails. A second check: compute the risk for the paper's proposed plan $[51,0]$ at the smallest lot in its range, $N=160$, with a $\\beta$-binomial prior $a=1.1$, $b=1$; the paper's sensitivity analysis implies it exceeds 10%.","tokens_in":10892,"feed_emoji":"🧪","tokens_out":17584,"duration_ms":156059,"temperature":0.7,"pith_summary":"Acceptance-sampling standards such as ISO 2859-2 promise to cap the consumer's risk of accepting a bad lot near 10%. The paper argues that this promise breaks when the inspection destroys the sample, as in seed-germination tests or checks on installed utility meters: what matters then is the quality of the lot that remains after the sample is removed, and the hypergeometric distribution cannot express the probability of accepting a bad remaining lot. Applying ISO 2859-2 plans to destructive sampling can push the Bayesian specific consumer's risk — the posterior probability that an accepted remaining lot still fails the quality limit — to 44%. The paper designs accept-zero sampling plans that limit this risk to 10% under a uniform reference prior, and proposes a new tabular form, fixing the remaining lot size $N-n$ rather than the sample size $n$, which keeps the plans near-optimal and easy to standardize. If the paper is right, the standard should not be used for destructive testing, and its future editions should say so explicitly.","feed_headline":"ISO sampling plans can hit 44% consumer risk in destructive testing","feed_subtitle":"ISO 2859-2 guards the whole lot; destructive testing leaves a smaller lot, and new tables limit its risk to 10%.","key_machinery":"Central machinery: the Bayesian specific consumer's risk (5), $P(k_{\\mathrm{rem}} \\ge \\lceil LQ\\cdot(N-n)\\rceil \\mid y \\le Ac)$, the posterior probability that an accepted remaining lot is still unsatisfactory. It is computed from the posterior (6): the hypergeometric sampling distribution (3) for the whole lot combined with a prior, taken as uniform $P(k_{\\mathrm{whole}})=1/N$, the reference prior for hypergeometric sampling. Conditioning on the observed sample sidesteps the difficulty that blocks the frequentist risk (4), whose conditioning event $k_{\\mathrm{rem}} = k_{\\mathrm{whole}} - Y$ involves the observed $Y$. The second piece is the square-bracket representation $[N-n, Ac]$: fixing","core_discovery":"Central claim: sampling plans indexed by whole-lot quality, in particular ISO 2859-2's, cannot assess the lot left after destructive sampling. The frequentist consumer's risk conditions on $k_{\\mathrm{rem}} = k_{\\mathrm{whole}} - Y$, which involves the sample count $Y$, so the hypergeometric distribution cannot describe it. The paper turns to the Bayesian specific consumer's risk, $P(k_{\\mathrm{rem}} \\ge \\lceil LQ\\cdot(N-n)\\rceil \\mid y \\le Ac)$, under a uniform reference prior. ISO 2859-2 plans then exceed the intended 10% risk for small lots, peaking at 44% for plan $(50,0)$ at $N=90$ with $LQ=2\\%$. The paper designs accept-zero plans holding this risk to 10%, tabulated in a new representa","pith_inferences":["The square-bracket representation likely transfers beyond destructive testing to any setting where sampled units leave the population of interest — audit sampling, food-safety testing that consumes the sample, or trials with destructive endpoints — and the Bayesian risk (5) would then define the protection guarantee there as well.","Because the 44% worst case is computed under a uniform prior, priors more pessimistic about lot quality would widen the gap between the standard's stated protection and what destructive sampling actually delivers; the failure is worst in exactly the small-lot regimes where destructive testing is common.","The reported linearity of lot-size ranges in the prior parameters suggests a simple on-the-fly recalibration: fit a beta-binomial prior to inspection history and adjust table ranges by the measured slopes, avoiding new optimization.","The tables assume the consumer adopts the same reference prior; in regulated settings (legal metrology, food safety) a standards body choosing the prior would effectively set the risk policy, a choice the paper leaves open."],"forward_implications":["If the paper is right, ISO 2859-2 plans should not be used when the sample destroys items, and future editions of the standard should state this limitation explicitly.","Destructive-sampling plans should be designed to limit the Bayesian specific consumer's risk (5); with the uniform reference prior, the required accept-zero sample sizes exceed the ISO 2859-2 sample sizes for whole lots up to $N=500$ at $LQ=2\\%$.","The square-bracket representation $[N-n, Ac]$ lets standard setters tabulate destructive-sampling plans whose sample size stays within $1/LQ$ of the optimal one, even for small lots.","The tabulated plans are stable under changes of the conjugate prior: the applicable lot-size ranges shift roughly linearly, about 10 in $N$ per 0.1 change in $a$ and about $-1$ in $N$ per unit change in $b$ at $LQ=2\\%$, and moving prior weight toward acceptable quality only improves consumer protection.","The same Bayesian construction and square-bracket format extend to informative priors, variable sampling, and to plans that limit producer's risk or minimize costs."],"supporting_citations":[{"why":"Supplies the ISO 2859-2 sampling plans and their claimed 10% consumer-risk limit that the paper tests against remaining-lot quality.","marker":"[ISO/TC 69/SC 5, 2020]"},{"why":"Provides the hypergeometric distribution (3) used as the sampling distribution throughout the argument.","marker":"[Johnson, Kotz and Kemp, 1992]"},{"why":"Defines the specific consumer's risk used in (5), the quantity the new plans limit.","marker":"[BIPM et al., 2012]"},{"why":"Justifies the uniform prior $P(k_{\\mathrm{whole}})=1/N$ as the reference prior for hypergeometric sampling.","marker":"[Bernardo, 2003]"},{"why":"Gives the beta-binomial family as conjugate to the hypergeometric sampling distribution, underpinning the posterior (6) and the prior-sensitivity analysis.","marker":"[Dyer and Pierce, 1993]"},{"why":"Precursor observation that the hypergeometric distribution overestimates confidence in low nonconforming counts in residual seed lots.","marker":"[Herman and Robbins, 2013]"},{"why":"Review of prediction intervals, cited to argue that interpreting the frequentist risk (4) as a prediction interval yields no unique answer.","marker":"[Patel, 1989]"}],"fun_headline_variants":["ISO 2859-2 fails destructive sampling: 44% risk","New sampling plans cap destructive-test risk at 10%","ISO plans miss remaining-lot quality, hit 44% risk","Destructive testing: ISO plans overstate safety, new fix","For destructive tests, ISO 2859-2 risk spikes to 44%"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The plans cap the risk at 10% only under the uniform prior $P(k_{\\mathrm{whole}})=1/N$; if the consumer's true prior differs — the sensitivity analysis (Section 4, Figures 5a-5c) shows a shift from $a=1.0$ to $a=1.1$ in a $\\beta$-binomial prior can already push the risk over the limit for the smallest lot sizes — the promised protection weakens.","fun_headline_variants_meta":{"raw":{"variants":["ISO 2859-2 fails destructive sampling: 44% risk","New sampling plans cap destructive-test risk at 10%","ISO plans miss remaining-lot quality, hit 44% risk","Destructive testing: ISO plans overstate safety, new fix","For destructive tests, ISO 2859-2 risk spikes to 44%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1241,"prompt_tokens":817,"completion_tokens":424,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":332}},"tokens_in":561,"tokens_out":424,"duration_ms":4316,"temperature":1.0,"reasoning_tokens":332,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:31:13.217750+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the Bayesian specific consumer's risk (5) for the ISO 2859-2 plan $n=50$, $Ac=0$, $LQ=2\\%$, $N=90$ with the uniform prior: the paper predicts 44%. If an independent calculation, or a simulation of repeated destructive-sampling campaigns following ISO 2859-2, puts that risk at or below 10%, the central claim fails. A second check: compute the risk for the paper's proposed plan $[51,0]$ at the smallest lot in its range, $N=160$, with a $\\beta$-binomial prior $a=1.1$, $b=1$; the paper's sensitivity analysis implies it exceeds 10%.","supporting_citations":[],"review_version":1}