{"id":"f75d8df8-7493-45ff-b90f-36c63667e85d","arxiv_id":"2510.22688","paper_version":3,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review surveys sequential stopping rules for Monte Carlo mean estimation, organizing assumptions, algorithms, convergence properties, and trade-offs across roughly 110 references.","lead":"This paper is a review of sequential stopping rules—methods that decide when a Monte Carlo simulation has run long enough—for estimating an unknown mean, with special attention to binomial proportions and mildly generalized settings. It is useful to generalists because it organizes roughly a hundred scattered results and attempts a guidebook for practitioners and researchers.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The review's value as a guide hinges on unverifiable summaries: key recent advances are author preprints [58,59,113] and at least one displayed algorithm (§3.3.2) uses an undefined λ, so a reader cannot check or implement the core material.","rationale":"The reader's UNVERDICTED is policy-appropriate because this is a review, not a new result. My concern does not move the verdict; it strengthens the reason for not treating the survey as settled. I focus on auditability rather than on the excluded topics (confidence sequences, resampling), because those exclusions are explicit and defensible. The most concrete point is that three of the newest pillars of the review are self-cited preprints and one algorithm formula is internally incomplete; both are checkable. This is a substantive but not fatal issue, hence UNCHANGED.","tokens_in":20644,"tokens_out":9924,"duration_ms":105641,"concrete_test":"Request the PDFs of [58], [59], and [113]; for each, locate the exact statements summarized in §3.1.1 and §3.4.2 and compare displays word-for-word. Separately, search the manuscript for any definition of λ in §3.3.2. If the preprints are unobtainable or the λ definition is absent, the affected sections should be revised or marked non-verifiable before the review is used as a guide.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this review is a comprehensive, accurate guide to sequential stopping rules. What would have to be true is that the summaries of primary sources are faithful and that the 'recent developments' are accessible enough to be checked. That condition is weakest in exactly the places where the review is newest: §3.1.1 leans on [58] and [59], and §3.4.2 on [113], all unnumbered preprints (two by the corresponding author, one by both authors) without page or theorem numbers. Also, most displayed results elsewhere, e.g. the second-order expansions in §3.1.2 and the Berry–Esseen rule in §3.2.2, are stated without theorem numbers. A concrete internal defect: in §3.3.2 the formula for ϒ, '4(e−2)λ ln(2/δ)/ε^2', introduces λ but λ is never defined, so the described Dagum et al. algorithm cannot be implemented from the review. These are verification gaps, not evidence of error; but a guide whose core is not independently checkable cannot fully support the claim that it directs researchers reliably.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a review article on sequential stopping rules for Monte Carlo estimation of expectations, with emphasis on iid sampling and lightly generalized settings (bounded, Bernoulli, martingale-difference, low-discrepancy, multiple means). It first lays out essential components—error tolerances, confidence/significance levels, multivariate outputs, non-iid sequences, preliminary sample sizes, large-sample theory, normal approximation, and single- versus two-stage procedures—then surveys recent developments organized by asymptotic analyses, moment-based rules, distributional assumptions, non-iid sequences, multiple means, quality assessment, and unconventional criteria. The stated aim is to provide a comprehensive, up-to-date guidebook of core assumptions, algorithms, convergence properties, and practical trade-offs.","tokens_in":20947,"tokens_out":6394,"duration_ms":66343,"significance":"If taken at face value, the review fills a real gap: the sequential-stopping-rule literature for Monte Carlo is scattered, and the paper provides a structured map of it. The classical summaries in Sections 2 and 3 (Chow–Robbins asymptotics, Glynn–Whitt framework, two-stage kurtosis-based rules, Berry–Esseen refinements, relative-precision rules) are consistent with the original sources and will be useful to both practitioners and researchers entering the area. The paper's contribution is synthesis and organization rather than new mathematics or reproducible code. Its value therefore depends on the accuracy and verifiability of the summaries, and it is precisely there that the manuscript needs work: one displayed algorithm uses an undefined parameter, and several key recent results are drawn from unnumbered preprints without accessible details.","major_comments":[{"comment":"The first stopping rule of [27] is presented with ϒ = 4(e−2)λ ln(2/δ)/ε², but λ is never defined anywhere in the manuscript. Since τ1 = argmin{n : X1+...+Xn ≥ ϒ1} and the displayed bound E[τ1] ≤ ϒ1/E[X1] both depend on λ, a reader cannot implement or check the algorithm. This is a concrete internal defect in a section whose purpose is to summarize numerical algorithms. Please define λ (or state its range/value as in the source) and reconcile the notation with [27].","section":"§3.3.2"},{"comment":"The review's most recent content—non-asymptotic malfunction of stopping rules and a martingale-difference stopping rule—rests on two unnumbered preprints by the corresponding author ([58], [59]) and one by both authors ([113]), cited as 'available here' without URLs or theorem/page numbers. These are not standard archival sources, so the accuracy of the summarized claims cannot be independently checked from the manuscript. Since the paper's central claim is to be a reliable up-to-date guide, please either state the precise results with assumptions and theorem numbers, or provide public links and identify the exact statements being summarized.","section":"§3.1.1, §3.4.2, Refs. [58], [59], [113]"},{"comment":"The introduction promises 'comparisons' and a guidebook covering 'core assumptions, numerical algorithms, convergence properties, and practical trade-offs', but the review contains no synthesis table or side-by-side comparison of the surveyed rules beyond descriptive prose. Given the stated aim, a summary table (columns: assumptions, algorithm type, convergence properties, required inputs, limitations) would make the comparison and selection criteria explicit. This is not a correctness issue but it is load-bearing for the paper's contribution as a review.","section":"§3 (overall)"}],"minor_comments":[{"comment":"Notation for the significance level changes from δ to α without explanation; e.g., α appears in the second-order expansions after δ was used in §2.6. Please harmonize.","section":"§3.1.2"},{"comment":"In the second stopping rule, S_n := 2^{-1}∑_{k=1}^n (X_{2k-1}-X_{2k})² uses n for a number of pairs, while τ2 and τ3 are sample sizes; clarify the index conventions.","section":"§3.3.2"},{"comment":"Refs. [58], [59], [113] should include arXiv IDs or URLs if they are to be usable. There are also typos in [2] ('randam'), [35] ('Marcel Dekkerr'), and [55] (author list formatting).","section":"References"},{"comment":"Several displayed results (e.g., the Berry-Esseen stopping criterion and the [61] complexity bound) are given without theorem or proposition numbers from the sources; this makes verification unnecessarily difficult for a review.","section":"§3.2.2, §3.2.3"},{"comment":"The statement that resampling methods are not explored is helpful, but the exclusion criteria for confidence sequences and group-sequential clinical-trial methods could be stated more explicitly in one place.","section":"§2.4"}],"recommendation":"major_revision","confidential_remarks":"The paper's main weaknesses are verification gaps rather than demonstrated errors. The undefined λ in §3.3.2 and the reliance on unnumbered preprint references [58], [59], [113] should be addressed before publication. I do not see a need to reject, as the classical material is accurately summarized and the review fills a real gap."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a review paper, not a new-results paper, and that is fine. The value is organizational: it brings together the classical Chow–Robbins asymptotics, two-stage kurtosis rules, Berry–Esseen refinements, GBAS, and the recent martingale-difference and QMC stopping rules into a single coherent framework. The taxonomy by assumptions, algorithms, and convergence properties is genuinely useful, and the emphasis on coverage profiles and the non-asymptotic coverage gap (Sections 3.6 and 2.7) is the right emphasis. I would point a new PhD student here before sending them into the primary literature.\n\nThe soft spots are real but not fatal. First, the stress-test note is correct: in §3.3.2 the formula for ϒ is “4(e−2)λ ln(2/δ)/ε^2” and λ is never defined. That is a concrete defect—a reader trying to implement the Dagum et al. algorithm cannot do so from the review. It also suggests the summaries were not checked line-by-line against the sources. Second, the review leans on three unnumbered preprints ([58], [59], [113]), two by the corresponding author, with “available here” but no URL, page, or theorem numbers. Those preprints carry some of the newest material in the survey, so the most novel parts are the least verifiable. I am not accusing the authors of misrepresentation—the summaries look plausible—but a review whose stated purpose is to guide researchers reliably should give readers enough to check the claims. Third, the selection criteria for “recent” developments after 1991 are vague; the paper explicitly excludes confidence sequences, resampling, and group-sequential methods, which is a reasonable scope choice, but it would help to say why.\n\nThe mathematical summaries of the classical results in Section 2 and most of Section 3 appear consistent with the original sources, as far as I can tell without pulling them all. The paper does not fabricate new theorems, and it is honest about what is an open challenge (e.g., relative-error Bernoulli stopping). The citation pattern is not abusive; the self-citations are to the authors’ own recent work, which is relevant to the topic, though the incomplete bibliographic details are a problem.\n\nWho is this for? Practitioners and students who need a map of stopping rules and their failure modes. It deserves a serious referee, not a desk reject, but it needs revision: define λ, expand the preprint references, and add a short statement on how the recent literature was selected. I would accept an invitation to review it, but I would ask for those fixes first.\n\nWould I cite it? Yes, as the standard survey of the area, once it is cleaned up.","headline":"A competent but uneven survey: its organizational value is real, yet the most novel sections rest on unverifiable author preprints and at least one displayed algorithm is missing a parameter.","tokens_in":21410,"tokens_out":1493,"would_cite":true,"duration_ms":18390,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65C05","62E20","62L05","62L15","60G42"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the scattered literature on sequential stopping rules for Monte Carlo mean estimation becomes coherent once organized by a few core choices: absolute versus relative accuracy, iid versus dependent samples, and asymptot","keywords":["sequential stopping rules","Monte Carlo methods","fixed-width confidence intervals","absolute precision","relative precision","Bernoulli trials","martingale difference","coverage probability"],"falsifier":"A concrete check: recompute the stated cost factor for the bounded-variable Bernoulli rule (claimed between 3.64 and 5.09 times the normal-approximation sample size for significance levels between 0.001 and 0.1); if the factor falls outside that range, the summary is wrong. For the comprehensiveness claim, find a stopping rule in active use for Monte Carlo mean estimation that fits no category in the review—say a confidence-sequence rule with no fixed width—and show it delivers better expected run length in the same iid setting.","tokens_in":20551,"feed_emoji":"🎯","tokens_out":4774,"duration_ms":53256,"temperature":0.7,"pith_summary":"The paper tries to establish that the apparently fragmented literature on sequential stopping rules for Monte Carlo mean estimation is actually organized around a small set of choices: whether accuracy is measured in absolute or relative terms, whether the samples are independent, whether guarantees are asymptotic or finite-sample, and what extra information (moment bounds, boundedness) is available. By collecting and comparing recent work under these headings, it aims to give practitioners and researchers a usable guide for choosing a stopping rule and a baseline for developing new ones. A sympathetic reader would care because simulations routinely waste compute or terminate early with misleading coverage, and the relevant results are currently scattered across specialized papers.","feed_headline":"Know when to stop: a review maps Monte Carlo rules","feed_subtitle":"Sequential stopping rules sorted by accuracy type, dependence, and finite-sample guarantees for a practical pick.","key_machinery":"The central object is the stopping time, the smallest sample size such that the probability the empirical mean misses the true mean by more than a tolerance is below a significance level (or the relative version with the true mean in the denominator). Because exact determination is impossible without extra information, each rule in the review is a way of approximating that stopping time: two-stage designs estimate the variance first, moment-based corrections improve the normal approximation, concentration inequalities give non-asymptotic bounds for bounded variables, and the martingale-difference rule extends the approach to dependent sequences with batch indices. The review also uses covera","core_discovery":"On its own terms, the review's central claim is that the many recent stopping rules—two-stage rules that inflate the variance estimate to guarantee coverage, rules that add third- and fourth-moment corrections to the normal approximation, relative-precision rules for bounded and Bernoulli variables, rules for martingale-difference and quasi-Monte Carlo sequences, and rules for estimating several means at once—are not isolated tricks but instances of a common design problem: finding the smallest sample size that keeps the error probability under a prescribed threshold. The paper's contribution is to state that problem once, to classify the solutions by their assumptions and guarantees, and to","pith_inferences":["A direct extension of the review's map is a decision procedure: if a practitioner can bound the support or moments, use a finite-sample rule; otherwise use an asymptotic rule with a coverage-profile pre-check and a variance inflation factor.","The deliberate exclusions—confidence sequences, resampling, and group-sequential trials—suggest a follow-up review that integrates those with the current taxonomy, since betting-based confidence sequences share the same goal with different machinery.","The coverage-profile diagnostic could be turned into a standardized benchmark: run each candidate rule on a suite of distributions and report coverage against run length, giving an empirical table that the field currently lacks."],"forward_implications":["If the classification is right, then selecting a stopping rule is a matching problem: check boundedness and moment information, decide absolute vs relative accuracy, and only then choose between asymptotic and finite-sample rules.","The review makes the classic fixed-width interval asymptotics the common benchmark, so different rules can be compared in terms of asymptotic consistency and efficiency.","Finite-sample reliability is not free: two-stage and moment-based rules require prior information such as a kurtosis bound or a bounded support, and coverage profiles show that otherwise early stopping can badly undercut coverage.","Relative-precision rules can cut run length substantially but are risky when the mean is near zero, since the procedure may stop too early; the review flags this trade-off.","Valid stopping rules exist beyond iid data, including martingale-difference sequences and low-discrepancy quasi-Monte Carlo, so dependent-sample users are not restricted to the iid toolbox."],"fun_headline_variants":["Monte Carlo's stop sign: a review of stopping rules","When to stop sampling: Monte Carlo rules dissected","Stop when it's safe: Monte Carlo stopping rules compared","How much data is enough? Monte Carlo stopping rules reviewed","The art of knowing when to stop: Monte Carlo rules"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the post-1991 papers the review selected are representative of the state of the art for Monte Carlo mean estimation, and that the authors' short summaries of each cited result are accurate.","fun_headline_variants_meta":{"raw":{"variants":["Monte Carlo's stop sign: a review of stopping rules","When to stop sampling: Monte Carlo rules dissected","Stop when it's safe: Monte Carlo stopping rules compared","How much data is enough? Monte Carlo stopping rules reviewed","The art of knowing when to stop: Monte Carlo rules"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000272,"raw_usage":{"total_tokens":1428,"prompt_tokens":665,"completion_tokens":763,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":409,"completion_tokens_details":{"reasoning_tokens":682}},"tokens_in":409,"tokens_out":763,"duration_ms":8089,"temperature":1.0,"reasoning_tokens":682,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T07:59:46.587897+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: recompute the stated cost factor for the bounded-variable Bernoulli rule (claimed between 3.64 and 5.09 times the normal-approximation sample size for significance levels between 0.001 and 0.1); if the factor falls outside that range, the summary is wrong. For the comprehensiveness claim, find a stopping rule in active use for Monte Carlo mean estimation that fits no category in the review—say a confidence-sequence rule with no fixed width—and show it delivers better expected run length in the same iid setting.","supporting_citations":[],"review_version":1}