{"id":"13946149-56f8-4f02-80ee-08971dbb5873","arxiv_id":"2608.01248","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A tolerance-based sequential test, the twin-peaks test, certifies a continuous parameter with asymptotic sample cost 2 log(1/epsilon)/(I(theta) delta^2) and is demonstrated for qubit phase and purity.","lead":"The paper introduces sequential parameter testing, a method that keeps measuring a continuous parameter until every competing value outside a chosen tolerance is ruled out with a target confidence. It applies the method to estimating the phase and purity of a qubit, showing sample savings over fixed-sample protocols.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Adaptive phase-testing stopping-time asymptotics rest on unproven a.s. convergence and first-passage exchange; the adaptive-vs-collective equivalence is supported only at one threshold.","rationale":"The reader's weakest_assumption is the same as mine: the stopping-time formulas are proved only under regularity and first-passage conditions that are not verified for adaptive designs. I agree because every quantitative headline—Eq. (35), Eq. (40), Eq. (67), and Fig. 3—inherits this gap. I do not instead anchor on the finite-sample calibration caveat, because the paper explicitly states that A_epsilon is an asymptotic calibration parameter rather than an exact posterior error bound; that is a disclosed limitation rather than a hidden gap. The paper deserves credit for the exact KL computation for the covariant phase POVM, the analytic purity estimates, and the numerical phase and purity studies. However, the greedy adaptive phase protocol lies outside the i.i.d. regime of Theorem A.1, and a single-threshold numerical comparison cannot substitute for the missing first-passage analysis. No change to the reader's CONDITIONAL verdict is needed; the condition should be that the authors supply rigorous conditions, or a proof, for the adaptive stopping-time asymptotics and release the simulation code.","tokens_in":33982,"tokens_out":16759,"duration_ms":174736,"concrete_test":"Set delta=0.1, fix true theta=0, and run the greedy adaptive phase protocol with the twin-peaks stopping rule disabled. For a grid of competitor values theta0, record S_n(theta0)=(1/n) log[p(m_1..m_n|theta)/p(m_1..m_n|theta0)] for n up to 10^7. Check that S_n -> -log cos(theta0-theta) almost surely, the KL rate of the perpendicular PVM. Separately, run the full test at thresholds A=10,10^2,...,10^6 and check whether E(N)-log A/(2 sin^2(delta/2)) stays O(1). If S_n fails to converge, Eq. (37) is violated for the adaptive strategy; if E(N) has an excess growing in log A, the 'additional first-passage and integrability conditions' of App. A.4 are not satisfied and the asymptotic stopping-time claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims—Eq. (35), the small-delta form Eq. (40), and the phase adaptive/collective equivalence—require the twin-peaks stopping time to obey classical SPRT asymptotics. For i.i.d. measurements this is plausible; for the greedy adaptive phase protocol of Sec. VII, the measurement axis at round k+1 is phi_{k+1}=theta_hat_k+pi/2, a function of the accumulated data. The normalized log-likelihood process (1/n)log[p(m_n|theta)/p(m_n|theta0)] is then a sum of conditionally dependent, design-dependent increments, and the uniform a.s. convergence assumed in Theorem A.1 (Eq. A27) is not proved for it. The paper's phrase 'adaptive experiments with stabilized likelihood increments' asserts rather than establishes this condition. Worse, App. A.4 derives only N/A_epsilon -> 1/inf D a.s. and then invokes 'additional first-passage and integrability conditions' to exchange the large-threshold limit with the expectation; these conditions are never stated or checked. For a stopping time defined as the first passage of a supremum over a continuum of correlated log-likelihood ratios, boundary overshoot and selection effects need not be O(1); if they contribute a growing term, Eq. (35) and Eq. (40) would overstate the resource savings. The key phase claim—adaptive projective measurements match collective covariant cost—is illustrated in Fig. 3 only at A_epsilon=19, so the missing proof is the load-bearing weak point.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a sequential framework called parameter testing for continuous parameters, in which the goal is to certify that an unknown parameter lies within a prescribed tolerance δ while ruling out δ-separated alternatives at a target error calibration ϵ. Three stopping rules are proposed (complement, concentration, and twin-peaks tests), and the twin-peaks test is analyzed in depth: it stops when the posterior-density ratio between the current MAP estimate and its strongest δ-separated competitor exceeds a threshold A_ϵ. The main analytical results are the asymptotic mean stopping time in Eq. (35), the small-δ Fisher-information form in Eq. (40), and the claim that all three tests are exponentially equivalent under the Laplace-principle conditions of App. A.4. The framework is applied to two qubit tasks: phase testing, where i.i.d. POVMs, random projective measurements, a greedy adaptive protocol, and a collective covariant strategy are compared numerically; and purity testing, where local measurements along a known Bloch direction are shown to saturate the KL bound, direction-agnostic weak Schur measurements asymptotically match this performance, and sequential strategies yield constant-factor sample savings over fixed-sample worst-case benchmarks.","tokens_in":34237,"tokens_out":12494,"duration_ms":112641,"significance":"If the central claims hold, the paper provides a computationally simple and operationally motivated sequential procedure for continuous-parameter certification, with explicit parameter-free predictions such as Eqs. (59) and (72). The purity analysis is the strongest part: the KL-rate calculation is explicit, the numerical agreement is good in the asymptotic regime, and the comparison against fixed-sample strategies is concrete. The phase analysis is suggestive but less conclusive, because the headline adaptive/collective equivalence currently rests on numerical evidence at one threshold and on an unproven extension of classical SPRT asymptotics to data-dependent measurements. The paper is also useful in clarifying that the twin-peaks threshold is an asymptotic calibration parameter rather than an exact finite-sample posterior-error probability.","major_comments":[{"comment":"The stopping-time asymptotics are not established for the adaptive phase protocol of Sec. VII. Theorem A.1 assumes almost-sure uniform convergence of the normalized log-likelihood to a deterministic limit, and the text says this holds for \"adaptive experiments with stabilized likelihood increments,\" but for the greedy rule φ_{k+1}=θ̂_k+π/2 of Eq. (57) the conditional outcome distribution at each round depends on the full past. No argument is given for uniform a.s. convergence of (1/n)log[p(m_n|θ)/p(m_n|θ0)] over θ0 outside B_δ(θ), nor for the first-passage properties needed to convert this into a mean stopping time. For i.i.d. strategies the argument is plausible, but the adaptive/collective equivalence requires either a proof for this specific design or an explicit statement that it is a numerical conjecture.","section":"Sec. V and App. A.4 (Eqs. (35), (40), (A27))"},{"comment":"The passage from the almost-sure statement N/A_ϵ → 1/inf D to the expectation E(N|θ*) = (log A_ϵ)/inf D + o(log ϵ^{-1}) is made by invoking \"additional first-passage and integrability conditions\" that are never stated or checked. For the twin-peaks stopping time, which is the first passage of a supremum over a continuum of correlated log-likelihood processes, boundary overshoot and the data-dependent selection of the maximizing competitor need not be O(1); if they contribute a growing term, Eq. (35) and Eq. (40) would overstate the resource savings. These conditions should be stated explicitly and verified for the i.i.d. models studied, and at least discussed for the adaptive protocol.","section":"App. A.4, Eqs. (A43)-(A44)"},{"comment":"The claim that adaptive projective measurements achieve the same average sample cost as collective covariant measurements is demonstrated only at the single threshold A_ϵ=19. Since the asymptotic proof for the adaptive protocol is missing (see the first major comment), one numerical point cannot establish the claimed equivalence as a general statement. Additional thresholds and values of δ, together with an analytic argument, are needed before this claim can be presented as a theorem rather than as numerical evidence.","section":"Sec. VII, Fig. 3"}],"minor_comments":[{"comment":"Several equations contain corrupted symbol artifacts such as \"/leftr⫯g⊸tl⫯ne\" and \"⌟⟨rro⟪⟪⟩r⟪\", which make parts of the manuscript unreadable. These need to be fixed in the final version.","section":"Throughout (Eqs. (37), (66), (A27))"},{"comment":"For small δ and small r the simulations hit the 15000-sample cap, and the flat portions of the curves are an artifact of this cap. The captions should state this explicitly so that these regions are not read as physical predictions.","section":"Figs. 4 and 5"},{"comment":"The notation uses the same symbol θ for the true parameter and for the argument of the maximum; using θ* for the true value and θ̂ for the MAP estimate would avoid ambiguity.","section":"Eq. (35)"},{"comment":"The sentence \"we set A_ϵ = log(1−ϵ)−logϵ\" is inconsistent with Eq. (9), where A_ϵ=(1−ϵ)/ϵ is a ratio; the logarithm belongs in the stopping-time formula rather than in the definition of the threshold.","section":"App. A.4"},{"comment":"Several references (e.g., [38], [39], [45]) are arXiv preprints without journal identifiers; these should be updated if published versions are available.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The i.i.d. and purity parts of the paper are solid and likely publishable, but the adaptive phase-testing claim is the main risk. The missing proof of the SPRT asymptotics for the data-dependent greedy measurement is load-bearing; either it should be supplied or the claim should be explicitly softened to numerical evidence. A revision focused on this point is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this paper introduces a genuinely new primitive — sequential parameter testing, with the twin-peaks test as the natural continuous analogue of an SPRT — and it probably works. The stopping-time formulas are internally consistent, reproduce the Fisher-information limit in the small-δ regime, and match the numerics in the asymptotic regime. No fitted parameters, no circularity.\n\nWhat is actually new: the framework extends sequential hypothesis testing from discrete or composite hypotheses to a continuum of tolerance regions, and the twin-peaks test is a clean way to do that without discretization artifacts. The applications to qubit phase and purity certification are well chosen. The purity section is the strongest part: the analysis is classical, the KL prediction is explicit, the fluctuation calculation (relative variance vanishing like δ) is a nice touch, and the comparison to fixed-sample benchmarks shows a constant-factor saving. The phase section also gives a clear example where adaptive measurements beat i.i.d. ones, and where collective covariant measurements set the benchmark.\n\nThe soft spots are real, but not uniformly fatal. The central stopping-time asymptotics rely on unproven regularity and first-passage conditions. Theorem A.1 assumes uniform a.s. convergence of the normalized log-likelihood; for the adaptive phase protocol, the measurement axis depends on the accumulated data, so the i.i.d. justification does not apply. The phrase 'adaptive experiments with stabilized likelihood increments' asserts exactly what needs to be proven. The exchange of the large-threshold limit with the expectation is also hand-waved via 'additional first-passage and integrability conditions' that are never stated. If those conditions fail, the predicted E(N) and the adaptive-vs-collective equivalence would overstate the resource savings.\n\nThe phase adaptive/collective equivalence is numerical, and the plot is shown at only one threshold (A=19). The purity simulations are censored at 15000 samples, which matters most near r=0. No code is released. And the paper does not cite the confidence-sequence/e-value literature, which is a genuine omission given the overlap in methods.\n\nNone of this kills the core idea. The i.i.d. results and the purity analysis are on solid ground, and the adaptive claim is plausible. But the paper currently presents conjectures as theorems in the adaptive case. That is fixable, and the fix is important.\n\nWho gets value: anyone working on certification thresholds (QKD, fault-tolerance, device verification) or sequential quantum metrology. I would send this to a serious referee, but with a clear request: either prove the adaptive convergence, or state it as a conjecture and soften the equivalence claim. A reproducibility pass — code, or at least a finite-sample calibration statement — would also strengthen it. I would cite the twin-peaks test in my own work; it is a useful concept even before the adaptive proof is settled.","headline":"Genuinely new sequential parameter-testing framework with a clean twin-peaks test; the i.i.d. and purity results hold up, but the adaptive phase asymptotics are asserted more than proven.","tokens_in":34810,"tokens_out":1792,"would_cite":true,"duration_ms":18975,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A twin-peaks stopping rule certifies an unknown continuous parameter within a prescribed tolerance, with expected sample cost set by the hardest δ-separated alternative and, for small δ, by the Fisher information.","keywords":["sequential parameter testing","twin-peaks test","sequential hypothesis testing","quantum state certification","phase estimation","purity estimation","Fisher information","quantum relative entropy"],"falsifier":"Simulate the adaptive greedy phase-testing protocol at fixed threshold $A_\\epsilon$ and small $\\delta$ for many true phases, and compare the empirical mean stopping time with $$\\max_{\\theta_0:\\|\\$\\theta$-\\theta_0\\|\\geq\\delta}\\frac{\\log A_\\epsilon+\\log(\\pi(\\theta_0)/\\pi(\\$\\theta$))}{D(\\$\\theta$\\|\\theta_0)}.$$ A discrepancy that does not shrink as $\\epsilon\\to0$ would show that the assumed convergence or the exchange of limits fails.","tokens_in":33780,"feed_emoji":"🎯","tokens_out":12066,"duration_ms":100191,"temperature":0.7,"pith_summary":"This paper introduces sequential parameter testing: rather than estimating a continuous parameter as closely as possible, one certifies that it lies within a prescribed tolerance $\\delta$ by ruling out all parameter values farther than $\\delta$ with a chosen evidence threshold. The central device is the twin-peaks test, which stops as soon as the most likely parameter value beats its strongest competitor outside the tolerance region by a factor $A_\\epsilon$. The paper derives the asymptotic expected number of measurement rounds for this test, shows that the small-tolerance cost is set by the Fisher information, and applies the framework to the phase and purity of a qubit. If the analysis is right, the approach provides a run-by-run evidence guarantee and a resource-efficient route to certification tasks in which only crossing a threshold matters.","feed_headline":"Twin-peaks test certifies a quantum parameter to tolerance δ","feed_subtitle":"The test stops when the best estimate beats every distant rival, setting sample cost by Fisher information.","key_machinery":"The engine is the twin-peaks statistic $$\\$Lambda_n^{{\\mathrm{TP}}$}(\\hat{\\$\\theta$})=\\frac{p(m_n\\mid\\hat{\\$\\theta$})\\pi(\\hat{\\$\\theta$})}{\\sup_{\\theta_0\\in B_\\delta^c(\\hat{\\$\\theta$})} p(m_n\\mid\\theta_0)\\pi(\\theta_0)},$$ checked after every round against $A_\\epsilon$. It collapses the continuum of competing hypotheses into a single comparison between the current maximum-a-posteriori estimate and its strongest $\\delta$-separated rival, avoiding the region integrals needed by the complement and concentration tests. The asymptotic analysis rests on a Laplace principle: if the normalized log-likelihood $\\ell_n(\\theta)=n^{-1}\\log p(m_n\\mid\\theta)$ converges almost surely and uniformly on compact sets, then all three tests are exponentially equivalent and the mean stopping time is governed by the minimum relative-entropy rate over the $\\delta$-separated alternatives. The small-$\\delta$ law follows from the curvature identity $D(p_\\theta\\|p_{\\theta+\\delta})=\\frac12 I(\\theta)\\delta^2+o(\\delta^2)$, which makes the Fisher information the cost coefficient for high-resolution certification.","core_discovery":"The central claim is that certification of a continuous parameter can be treated as a continuous analogue of sequential likelihood-ratio testing. For a parameter $\\theta$, tolerance $\\delta$, and threshold $A_\\epsilon=(1-\\epsilon)/\\epsilon$, the twin-peaks test stops once $$\\frac{p(\\hat{\\$\\theta$}\\mid m_n)\\pi(\\hat{\\$\\theta$})}{\\sup_{\\theta_0:\\|\\theta_0-\\hat{\\$\\theta$}\\|\\geq\\delta} p(\\theta_0\\mid m_n)\\pi(\\theta_0)}\\geq A_\\epsilon.$$ Under almost-sure convergence of the normalized log-likelihoods, the mean stopping time satisfies $$E(N\\mid\\$\\theta$)\\sim \\max_{\\theta_0:\\|\\$\\theta$-\\theta_0\\|\\geq\\delta}\\frac{\\log A_\\epsilon+\\log(\\pi(\\theta_0)/\\pi(\\$\\theta$))}{D(\\$\\theta$\\|\\theta_0)},$$ where $D(\\theta\\|\\theta_0)$ is the asymptotic relative-entropy rate; for small $\\delta$ this reduces to $E(N\\mid\\theta)\\sim 2\\log(1/\\epsilon)/(I(\\theta)\\delta^2)$. For the equatorial qubit phase the paper reports numerically that adaptive projective measurements attain the same average sample count as collective covariant measurements on a fixed number of copies. For qubit purity, local projective measurements attain the asymptotic stopping time, batch Schur measurements recover the same cost when the Bloch direction is unknown, and sequential stopping saves a constant fraction over fixed-sample worst-case designs.","pith_inferences":["A pragmatic follow-up would be to calibrate $A_\\epsilon$ numerically in finite samples, rather than using the asymptotic value $(1-\\epsilon)/\\epsilon$, to turn the twin-peaks evidence ratio into an exact posterior error guarantee.","If the almost-sure convergence assumption fails for the adaptive phase protocol, a block-adaptive variant that keeps the measurement fixed for a block of rounds before re-estimating would likely restore the guarantee at a finite sample-cost premium; this is testable by simulation.","The small-$\\delta$ scaling suggests a dimension-free benchmark for certification: the sample cost per unit of guaranteed tolerance scales as $\\delta^{-2}$, independent of Hilbert-space dimension, which could be used to compare threshold tasks across different quantum platforms.","For purity with unknown Bloch direction, the paper's large-block Schur limit implies that a simple two-stage protocol (estimate the direction with a vanishing fraction of copies, then run the local twin-peaks test) should match the asymptotic cost; working out the finite-sample trade-off between the two stages is left to future work."],"forward_implications":["Every stopped run of the twin-peaks test carries a trajectory-wise evidence guarantee: the reported tolerance region contains the MAP estimate, and every parameter value outside it is suppressed by at least $A_\\epsilon$ in posterior density ratio.","The mean sample count is fixed by the hardest $\\delta$-separated alternative, $E(N\\mid\\theta)\\sim \\max_{\\theta_0:\\|\\theta-\\theta_0\\|\\geq\\delta}(\\log A_\\epsilon+\\log(\\pi(\\theta_0)/\\pi(\\theta)))/D(\\theta\\|\\theta_0)$, so the test spends more rounds exactly where distinguishability is low.","For small tolerance, $E(N\\mid\\theta)\\sim 2\\log(1/\\epsilon)/(I(\\theta)\\delta^2)$: the Fisher information of the chosen measurement is the resource coefficient, and the leading cost in $\\delta$ is universal.","For equatorial qubit phase, adaptive greedy projective measurements match the average cost of collective covariant measurements, so adaptivity can substitute for quantum memory in this certification task.","For qubit purity, sequential stopping uses a constant fraction of the fixed-sample worst-case budget: $2/3$ of that budget for a flat prior, $2/5$ for a Hilbert-Schmidt-uniform prior, and $1/4$ for a Bures-uniform prior."],"supporting_citations":[{"why":"supplies the sequential probability ratio test whose expected stopping-time formula anchors every asymptotic expression in the paper.","marker":"[1]"},{"why":"gives the multihypothesis sequential procedure in which the hardest competitor governs stopping, the discrete prototype of Eq. (35).","marker":"[19]"},{"why":"establishes asymptotic optimality of maximum-competitor sequential tests, justifying the twin-peaks stopping rule.","marker":"[23]"},{"why":"provides the universal quantum lower bound on mean stopping time in terms of quantum relative entropy for general sequential instruments.","marker":"[33]"},{"why":"shows adaptive switching between directional measurements attains both relative-entropy rates, the model for the greedy phase protocol.","marker":"[34]"},{"why":"proves that the measured relative entropy per copy converges to the quantum relative entropy, used for collective block strategies.","marker":"[27]"},{"why":"supplies an explicit sequence of collective measurements attaining the quantum relative-entropy rate.","marker":"[29]"},{"why":"identifies Fisher information as the local curvature of the Kullback-Leibler divergence, yielding the small-delta stopping law.","marker":"[51]"},{"why":"provides the Gaussian approximation of the weak Schur measurement used for direction-agnostic purity testing.","marker":"[53]"}],"fun_headline_variants":["Twin-peaks test certifies quantum parameters to tolerance","Sequential quantum testing stops when distant rivals fail","Adaptive quantum phase testing matches collective efficiency","Quantum purity certification: sequential beats fixed-sample","Continuous hypothesis testing for quantum parameters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The predicted stopping times hold only if the normalized log-likelihood ratios converge uniformly almost surely and the expected stopping threshold can be exchanged with the large-threshold limit; for the adaptive phase protocol this convergence is assumed without proof.","fun_headline_variants_meta":{"raw":{"variants":["Twin-peaks test certifies quantum parameters to tolerance","Sequential quantum testing stops when distant rivals fail","Adaptive quantum phase testing matches collective efficiency","Quantum purity certification: sequential beats fixed-sample","Continuous hypothesis testing for quantum parameters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000303,"raw_usage":{"total_tokens":1797,"prompt_tokens":1055,"completion_tokens":742,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":671,"completion_tokens_details":{"reasoning_tokens":674}},"tokens_in":671,"tokens_out":742,"duration_ms":7331,"temperature":1.0,"reasoning_tokens":674,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:10:31.562926+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the adaptive greedy phase-testing protocol at fixed threshold $A_\\epsilon$ and small $\\delta$ for many true phases, and compare the empirical mean stopping time with $$\\max_{\\theta_0:\\|\\$\\theta$-\\theta_0\\|\\geq\\delta}\\frac{\\log A_\\epsilon+\\log(\\pi(\\theta_0)/\\pi(\\$\\theta$))}{D(\\$\\theta$\\|\\theta_0)}.$$ A discrepancy that does not shrink as $\\epsilon\\to0$ would show that the assumed convergence or the exchange of limits fails.","supporting_citations":[{"cited_title":"Wald, Sequential Tests of Statistical Hypotheses, Ann","cited_arxiv_id":null,"evidence_quote":"supplies the sequential probability ratio test whose expected stopping-time formula anchors every asymptotic expression in the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"gives the multihypothesis sequential procedure in which the hardest competitor governs stopping, the discrete prototype of Eq. (35)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"establishes asymptotic optimality of maximum-competitor sequential tests, justifying the twin-peaks stopping rule."},{"cited_title":"Mart´ ınez-Vargas, C","cited_arxiv_id":null,"evidence_quote":"provides the universal quantum lower bound on mean stopping time in terms of quantum relative entropy for general sequential instruments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"shows adaptive switching between directional measurements attains both relative-entropy rates, the model for the greedy phase protocol."},{"cited_title":"Hiai and D","cited_arxiv_id":null,"evidence_quote":"proves that the measured relative entropy per copy converges to the quantum relative entropy, used for collective block strategies."},{"cited_title":"Hayashi, Optimal sequence of quantum measurements in the sense of Stein’s lemma, J","cited_arxiv_id":null,"evidence_quote":"supplies an explicit sequence of collective measurements attaining the quantum relative-entropy rate."},{"cited_title":"Amari and H","cited_arxiv_id":null,"evidence_quote":"identifies Fisher information as the local curvature of the Kullback-Leibler divergence, yielding the small-delta stopping law."},{"cited_title":"Gut ¸˘ a and J","cited_arxiv_id":null,"evidence_quote":"provides the Gaussian approximation of the weak Schur measurement used for direction-agnostic purity testing."}],"review_version":1}