{"id":"b79ca328-e2be-4f0d-b57b-e922e9521c4b","arxiv_id":"1908.07979","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper proposes a generalized pivotal test for the root mean square parameter ρ in a random-effects model, showing better type I error control and power than large-sample Z-tests in simulations and an oximetry example.","lead":"This paper develops a new statistical test for showing two medical devices agree, by testing the root mean square of paired differences. It gives more reliable and powerful results than existing methods for small studies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Qb in §2.1 is undefined for a positive fraction of Monte Carlo draws: simulated SSR can exceed the maximum possible weighted between-group sum of squares, so no nonnegative σ_b^2 solves eq. (2), and the paper supplies no fallback or existence proof.","rationale":"The reader's weakest assumption correctly targets the construction of Qb in §2.1, but it frames the risk as non-existence or non-uniqueness of the solution to eq. (2). In fact, for fixed ȳ and Qw the left side of eq. (2) is strictly decreasing in σ_b^2, so uniqueness is guaranteed; the real gap is existence. Simulated SSR values can exceed the maximum L(0) attainable with σ_b^2 ≥ 0, and this happens with substantial probability near the boundary σ_b^2 = 0, a region the paper's simulations do not explore. Without an explicit algorithm and distributional analysis for such draws, the claimed parameter-free distribution of Q and the exactness of Pr(Q ≥ ρ_0^2) are not established. This is a concrete, addressable gap rather than a demonstrated failure: the R package may already use a sensible fallback, and the method's extensive simulations show controlled type I error away from the boundary. I therefore keep the conditional verdict: the paper should supply the solving algorithm, the fallback rule, and either a proof or a numerical verification that the resulting Q retains the pivotal property for small and zero σ_b^2. This concern does not change the reader's overall assessment, but it sharpens the condition on which acceptance should depend.","tokens_in":18517,"tokens_out":19620,"duration_ms":201177,"concrete_test":"Instrument the public R package RAMgt (github.com/baolinwu/RAMgt) to record, for every Monte Carlo draw in step (ii) of §2.2, whether the equation SSR = Σ w_i(Qb)(ȳ_i − ȳ_w)^2, w_i(Qb) = 1/(Qb + Qw/m_i), has a nonnegative solution Qb. Run this diagnostic on 10^4 datasets simulated from balanced designs with n = 10, m = 5, σ_w^2/σ_b^2 in {3, 1/3}, and true σ_b^2 set to 0, 0.1, and 1.0 times the nominal scale used in §3. If the non-solvable fraction exceeds 0.01 in any configuration, recompute the p-values with the package's actual handling of non-solvable draws and compare them against p-values computed after explicitly setting Qb = 0 for those draws; a difference beyond Monte Carlo error would demonstrate that the omitted algorithm is load-bearing for the claimed exact inference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The claimed exactness of the generalized p-value in §2.2 requires Qb = h(ȳ, Qw, SSR) to be a well-defined function of parameter-free random variables with the property Qb = σ_b^2 at the observed data. The paper only states that eq. (2) can be solved for σ_b^2; it gives no algorithm, no proof of existence, and no discussion of what is done when a solution does not exist. For fixed observed group means ȳ_i and a simulated Qw, the left side of eq. (2), L(b) = Σ w_i(b)(ȳ_i − ȳ_w(b))^2 with w_i(b) = 1/(b + Qw/m_i), is strictly decreasing in b ≥ 0 (its derivative is −Σ w_i^2(ȳ_i − ȳ_w)^2 < 0), so a solution is unique when it exists. However, L(b) is bounded above by L(0) = Σ m_i(ȳ_i − ȳ_w)^2 / Qw. A simulated SSR drawn from χ^2_{n−1} can exceed L(0), in which case no nonnegative σ_b^2 solves eq. (2). This is not a measure-zero event: when σ_b^2 is small (e.g., at the boundary σ_b^2 = 0), L(0) is itself on the scale of a χ^2_{n−1} variable, so the conditional probability that an independent simulated SSR exceeds L(0) can be about one-half. The simulation study in §3 only uses σ_w^2/σ_b^2 ratios 1/3, 1, 3, so the problematic small-σ_b^2 region is never exercised. If the R package silently handles non-solvable draws by clipping Qb to 0, resampling SSR, or allowing negative values, that fallback changes the distribution of Qb and must be stated and validated before the parameter-free-distribution and exactness claims in §2.2 can be accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript develops a generalized pivotal quantity test for the equivalency hypothesis H0: ρ ≥ ρ0 versus Ha: ρ < ρ0, where ρ = sqrt(μ² + σ_b² + σ_w²) is the root mean square of paired differences in a one-way random-effects model. The method constructs pivots Qw for σ_w², Qb for σ_b² (by inverting the between-group weighted sum-of-squares identity), and Qμ for μ², and uses Q = Qw + Qb + Qμ as the test statistic. The authors report simulation studies with 10^4 replications showing controlled type I error and improved power relative to a Z-score test, illustrate the method on a pulse oximetry equivalency study, and provide an R package.","tokens_in":18986,"tokens_out":12311,"duration_ms":119383,"significance":"If the pivot construction is rigorously valid, this paper fills a practically important gap: a small-sample test for the FDA-relevant RMS parameter in paired repeated-measures device comparison studies. The simulation evidence is extensive and the publicly available R package is a valuable contribution; the power advantage over the Z-score test, especially at α = 0.01, is substantial in the studied settings. However, the load-bearing definition of Qb has a gap that must be addressed before the claims of a parameter-free pivot distribution and exact generalized inference can be accepted.","major_comments":[{"comment":"The existence of the function h used to define Qb is not established for all Monte Carlo draws. For fixed observed group means ȳ_i and a simulated Qw, the right-hand side of Eq. (2) is a strictly decreasing function of σ_b² on [0,∞) with supremum L(0) = Σ m_i(ȳ_i − ȳ_w)²/Qw; therefore a nonnegative solution exists only when the simulated SSR is no larger than L(0). Simulated SSR ∼ χ²_{n−1} can exceed L(0) with non-negligible probability, especially when σ_b² is near zero, so the non-solvable set is not measure zero. Since Qb is undefined on those draws, the statement in §2.1 that Qb has a parameter-free distribution and the generalized p-value construction in §2.2 are not justified. The authors must either prove the event has probability zero (which is not the case), or redefine the pivot so it is always well-defined, and they must specify the numerical algorithm used in the R package.","section":"§2.1, Eq. (2)–(3)"},{"comment":"The simulation grid never exercises the σ_b² near zero boundary where the existence failure above is most severe; the smallest σ_b² configuration used is σ_w²:σ_b² = 3:1, which keeps σ_b² at 25% of the total variance component. If the implementation silently handles non-solvable draws (for example by clipping Qb to zero or resampling SSR), the distribution of Q changes and the reported type I error and power results do not necessarily describe the method as defined in §2.1. The paper should report the frequency of non-solvable draws across configurations and validate any fallback rule.","section":"§3.1–3.2, Tables 1–4 and S1–S6"}],"minor_comments":[{"comment":"The text says \"Table S6 summarizes the power for the unbalanced study,\" but Table S6 reports balanced scenarios with n=30; the unbalanced-study power appears in Table 3. The cross-reference should be corrected.","section":"§3.2"},{"comment":"The Z-score and Z-Wald tests are named but their construction is not fully specified; equation (8) gives Var(R), but the paper should state explicitly how the variance is estimated under the null for the score test and at the MLE for the Wald test.","section":"§2.4"},{"comment":"The notation Sχ²_1(·|λ) is introduced without a clear definition of the argument order; the displayed formula for Pr(Q ≥ ρ0² | Qw,Qb) should define whether the first argument is the threshold or the noncentrality parameter.","section":"§2.3"},{"comment":"The appendix uses sse for the within-group sum of squares while the main text uses SSE for the random variable; the notation should be made consistent.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The gap in the definition of Qb seems genuinely repairable, but it is not a copy-editing issue: the authors need to provide a valid pivot or an explicit, validated fallback. I recommend asking the authors to address the existence problem and to supplement or rerun the simulations in the small-σ_b² region."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear X,\n\nYou should know about this paper because its core statistic Qb is not actually defined for all Monte Carlo draws. In Section 2.1 the authors solve eq. (2) for sigma_b^2 as a function of observed group means, Qw, and a simulated chi-square SSR, but as written the equation only has a nonnegative solution when SSR is no larger than the maximum L(0) = sum_i m_i(bar y_i - bar y_w)^2 / Qw. Simulated SSR can exceed L(0), and in small-sigma_b^2 regimes it will do so often. The paper gives no fallback, no algorithm, and no proof of existence; it simply states 'we can solve.' That is a real gap in the generalized inference claim, because the distribution of Q would not be parameter-free if a solution sometimes fails to exist.\n\nThat said, the idea is genuinely new and the paper does much right. Combining generalized pivotal quantities for the mean and variance components to test the root-mean-square parameter in a one-way random effects model is not in the cited literature, and the simulation study is extensive: 10^4 reps, unbalanced and balanced designs, multiple variance ratios, and the results consistently show better power than the Z-score test at small samples while keeping type I error below nominal. The oximetry example and the public R package are useful contributions. The conservative type I rates are a minor issue, not a fatal one, since conservatism is generally acceptable in equivalence testing.\n\nThe soft spot is the construction of Qb. I checked the stress-test note and it holds up: for fixed observed group means and simulated Qw, the weighted between-group sum of squares is strictly decreasing in sigma_b^2 and bounded above by its value at zero, so a solution exists only if the simulated SSR falls below that bound. The simulation configurations keep sigma_b^2 well away from zero (ratios 1/3, 1, and 3), so the failure region is never exercised. The authors need to state how the R package handles non-solvable draws — clipping, rejecting, re-drawing — and show that the resulting distribution of Qb is still correct, or modify the pivot so a solution always exists. Without that, the exactness claim in Section 2.2 is not established.\n\nWho benefits: statisticians working on FDA submissions for pulse oximeters and similar paired device comparisons. It deserves a serious referee, but the referee should ask for a precise existence proof or a corrected algorithm, plus simulations near sigma_b^2 = 0. I would not desk-reject it. Recommend peer review with major revision.","headline":"Promising generalized pivotal test for RMS equivalence, but the Qb construction can fail to exist and needs referee attention.","tokens_in":19461,"tokens_out":5634,"would_cite":false,"duration_ms":50954,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62J10","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper develops a generalized pivotal test that claims exact parameter-free inference for the root mean square \\(\\rho\\) in paired repeated-measures device comparisons, outperforming large-sample approximations.","keywords":["equivalence testing","generalized pivotal quantity","root mean square","paired repeated measures","linear mixed effects model","pulse oximetry","type I error","generalized confidence interval"],"falsifier":"Simulate null data with small n (for example n=10 subjects with m=5 repeated measures), \\(\\rho=\\rho_0\\), and between-subject variance large relative to within-subject variance, compute the generalized p-value \\(\\Pr(Q \\ge \\$rho_0^{2}$)\\) over many replicates, and check whether the empirical rejection rate at \\(\\$\\alpha$=0.05\\) stays within binomial tolerance of 0.05. Also record the fraction of Monte Carlo draws of \\((SSR, Q_w)\\) for which the equation defining \\(Q_b\\) has no positive real solution or yields multiple solutions; a non-negligible fraction would break the distributional claim.","tokens_in":18340,"feed_emoji":"🩺","tokens_out":7423,"duration_ms":66281,"temperature":0.7,"pith_summary":"The paper develops a test for whether two measurement methods agree, judged by the root mean square \\(\\rho\\) of their paired differences, in studies with repeated measures per subject. It claims that \\(\\$rho^{2}$\\) can be decomposed into three generalized pivotal quantities whose joint distribution contains no unknown parameters once the observed summary statistics are fixed, so the tail probability \\(\\Pr(Q \\ge \\$rho_0^{2}$)\\) supplies a valid p-value for the equivalence hypothesis \\(H_0: \\rho \\ge \\rho_0\\). Existing approaches rely on large-sample normal approximations and are conservative, especially at \\(\\$\\alpha$=0.01\\); simulations show the new test controls type I error and has higher power. Applied to a pulse oximetry comparison, it gives a p-value of 0.006 and a narrower confidence interval than the Z-score test, supporting device equivalency.","feed_headline":"Oximeter equivalence test holds type I error without large samples","feed_subtitle":"A generalized pivotal statistic for the root mean square gives controlled type I error and higher power.","key_machinery":"The central object is the generalized test statistic \\(Q = Q_w + Q_b + Q_\\mu\\). \\(Q_w = sse/(SSE/\\$sigma_w^{2}$)\\) has a scaled inverse chi-square distribution; \\(Q_b = h(\\bar{y}, Q_w, SSR)\\) is defined as the solution to the weighted between-subject sum-of-squares identity (2) for \\(\\$sigma_b^{2}$\\); \\(Q_\\mu = [\\tilde{y} - Z(\\sum_i \\tilde{W}_i)^{-1/2}]^2\\) is the squared normal pivot for the mean using the same simulated variance components. The machinery works because, conditional on the observed summary statistics, the randomness in \\(Q\\) comes only from \\(\\$chi^{2}$_{N-n}\\), \\(\\$chi^{2}$_{n-1}\\), and \\(N(0,1)\\) variables, so the distribution of \\(Q\\) is free of unknown parameters; and because \\(Q\\) collapses to \\(\\$rho^{2}$\\) at the observed data, tail probabilities of \\(Q\\) serve as p-values. A numerical step integrates out \\(Z\\) analytically, reducing the tail probability to a noncentral chi-square calculation.","core_discovery":"The paper claims that the root mean square parameter \\(\\rho = \\sqrt{\\$mu^{2}$ + \\$sigma_b^{2}$ + \\$sigma_w^{2}$}\\) in a one-way random-effects model can be tested by a generalized pivotal quantity \\(Q = Q_w + Q_b + Q_\\mu\\). \\(Q_w\\) follows a scaled inverse chi-square distribution conditional on the observed within-subject sum of squares; \\(Q_b\\) is obtained by solving equation (2) for \\(\\$sigma_b^{2}$\\) as a deterministic function of the observed sample means, the simulated \\(Q_w\\), and a chi-square between-subject sum of squares; \\(Q_\\mu\\) is the squared normal pivot for the mean. The distribution of \\(Q\\) depends only on known distributions once the observed summary statistics \\((\\bar{y}, sse)\\) are fixed, and \\(Q\\) reduces to \\(\\$rho^{2}$\\) at the observed data. Therefore the paper asserts that \\(\\Pr(Q \\ge \\$rho_0^{2}$)\\) is a valid generalized p-value for \\(H_0: \\rho \\ge \\rho_0\\), and that the percentiles of \\(Q\\) give a generalized confidence interval for \\(\\rho\\).","pith_inferences":["Editorial inference: the same three-part pivot decomposition could be adapted to test other composite parameters that are sums of a squared mean and variance components, such as the coefficient of variation, provided the solving step for the variance component remains tractable.","Editorial inference: the observed-data collapse \\(Q \\to \\rho^2\\) suggests the generalized confidence interval could be inverted for sample-size planning of equivalence studies, a use the paper does not develop.","Editorial inference: because the paper does not prove uniqueness of the solution to equation (2), a numerical study of how often the solving step produces no root or multiple roots across parameter configurations would clarify the method's practical domain.","Editorial inference: the analytical integration over \\(Z\\) implies the method's Monte Carlo error comes only from the chi-square draws, so the p-value could in principle be computed deterministically to arbitrary precision."],"forward_implications":["Equivalence testing of pulse oximeters against an FDA threshold can be performed at small to medium sample sizes with type I error near the nominal level, instead of conservative Z-score tests.","The generalized confidence interval for \\(\\rho\\) has coverage close to the nominal 90% and is narrower on average than intervals from the large-sample Z-score approach.","Because the method needs only the per-subject sample sizes, means, and the residual sum of squares, it can be applied to published summary statistics without raw data.","At the stringent \\(\\alpha=0.01\\) level, the proposed test preserves power in settings where the Z-score test nearly loses power, for example 49.3% versus 3.2% in the unbalanced n=16 scenario."],"supporting_citations":[{"why":"Supplies the generalized inference framework that the constructed test statistic \\(Q\\) builds on.","marker":"Weerahandi, 1995"},{"why":"Introduces generalized p-values, which justify computing the test p-value as \\(\\Pr(Q \\ge \\rho_0^2)\\).","marker":"Tsui and Weerahandi, 1989"},{"why":"Introduces generalized confidence intervals, which the paper uses to derive the interval for \\(\\rho\\).","marker":"Weerahandi, 1993"},{"why":"Provides the model and pivot used to construct \\(Q_\\mu\\) for the mean parameter.","marker":"Iyer et al., 2004"},{"why":"Defines the FDA-motivated RMS testing problem and the large-sample approach the new method is compared against.","marker":"Pennello, 2002, 2003"},{"why":"Supplies the Z-score large-sample test that serves as the main numerical baseline.","marker":"Ndikintum and Rao, 2016"},{"why":"Frames the paired repeated-measures data as a linear mixed-effects model with between- and within-subject variance components.","marker":"Laird and Ware, 1982"}],"fun_headline_variants":["RMS equivalence test for oximeters without large-sample approximation","Generalized pivotal test controls error in RMS equivalence studies","Improved equivalency test for device comparisons using RMS","Small-sample RMS equivalence test for medical devices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the mathematical relationship used to pin down the between-subject variance \\(\\$sigma_b^{2}$\\) has a unique solution for every simulated draw and that the resulting quantity \\(Q_b\\) has a distribution free of unknown parameters; the paper assumes this without giving a proof or explicit algorithm.","fun_headline_variants_meta":{"raw":{"variants":["RMS equivalence test for oximeters without large-sample approximation","Generalized pivotal test controls error in RMS equivalence studies","Improved equivalency test for device comparisons using RMS","Small-sample RMS equivalence test for medical devices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000267,"raw_usage":{"total_tokens":1659,"prompt_tokens":1032,"completion_tokens":627,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":566}},"tokens_in":648,"tokens_out":627,"duration_ms":6899,"temperature":1.0,"reasoning_tokens":566,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:52:14.082885+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate null data with small n (for example n=10 subjects with m=5 repeated measures), \\(\\rho=\\rho_0\\), and between-subject variance large relative to within-subject variance, compute the generalized p-value \\(\\Pr(Q \\ge \\$rho_0^{2}$)\\) over many replicates, and check whether the empirical rejection rate at \\(\\$\\alpha$=0.05\\) stays within binomial tolerance of 0.05. Also record the fraction of Monte Carlo draws of \\((SSR, Q_w)\\) for which the equation defining \\(Q_b\\) has no positive real solution or yields multiple solutions; a non-negligible fraction would break the distributional claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the generalized inference framework that the constructed test statistic \\(Q\\) builds on."},{"cited_title":"and Weerahandi, S","cited_arxiv_id":null,"evidence_quote":"Introduces generalized p-values, which justify computing the test p-value as \\(\\Pr(Q \\ge \\rho_0^2)\\)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces generalized confidence intervals, which the paper uses to derive the interval for \\(\\rho\\)."},{"cited_title":"K., Wang, C","cited_arxiv_id":null,"evidence_quote":"Provides the model and pivot used to construct \\(Q_\\mu\\) for the mean parameter."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the FDA-motivated RMS testing problem and the large-sample approach the new method is compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Z-score large-sample test that serves as the main numerical baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames the paired repeated-measures data as a linear mixed-effects model with between- and within-subject variance components."}],"review_version":1}