{"id":"40fc78d1-31ac-4e1a-a4bd-2e2c6b15b54e","arxiv_id":"2501.14075","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new class of unbiased and consistent tests for convex-ordered distribution families, built from expected order statistics and valid for any reference distribution, including heavy-tailed cases.","lead":"This paper introduces a general nonparametric test for whether an unknown distribution lies in a class defined by a known reference distribution, such as increasing hazard rate or increasing odds rate families. The test relies on expected order statistics and remains consistent even when distributions have infinite means, making it suitable for heavy-tailed data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proof of Theorem 2 misuses an independence-requiring theorem to pass from componentwise to joint stochastic order; the central unbiasedness and monotone-power claims are not established as written.","rationale":"The reader identified the linearly interpolated quantile composition in Lemma 1 as the weakest assumption. That interpolation step is actually sound: the points (x_i, y_i) lie on the convex function H^{-1}∘F, and the secant slopes of a convex function are nondecreasing, so the piecewise-linear interpolant eH_n^{-1}∘eF_n is convex. The genuine load-bearing issue lies later in Theorem 2, where the proof converts componentwise stochastic order into multivariate stochastic order for the test statistic. The cited Theorem 1.A.3 is the standard result requiring independence of the coordinates; the coordinates here are dependent because they derive from the same empirical CDF. This is not a minor exposition issue: it is the logical bridge to the paper's headline claims. However, the gap is likely fixable by exploiting the coupling already present in Lemma 1, so the appropriate response is a conditional acceptance pending a corrected proof. My disagreement with the reader is about which assumption is weakest; the interpolation is not the problem.","tokens_in":18536,"tokens_out":38440,"duration_ms":305739,"concrete_test":"Check the exact statement of Theorem 1.A.3 in Shaked and Shantikumar (2007) to confirm it requires independence of the coordinates. Then re-derive Theorem 2 using the coupling from Lemma 1: for a fixed realization of H_n, define F_n = H_n ∘ H^{-1}∘F and verify the pointwise inequality T^{G+}_{m,p}(F_n) ≥ T^{G+}_{m,p}(H_n) for all j simultaneously. If the pointwise inequality holds, the theorem is true and the proof should be amended to cite the coupling; if it fails, the theorem is false.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The finite-sample claims of unbiasedness and monotone power (Corollary 2 and Theorem 2) rest on a stochastic-order argument that is incomplete as written. The proof of Theorem 2 establishes only componentwise stochastic inequalities eU_j ≤st eZ_j for j=1,...,m via Lemma 1, then invokes Theorem 1.A.3 of Shaked and Shantikumar (2007) to conclude that the increasing function ψ(z)=||z||_p applied to the vectors is stochastically ordered. Theorem 1.A.3, however, requires the coordinates within each vector to be independent. Here the eU_j are all functions of the same empirical CDF H_n (and eZ_j of F_n), so they are dependent; componentwise stochastic dominance does not imply multivariate stochastic dominance for dependent coordinates. The gap is repairable: the coupling in the proof of Lemma 1 (Fn = Hn ∘ H^{-1}∘F) yields a pointwise componentwise inequality for all j simultaneously, which would establish the needed joint stochastic order. But the paper does not present this argument; it relies on a theorem whose independence condition is not met. Without a corrected proof, the central advertised properties are not logically guaranteed by the text.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a class of nonparametric tests for convex-ordered families F^cx_G and F^cv_G, with null hypothesis H^G_0: F belongs to the location-scale family of a known absolutely continuous reference distribution G. The test statistics are L_p-norms of the positive or negative parts of the vector pi^G_{j:m} - \\tilde F_n(\\hat\\mu_{j:m}(F_n)), where pi^G_{j:m} = G(E(G^{-1}(B_{j:m}))) and \\hat\\mu_{j:m} is an L-estimator of the expected order statistic. The paper proves an almost-sure limit theorem for these L-estimators under extreme-value domain-of-attraction conditions, including cases where the mean is infinite, and claims finite-sample unbiasedness and monotone power through Lemma 1 and Theorem 2, plus consistency through Propositions 5 and 6. Simulations compare the method with the Proschan-Pyke test and apply it to river flow data.","tokens_in":18777,"tokens_out":14929,"duration_ms":141484,"significance":"If the finite-sample claims are properly established, this is a useful contribution: it provides a unified testing method for several shape-constrained families, including IOR, DOR, DRHR and IRHR, for which direct tests are largely unavailable, and it does so without support restrictions and with robustness to infinite means. Theorem 1, the strong law for L-estimators under heavy tails, is a valuable standalone result, and the paper is transparent about code and simulations. The central inequality Proposition 1 is taken from the authors' prior published work, but it is a Jensen-type bound rather than a fitted quantity, so I do not see a circularity problem. The main caveat is that the proof of the finite-sample stochastic-order result is incomplete as written, so the advertised unbiasedness and monotone power properties are not yet logically guaranteed by the text.","major_comments":[{"comment":"The proof passes from componentwise stochastic inequalities \\tilde U_j ≤_st \\tilde Z_j, obtained from Lemma 1, to stochastic ordering of the increasing function ψ_p(z) = ||z||_p by invoking Theorem 1.A.3 of Shaked and Shanthikumar. That theorem requires the coordinates within each vector to be independent. Here the coordinates are all functions of the same empirical CDF H_n (or F_n), so they are dependent, and componentwise stochastic dominance does not imply multivariate stochastic dominance for dependent coordinates. This is a load-bearing gap because Corollary 2 and the advertised finite-sample unbiasedness and monotone power rest on Theorem 2. The gap appears repairable: the proof of Lemma 1 constructs a coupling F_n = H_n ∘ H^{-1} ∘ F for which the inequalities hold pointwise for all j simultaneously, and applying ψ_p to both sides of that pointwise vector inequality would establish the desired stochastic order directly. The repair should be written into the proof.","section":"Section 4.1, proof of Theorem 2"},{"comment":"The second displayed inequality in Theorem 2 has the wrong direction. Under F ≤_c H, Lemma 1 gives \\tilde F_n(μ_{j:m}(F_n)) ≤ \\tilde H_n(μ_{j:m}(H_n)). Since x ↦ (π^G_{j:m} - x)_- is decreasing, the componentwise inequality reverses for the statistic T^{G-}_{m,p}. The paragraph after Lemma 1 also states that under F ≤_c H the rejection probability for H^{G}_{1-} is smaller under F than under H, and Corollary 2(2) requires P(T^{G-}_{m,p}(F_n) ≥ c^-_{α,n}) ≤ P(T^{G-}_{m,p}(H_n) ≥ c^-_{α,n}) when F ≤_c H. Therefore part 2 of Theorem 2 should read ≤, not ≥, or should be reformulated with the order reversed. As written, the theorem contradicts Corollary 2(2) when H = G and F ≤_c G.","section":"Section 4.1, Theorem 2(2)"}],"minor_comments":[{"comment":"The statement that \\tilde H_n^{-1} ∘ \\tilde F_n 'is a convex function, interpolating H_n^{-1} ∘ F_n' needs one explicit justification: this composition is the piecewise-linear interpolant of the convex function H^{-1} ∘ F at the increasing knots x_i, and a piecewise-linear interpolant of a convex function with increasing knots is convex because its successive slopes are nondecreasing. Adding this sentence would make the proof easier to follow.","section":"Section 4.1, proof of Lemma 1"},{"comment":"Several displays in the proof of Theorem 1 have lost their superscript formatting. In particular, the expression X_{n-k+1:n} n^{-1-(m-j)} k^{m-j} and the conditions E(h,j,m) and Q(h,j,m) are hard to parse; the exponents should be typeset explicitly so the reader can verify the convergence conditions.","section":"Section 4.2, proof of Theorem 1"},{"comment":"The references to 'Table 2b-(a)' and 'Table 2b-(b)' are confusing; the table should be given distinct sublabels or separate tables. There are also small typos, such as 'well-kown' in Section 2.1 and 'Kologorov-Smirnov' in Section 6, which should be corrected.","section":"Section 5.4 and Section 6"},{"comment":"The constrained test statistic T^{G+}_{m,ℓ,p} is introduced with notation that is not fully defined before first use; please state explicitly that ℓ is the number of selected indices j_k and that the indices are chosen so that the corresponding L-estimators converge under Theorem 1.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the central idea is attractive. The two major issues are both fixable: replace the invalid componentwise-to-joint stochastic-order step in Theorem 2 with the coupling argument already implicit in Lemma 1, and correct the sign in Theorem 2(2). Once those are addressed, the finite-sample claims would be supported by the text. The self-citation to Arab et al. (2025) for Proposition 1 is legitimate and does not amount to circularity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a solid, genuinely new testing method for convex-ordered families (IOR, DOR, DRHR, IHR/DHR), and it fills a real gap—most families in this class had no tests at all. The central idea is simple and sound: if G^{-1}∘F is convex, Jensen gives bounds on P(X ≤ μj:m) in terms of the reference distribution G, and the paper builds a test by comparing empirical non-exceedance probabilities at estimated expected order statistics against those bounds. The heavy-tail extension (a strong law for L-estimators when the mean is infinite, under max-domain-of-attraction conditions) is a genuine contribution, and the simulations are honest: they openly report that Proschan–Pyke has higher power under genuinely monotone hazard rates, while their test handles non-monotone hazard rates and new families. Code is on GitHub, and the real-data example is a nice use case. The citation pattern is mostly the authors' own prior work—the key bound is from their JAP paper—but it is a published, parameter-free Jensen inequality, and the testing machinery built on it is new, so the self-citation does not bother me.\n\nThe soft spots, in proportion. First, the proof of Theorem 2, which carries the finite-sample unbiasedness and monotone-power claims, is not valid as written. It establishes componentwise stochastic inequalities via Lemma 1 and then invokes Theorem 1.A.3 of Shaked–Shantikumar, which requires the coordinates within each vector to be independent. Here the coordinates are all functions of the same empirical CDF. The gap is real and repairable: the coupling inside the proof of Lemma 1 (F_n = H_n∘H^{-1}∘F) gives the inequality for all j simultaneously on the same probability space, so the joint stochastic order follows in one line. But the paper does not present that argument, and a referee should insist on it. Second, the interpolation step in Lemma 1 is terse but actually fine—piecewise linear interpolation of points on a convex function is convex by the chord-slope property—so I would not worry about that one. Third, the constrained heavy-tail test needs the user to know or estimate the tail index α to choose ℓ; the paper says so in Section 5.3, but the abstract's \"broadly applicable\" claim is a bit strong. Fourth, there is a sign-flipped exponent in a displayed integral in the proof of Theorem 1; the stated convergence condition is correct, so this is a typo, not a mathematical error.\n\nNet: the central argument holds up, the proof gap is genuine but clearly fixable, and the paper is honest about its limitations. It deserves a serious referee. I would send it to review and ask for a corrected proof of Theorem 2, the typo fixed, and a short paragraph on how to choose ℓ when α is unknown.","headline":"Solid new tests for convex-ordered families with an honest heavy-tail extension; the finite-sample proof of monotone power has a real but repairable gap that needs a referee's attention.","tokens_in":19310,"tokens_out":13893,"would_cite":true,"duration_ms":109024,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62G20","62G30"],"pacs":[],"model":"deepseek-v4-flash","headline":"A general class of nonparametric tests uses expected order statistics to decide whether an unknown distribution lies in a convex-ordered family relative to a known reference, with unbiasedness, monotone power, and consistency for heavy…","keywords":["convex transform order","expected order statistics","L-estimator","nonparametric test","hazard rate","heavy tails","stochastic order","infinite mean"],"falsifier":"Simulate from a pair $F \\le_c H$ (for example exponential versus Weibull with shape 1.5) and check every sample path for whether $\\widetilde H_n^{-1}\\circ\\widetilde F_n$ is convex between consecutive order statistics; a single sample with a downward kink disproves the Lemma 1 premise, and Monte Carlo can then test whether rejection probabilities under $F$ and $H$ violate the claimed monotonicity.","tokens_in":18328,"feed_emoji":"📊","tokens_out":9641,"duration_ms":79654,"temperature":0.7,"pith_summary":"This paper proposes a general nonparametric test for whether an unknown distribution $F$ belongs to a family defined by the convex transform order with respect to a known reference distribution $G$, such as increasing hazard rate, increasing or decreasing odds rate, and decreasing reversed hazard rate families. The guiding observation is that if $G^{-1}\\circ F$ is convex or concave, then the probability that $X$ falls below a certain expected order statistic is bounded by a known number $\\pi^G_{j:m}$; the paper turns this into a test by estimating the expected order statistics with L-estimators and comparing the empirical non-exceedance probabilities to the bounds. The authors prove the resulting tests are unbiased, consistent, and have power that increases as $F$ moves away from $G$ in the convex order, and they show the consistency survives even when the mean is infinite by proving a strong law for the L-estimators under domain-of-attraction conditions. If the claims are correct, the method supplies the first tests for several recently studied families, works for any absolutely continuous $G$, and does not restrict the supports of $F$ and $G$.","feed_headline":"A single test recipe catches many distribution shape families","feed_subtitle":"Expected order statistics supply the bounds, so the same tests work for heavy tails and any known reference distribution.","key_machinery":"The engine is the identity $\\pi^G_{j:m} = G(E(G^{-1}(B_{j:m})))$, where $B_{j:m}$ is a $\\beta$$(j, m-j+1)$ random variable, together with the Jensen bound of Proposition 1 relating $P(X \\leq \\mu_{j:m})$ to this quantity. The expected order statistics $\\mu_{j:m}(F)$ are estimated by L-estimators, weighted linear combinations of sample order statistics, and Theorem 1 gives an almost-sure strong law for these estimators when $E|X|$ is infinite, under maximum-domain-of-attraction and regular-variation conditions. The finite-sample theory rests on Lemma 1, which uses the linearly interpolated empirical quantile composition to show that $F \\le_c H$ implies $\\widetilde F_n(\\mu_{j:m}(F_n)) \\le_{\\mathrm{st}} \\widetilde H_n(\\mu_{j:m}(H_n))$; this stochastic monotonicity is what carries the unbiasedness and monotone-power results for every sample size.","core_discovery":"The central claim is that, for any absolutely continuous reference distribution $G$, the test statistics $T^{G+}_{m,p}$ and $T^{G-}_{m,p}$ --- the positive and negative parts of the gap between the bounds $\\pi^G_{j:m}$ and the linearly interpolated empirical CDF evaluated at estimated expected order statistics --- solve the problem of testing $H_0: F$ belongs to the location-scale family of $G$ against the alternatives that $F$ is strictly convex- or concave-ordered with respect to $G$. Proposition 1 provides the load-bearing bound: $F \\in \\mathcal{F}^{cx}_G$ implies $P(X \\leq \\mu_{j:m}) \\leq \\pi^G_{j:m}$, and $F \\in \\mathcal{F}^{cv}_G$ implies the reverse inequality. The paper proves that these tests are unbiased and have monotone power at every fixed sample size, and that they are consistent both in the finite-mean case and, through a constrained version that discards non-converging ranks, for heavy-tailed distributions with infinite or undefined mean. It also reports simulations indicating that the tests work for the IHR/DHR, IOR/DOR, and DRHR/IRHR families, including detection of non-monotone hazard rates where the normalized-spacings test can be misleading.","pith_inferences":["A natural extension, not developed in the paper, is to automate the choice of safe ranks by estimating the tail index from the sample (for example with a Hill-type estimator) and then selecting $m$ and $\\ell$ data-adaptively rather than from a prior lower bound on $\\alpha$.","The same bounds could support confidence statements about shape rather than only tests: the distance between $\\pi^G_{j:m}$ and $F(\\mu_{j:m})$ is a population measure of how far $F$ is from the location-scale family, and its L-estimator could be used to construct an effect-size summary.","Because the tests are location- and scale-invariant and do not restrict support, they could be applied iteratively to narrow down the shape of a distribution, as the paper does for river-flow data; this suggests a model-selection workflow where families are accepted or rejected one by one.","The strictness of Jensen's inequality in the consistency proof means the tests distinguish the location-scale family from the rest of the convex-ordered family; they do not by themselves certify membership in the shape class, so a combined use with existing goodness-of-fit tests is the intended reading."],"forward_implications":["The same testing recipe applies to every absolutely continuous reference $G$, including uniform, exponential, negative exponential, log-logistic, Frechet, and Cauchy, with no support restrictions.","Families such as IOR, DOR, and DRHR, for which the paper says no comparable tests currently exist, become testable with guaranteed size control and consistency.","The constrained version of the test remains consistent when the mean is infinite, so a Cauchy or other very heavy-tailed $F$ can still be tested using ranks near the median, such as $\\hat\\mu_{2:3}$.","For the exponential reference, running the IHR and DHR versions together can detect a non-monotone hazard rate: the proposed tests reject in favour of both alternatives, while normalized-spacings tests may incorrectly report a single monotone direction.","The tests complement goodness-of-fit procedures: rejecting $H_0$ while a goodness-of-fit test accepts the shape class provides evidence that $F$ lies in the convex-ordered family outside the location-scale family."],"supporting_citations":[{"why":"Supplies Proposition 1, the Jensen bound on $P(X \\leq \\mu_{j:m})$ for convex-ordered families on which the test statistics are built.","marker":"Arab et al. (2025)"},{"why":"Introduces the convex transform order that defines the families being tested.","marker":"Van Zwet (1964)"},{"why":"Provides the stochastic-order results used to turn Lemma 1 into monotone power and unbiasedness.","marker":"Shaked and Shantikumar (2007)"},{"why":"Gives the strong-law characterizations for linear functions of order statistics used to prove Theorem 1 for infinite-mean cases.","marker":"Mason (1982)"},{"why":"Supplies the almost-sure convergence theorem for L-statistics used for finite-mean consistency.","marker":"Van Zwet (1980)"},{"why":"Provides the domain-of-attraction and regular-variation conditions that define which ranks converge under heavy tails.","marker":"de Haan and Ferreira (2006)"},{"why":"Supplies the regular-variation facts used in the proof of the strong law for the L-estimators.","marker":"Bingham et al. (1989)"},{"why":"Is the established normalized-spacings test for IHR/DHR used as the main comparison baseline in simulations.","marker":"Proschan and Pyke (1967)"}],"fun_headline_variants":["Tests for convex order families via expected order stats","Expected order stats power new distribution tests","Heavy-tail-safe tests for convex-ordered distributions","Unbiased, consistent tests for convex-ordered families","One test framework covers many shape families and heavy tails"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The finite-sample results depend on the claim that linearly interpolating the empirical distribution preserves the convexity of the transformed quantile function in Lemma 1; if a sample can make the smoothed quantile composition non-convex where the true composition is convex, unbiasedness and monotone power no longer follow.","fun_headline_variants_meta":{"raw":{"variants":["Tests for convex order families via expected order stats","Expected order stats power new distribution tests","Heavy-tail-safe tests for convex-ordered distributions","Unbiased, consistent tests for convex-ordered families","One test framework covers many shape families and heavy tails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000674,"raw_usage":{"total_tokens":3112,"prompt_tokens":1036,"completion_tokens":2076,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":2004}},"tokens_in":652,"tokens_out":2076,"duration_ms":13732,"temperature":1.0,"reasoning_tokens":2004,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:23:47.683444+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate from a pair $F \\le_c H$ (for example exponential versus Weibull with shape 1.5) and check every sample path for whether $\\widetilde H_n^{-1}\\circ\\widetilde F_n$ is convex between consecutive order statistics; a single sample with a downward kink disproves the Lemma 1 premise, and Monte Carlo can then test whether rejection probabilities under $F$ and $H$ violate the claimed monotonicity.","supporting_citations":[{"cited_title":", author Lando, T","cited_arxiv_id":null,"evidence_quote":"Supplies Proposition 1, the Jensen bound on $P(X \\leq \\mu_{j:m})$ for convex-ordered families on which the test statistics are built."},{"cited_title":", year 1964","cited_arxiv_id":null,"evidence_quote":"Introduces the convex transform order that defines the families being tested."},{"cited_title":", author Shantikumar, J.G","cited_arxiv_id":null,"evidence_quote":"Provides the stochastic-order results used to turn Lemma 1 into monotone power and unbiasedness."},{"cited_title":", year 1982","cited_arxiv_id":null,"evidence_quote":"Gives the strong-law characterizations for linear functions of order statistics used to prove Theorem 1 for infinite-mean cases."},{"cited_title":", year 1980","cited_arxiv_id":null,"evidence_quote":"Supplies the almost-sure convergence theorem for L-statistics used for finite-mean consistency."},{"cited_title":", author Ferreira, A","cited_arxiv_id":null,"evidence_quote":"Provides the domain-of-attraction and regular-variation conditions that define which ranks converge under heavy tails."},{"cited_title":", author Goldie, C.M","cited_arxiv_id":null,"evidence_quote":"Supplies the regular-variation facts used in the proof of the strong law for the L-estimators."},{"cited_title":", author Pyke, R","cited_arxiv_id":null,"evidence_quote":"Is the established normalized-spacings test for IHR/DHR used as the main comparison baseline in simulations."}],"review_version":1}