{"id":"4ecbf725-71f0-4492-925d-42c63de39c72","arxiv_id":"1908.04328","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A test and graphical device detect whether two convex regression curves are identical up to horizontal and vertical shifts, valid for dependent, non-stationary data.","lead":"This paper builds a statistical test and a plot for deciding whether two regression curves differ only by a shift in the horizontal or vertical direction, while allowing the data to be dependent and non-stationary. It derives the test from the observation that shifted convex curves have equal derivatives after inversion, and validates it with asymptotics and simulations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The implementable bootstrap test's level control is unproven, and the reported null simulations use c<0, outside the paper's c∈(0,1) null, so the advertised controlled level lacks support.","rationale":"The reader's stated weakest assumption is the geometric decay in Assumption 3.1(c). I agree that this is a strong modeling restriction, but it is explicit and internally consistent, so it is a limitation rather than a flaw in the central argument. I also examined the proof of Theorem 3.2 and found that the bound in Eq. (5.11) is algebraically questionable: with pi_n=o(h_d), the displayed O(pi_n^3/h_d^3) bound is not the correct size of the second-order Taylor remainder. However, a careful check shows that the bandwidth conditions in Assumption 3.4, especially the h_d term in the squared sum, are strong enough to control even the crude corrected remainder, so this appears to be a fixable proof blemish rather than a fatal gap. The most load-bearing concern is the bootstrap: the paper's recommended test is Algorithm 4.1, but no theorem justifies its critical values, and the simulation nulls fall outside the formal c∈(0,1) hypothesis. Since the reader already returned a CONDITIONAL verdict, my read does not change the verdict, but it shifts the emphasis from the dependence assumption to the missing bootstrap consistency and the simulation mismatch.","tokens_in":20491,"tokens_out":48146,"duration_ms":467624,"concrete_test":"Run the Table 1 level study with null models that genuinely satisfy c∈(0,1), for example m1(x)=(x-0.4)^2 and m2(x)=(x-0.5)^2 (so c=0.1, d=0), using Algorithm 4.1 with \\w(t) instead of the infeasible w(t). Check whether empirical rejection rates for n=100, 200, 500 at nominal levels 5% and 10% stay within two Monte Carlo standard errors. Separately, derive the conditional weak limit of n1 b_{n,1}^{9/2} W_B under Assumptions 3.1–3.4; if it matches the null limit of the scaled T_{n1,n2}, the bootstrap consistency gap is closed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central practical claim is that Algorithm 4.1 provides a test with controlled asymptotic level. This claim is not backed by a consistency theorem: no result shows that the bootstrap quantile W_{⌊B(1-α)⌋} converges to the null distribution of T_{n1,n2}. The heuristic in Section 4.1 connects W_B to the Gaussian process U_n in the proof of Theorem 3.2, but plugging in estimated m'_s and σ_s requires a separate uniform-in-bandwidth argument that the paper does not supply. Moreover, Algorithm 4.1(b) as written uses the deterministic weight w(t) from (2.16), which depends on unknown a=m'_1(0) and b=m'_1(1-c); taken literally the procedure is infeasible. The simulation study does not fill the gap: the two null models (4.4) and (4.5) both satisfy m1(t)=m2(t-0.1)+d, i.e., c=-0.1, violating the restriction c∈(0,1) in (2.2), and the real-data example has only n1=n2=37. Thus neither theory nor the reported experiments establish the claimed level of the bootstrap test.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the problem of testing whether two convex (or concave) regression functions differ by a horizontal shift c and a vertical shift d, i.e., m1(t)=m2(t+c)+d. A key observation is that under this null hypothesis the derivatives of the inverse derivatives, ((m1')^{-1})' and ((m2')^{-1})', coincide. The authors propose a graphical device based on estimating these quantities by kernel density estimators of local linear derivative estimates, and a formal L2 test statistic T_{n1,n2} constructed from the squared difference of the two estimated curves. For locally stationary, dependent errors, they derive asymptotic normality of n1 b_{n,1}^{9/2} T_{n1,n2} under the null and local alternatives. Because the asymptotic bias and variance are complicated, they propose a bootstrap procedure (Algorithm 4.1) and illustrate it with simulations and a real-data application on infant growth curves.","tokens_in":20751,"tokens_out":5403,"duration_ms":50892,"significance":"The theoretical result in Theorem 3.2, if correct, is a substantial contribution: it gives a nonparametric test for shift-invariant curve comparison under a general locally stationary error structure, going beyond the independence assumptions common in the shape-invariant-model literature. The inverse-derivative characterization in Lemma 2.1 is elegant and leads to a simple graphical tool. The paper also makes good use of existing Gaussian approximation results (Wu and Zhou, 2011) and inverse-regression estimation (Dette et al., 2006). However, the practical test advocated in the paper, the bootstrap in Algorithm 4.1, is not accompanied by a consistency theorem, and the simulation study does not evaluate the null hypothesis as defined because the two null models use a negative shift c, violating the c in (0,1) restriction in (2.2). These gaps substantially weaken the support for the paper's central practical claim of a test with controlled level.","major_comments":[{"comment":"There is no theorem establishing consistency of the bootstrap test. The heuristic preceding Algorithm 4.1 connects W_B to the Gaussian process U_n used in the proof of Theorem 3.2, but no argument shows that the bootstrap quantile W_{⌊B(1−α)⌋} converges to the null distribution of T_{n1,n2}. In particular, replacing the unknown functions m_s' and σ_s in U_n by the estimates \\mhat{m}_s' and \\mhat{σ}_s requires a uniform-in-bandwidth or continuity argument that the paper does not supply. Without such a result, the level control of the bootstrap test is unproven.","section":"Section 4.1, Algorithm 4.1"},{"comment":"As written, the bootstrap statistic W_B uses the deterministic weight function w(t) from (2.16), which depends on the unknown quantities a=m_1'(0) and b=m_1'(1−c). The test statistic T_{n1,n2} in (2.15) uses the estimated weight \\mhat{w}(t). The algorithm does not state how w(t) is to be obtained in practice; taken literally, the procedure is infeasible. The authors should either replace w(t) by \\mhat{w}(t) and prove that the replacement is asymptotically negligible in the bootstrap, or provide a feasible construction of the weight.","section":"Section 4.1, Algorithm 4.1(b)"},{"comment":"The two null models used in the simulation study satisfy m1(t)=m2(t−0.1)+d, i.e., c=−0.1, which violates the assumption c∈(0,1) in the null hypothesis (2.2). The construction of the test, including the estimate \\mhat{c} in (2.9) and the domain of the weight function in (2.16), is explicitly based on c>0. Therefore the reported empirical sizes in Table 1 do not provide evidence that the test controls the level under the null hypothesis as formulated. The simulations should be rerun with c∈(0,1), or the paper should justify that the procedure is invariant under the sign of c.","section":"Section 4.2, models (4.4) and (4.5)"}],"minor_comments":[{"comment":"The paper repeatedly calls c the 'vertical shift' (e.g., in Section 2.1 'let \\mhat{c} be an estimate of the vertical shift c' and in the sentence after (2.10)), but c is the horizontal shift in (2.2). This terminology should be corrected.","section":"Throughout, Section 2"},{"comment":"The text says the bootstrap 'does not require the estimation of the derivatives', but Algorithm 4.1(a) requires estimates of m_1' and m_2' via (2.6), which are then plugged into Ξ_1^{(B)} and Ξ_2^{(B)}. This statement is misleading and should be corrected.","section":"Section 4.1, first paragraph"},{"comment":"Assumption 3.1(b), sup_{0≤t<s≤1} ||G(t,F0)−G(s,F0)||_4<∞, is automatically implied by Assumption 3.1(a) and the triangle inequality; if a stronger smoothness condition is needed, it should be stated explicitly (e.g., a Hölder or Lipschitz condition in t).","section":"Section 3.1, Assumption 3.1(b)"},{"comment":"The reference 'Silverman (1998)' is incomplete; in the reference list it appears only as 'CHAPMAN & HALL, London.' The correct reference is B. W. Silverman, Density Estimation for Statistics and Data Analysis, Chapman and Hall, 1986.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The two major issues are both fixable but require real work: a bootstrap consistency proof and a corrected simulation design. The infeasibility of Algorithm 4.1(b) as written is a straightforward fix but must be addressed before publication. I would not recommend rejection, as the central asymptotic result in Theorem 3.2 is a valuable contribution; however, the paper's advertised practical test currently lacks both theoretical and numerical support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's core identity is elementary—if m1(t)=m2(t+c)+d then ((m1')^{-1})'=((m2')^{-1})'—but the extension to locally stationary dependent errors gives the result real value. The asymptotic normality result in Theorem 3.2 is proved in detail, with the dependence handled via the Wu–Zhou Gaussian approximation, and that part looks solid. The graphical device is simple and the paper honestly frames it as a visual aid. The soft spots are in the gap between the theorem and the advertised bootstrap test. First, there is no consistency theorem for the bootstrap: no result shows that the quantile W_{floor(B(1-alpha))} converges to the null distribution of T_{n1,n2}. The motivation from U_n is heuristic, and plugging in estimated m'_s and sigma_s requires a uniform-in-bandwidth argument that is not supplied. Second, Algorithm 4.1(b) literally uses the deterministic weight w(t) from (2.16), which depends on unknown a=m'_1(0) and b=m'_1(1-c). Taken at face value, the procedure is infeasible; presumably the authors meant hat w(t), but that's not what is written. Third, both null simulation models (4.4) and (4.5) satisfy H0 with c=-0.1, not c in (0,1). The paper says a similar treatment works for negative c, but no details are given, so the level simulations cannot be taken as evidence for the stated null. The real-data example has n1=n2=37, so it is only illustrative. None of this undermines Theorem 3.2 itself, which is a genuine contribution: a shift test for convex regression curves under locally stationary errors, with the proof laid out in detail. But the paper's central practical claim—a bootstrap test with controlled asymptotic level—is not supported by either theory or the reported experiments. This paper deserves a serious referee. A careful referee should ask for a bootstrap consistency theorem or an explicit statement that the bootstrap is heuristic, a corrected weight in Algorithm 4.1, and simulation null models that actually satisfy c in (0,1) (or a precise extension to negative c). I would accept it for peer review, but with major revision expected. I would not cite the bootstrap procedure as it stands, though the asymptotic result may be worth citing once the gaps are addressed.","headline":"The core inverse-derivative idea is elementary but the extension to dependent non-stationary errors makes it worth a look; however, the bootstrap test's level is unproven and the simulation nulls fall outside the paper's own c>0 restriction.","tokens_in":706,"tokens_out":1825,"would_cite":false,"duration_ms":49476,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G08","62G10","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A shift between two convex regression curves reduces to a derivative identity","keywords":["comparison of curves","nonparametric regression","hypothesis testing","shape invariant models","horizontal shift","vertical shift","locally stationary processes","bootstrap test"],"falsifier":"Simulate two samples with convex curves satisfying $m_1(t)=m_2(t+c)+d$, generating errors from a locally stationary process with polynomial decay such as $\\delta_4(k)=O(k^{-2})$ instead of $O(\\rho^k)$, and record the empirical rejection rate of the bootstrap test at nominal 5%; a rate clearly above 5% as $n$ grows would show the geometric-decay assumption is load-bearing.","tokens_in":20274,"feed_emoji":"📈","tokens_out":7015,"duration_ms":65899,"temperature":0.7,"pith_summary":"The paper asks when two nonparametric regression curves differ only by a horizontal shift $c$ and a vertical shift $d$, so that $m_1(t)=m_2(t+c)+d$ for all $t$ in an interval. Its first claim is an equivalence: under this hypothesis the functions $((m_1')^{-1})'$ and $((m_2')^{-1})'$ coincide, and conversely. This turns a shape-invariant comparison into a problem of comparing two curves that can be estimated directly. On that basis the paper constructs a graphical diagnostic and an $L^2$ test statistic, derives its asymptotic normal distribution under local alternatives, and proposes a bootstrap version that avoids estimating the complicated bias and variance. The intended payoff is a shift test that works for dependent and non-stationary errors, where most earlier shape-invariant methodology requires independence.","feed_headline":"A shift between two convex curves reduces to a derivative identity","feed_subtitle":"New test and graphical check work for dependent, non-stationary errors, unlike standard shape-invariant methods.","key_machinery":"The load-bearing object is $f_s=((m_s')^{-1})'$, the derivative of the inverse of the derivative of the regression function. A horizontal shift $c$ in the original curves becomes an additive constant in the inverse derivatives, so the derivative cancels it and $f_1-f_2$ vanishes exactly under the null. The paper estimates $f_s$ without explicit inversion by noting that if $U$ is uniform, $f_s$ is the density of $m_s'(U)$; hence $\\hat f_s(t)=(Nh_{d,s})^{-1}\\sum_{i=1}^N K_d((\\hat m_s'(i/N)-t)/h_{d,s})$, with $\\hat m_s'$ a local linear estimate. The test statistic is the integrated squared difference of $\\hat f_1$ and $\\hat f_2$, and the asymptotic argument combines a Gaussian approximation for locally stationary error processes with a central limit theorem for quadratic forms.","core_discovery":"The central discovery is Lemma 2.1: for regression functions with strictly increasing first derivative, $m_1(t)=m_2(t+c)+d$ on $(0,1-c)$ holds if and only if $((m_1')^{-1})'(u)=((m_2')^{-1})'(u)$ for $u$ in $(m_1'(0),\\,m_1'(1-c))$. The paper then estimates $f_s=((m_s')^{-1})'$ by a kernel density estimate of $\\hat m_s'(U)$ rather than by inverting a non-monotone estimate, and forms the statistic $T_{n_1,n_2}=\\int(\\hat f_1-\\hat f_2)^2\\hat w(t)\\,dt$. Theorem 3.2 shows that under Assumptions 3.1–3.4 and local alternatives with $((m_1')^{-1})'-((m_2')^{-1})'=\\rho_n g+o(\\rho_n)$, $\\rho_n=(n_1 b_{n,1}^{9/2})^{-1/2}$, the standardized statistic $n_1 b_{n,1}^{9/2}T_{n_1,n_2}-B_n(g)$ converges weakly to $N(0,V_T)$, with explicit formulas for $B_n(g)$ and $V_T$ in terms of long-run variances and second derivatives of the regression functions. Because $B_n$ and $V_T$ are hard to estimate, the paper proposes a bootstrap test, whose finite-sample level and power it studies by simulation and on infant growth data.","pith_inferences":["The equivalence may extend to testing whether two monotone regression functions coincide up to any increasing transformation of the covariate axis, not just an additive shift, by comparing derivative-inverse derivatives after a suitable transformation; the paper does not pursue this.","Because the construction only uses monotonicity of $m'$, the same estimator could be adapted to higher-dimensional index sets by comparing densities of gradient images, though the asymptotic theory would need new concentration arguments.","A practical robustness question the paper leaves open is whether the bootstrap still controls level under polynomially decaying dependence; a natural extension would be a block or sieve bootstrap under weaker conditions.","The graphical device could be used as a diagnostic before fitting parametric shape-invariant models, potentially improving model selection in growth-curve and production-function applications."],"forward_implications":["Testing the shift hypothesis does not require first estimating $c$ and $d$; the same statistic and graphical device can be used for any convex regression pair.","The test controls asymptotic level for locally stationary dependent errors, so it applies to time series settings where independence-based shape-invariant tests are not justified.","Under local alternatives the test detects deviations of order $(n_1 b_{n,1}^{9/2})^{-1/2}$, and the asymptotic power is approximately $\\Phi(\\int g^2 w / V_T^{1/2}-z_{1-\\alpha})$.","The bootstrap version avoids direct estimation of the bias term of order $1/\\sqrt{b_{n,1}}$ and the variance involving long-run variances and second derivatives, and the simulation study reports levels close to nominal for $n=100,200,500$.","For concave regression functions the same methodology applies after negating the responses."],"supporting_citations":[{"why":"Supplies the kernel-density estimator of $((m_s')^{-1})'$ via the density of $\\hat m_s'(U)$ and its Riemann approximation.","marker":"Dette et al. (2006)"},{"why":"Provides the uniform bounds for local linear derivative estimates and the consistent long-run variance estimator used in the bootstrap.","marker":"Dette and Wu (2019)"},{"why":"Gives the Gaussian approximation for non-stationary multiple time series that is the foundation of the proof of Theorem 3.2.","marker":"Wu and Zhou (2011)"},{"why":"Defines the locally stationary framework and physical dependence measure used in Assumption 3.1.","marker":"Zhou and Wu (2009)"},{"why":"Supplies the central limit theorem for generalized quadratic forms used to derive the asymptotic normality of the dominant term.","marker":"de Jong (1987)"},{"why":"Provides technical lemmas for sums of weighted normal variables and the variance calculations used in the proof.","marker":"Zhou (2010)"}],"fun_headline_variants":["Curve shifts? Check the derivative inverses.","Detect shifts in convex curves via derivative identity","Nonparametric test for curve shifts under dependent noise","Derivative identity reveals horizontal and vertical shifts","New test for curve shifts handles non-stationary data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes the errors form a locally stationary process whose dependence on a single past innovation decays geometrically fast; if the true errors have only polynomial dependence or are not locally stationary, the claimed null distribution and bootstrap calibration have no stated justification.","fun_headline_variants_meta":{"raw":{"variants":["Curve shifts? Check the derivative inverses.","Detect shifts in convex curves via derivative identity","Nonparametric test for curve shifts under dependent noise","Derivative identity reveals horizontal and vertical shifts","New test for curve shifts handles non-stationary data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1618,"prompt_tokens":1013,"completion_tokens":605,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":533}},"tokens_in":629,"tokens_out":605,"duration_ms":5877,"temperature":1.0,"reasoning_tokens":533,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:44:48.822606+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate two samples with convex curves satisfying $m_1(t)=m_2(t+c)+d$, generating errors from a locally stationary process with polynomial decay such as $\\delta_4(k)=O(k^{-2})$ instead of $O(\\rho^k)$, and record the empirical rejection rate of the bootstrap test at nominal 5%; a rate clearly above 5% as $n$ grows would show the geometric-decay assumption is load-bearing.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the kernel-density estimator of $((m_s')^{-1})'$ via the density of $\\hat m_s'(U)$ and its Riemann approximation."},{"cited_title":"and Wu, W","cited_arxiv_id":null,"evidence_quote":"Provides the uniform bounds for local linear derivative estimates and the consistent long-run variance estimator used in the bootstrap."},{"cited_title":"and Wu, W","cited_arxiv_id":null,"evidence_quote":"Defines the locally stationary framework and physical dependence measure used in Assumption 3.1."}],"review_version":1}