{"id":"cbf7dee6-100c-4fde-9d91-1045f0d1698a","arxiv_id":"2601.16095","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The spherical cardioid family on S^d, previously used only as an auxiliary construction, is given closed-form moments, convolution closure, asymptotically normal estimators, and a bootstrap goodness-of-fit test, with an order-2 fit to long-period comet orbits.","lead":"This paper develops a full statistical toolkit for the spherical cardioid distribution, a rotationally symmetric family on spheres that can be unimodal, multimodal, or axial. It gives formulas for moments, estimators, and a bootstrap goodness-of-fit test, then uses the order-2 model to describe the orientations of long-period comet orbits.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5.3's parametric bootstrap is assumed valid without a theorem or reference; if it fails under the composite null, the comet p-values and the C2 fit claim are unsupported.","rationale":"The paper's distribution-theoretic contributions—density, moments (Thm 3.1), convolution (Prop 3.1), characteristic function, simulation, and estimator asymptotics—appear correct and are supported by detailed proofs. The vectorized moment computations are lengthy but internally consistent. The ML and MM asymptotic normality theorems use standard delta-method/van der Vaart conditions and are credible. The one place where an inferential claim outruns the mathematics is the GOF test's bootstrap calibration. A composite null with estimated parameters requires a proof or citation (e.g., Beran's parametric bootstrap theory); the paper supplies neither. Table 1's size simulations are too limited to substitute for this. Since the comet application's headline conclusion (C2 fits long-period comet normals) rests directly on these bootstrap p-values, this is load-bearing. The reader's other points—post-hoc k selection and the minor typos in Table 2/Sec 6.1—are secondary. I agree with the reader's conditional acceptance, contingent on addressing the bootstrap justification.","tokens_in":49880,"tokens_out":8965,"duration_ms":84879,"concrete_test":"Run a Monte Carlo study under the composite null: for k=1,2, d=1,2, and ρ∈{0.25,0.5,0.75}, n=100,200, simulate S=5000 samples from C_k(μ,ρ) with fixed μ. For each sample, compute the bootstrap p-value via Algorithm 3 (using the exact P_n^{CvM,Unif} or Monte Carlo with K=50 as in the paper) with B=1000. Test whether the p-values are uniformly distributed using, e.g., a Kolmogorov–Smirnov test. If the p-values deviate significantly from uniform for any configuration, the bootstrap is not valid; if uniform, it provides empirical support (though not a proof) for the procedure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central methodological gap is the unproven validity of the parametric bootstrap in Algorithm 3 (Sec. 5.3). The test statistic P_n^{W,λ} uses the estimated projected cdf F̂_γ with parameters estimated from the same sample, so the null distribution of the statistic is not the simple-hypothesis distribution. The paper merely states the procedure is 'standard' and provides no theorem or reference ensuring that the bootstrap distribution of P_n^{*,b} under C_k(μ̂,ρ̂) converges to the null distribution. Without this, the p-values in Table 3 for the comet data—and the conclusion that C2 is an adequate model—are not justified. The empirical size checks in Table 1 cover only n=100 and few (ρ,k,d) values, and even those show deviations; they do not establish asymptotic validity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the spherical cardioid distribution C_k(μ,ρ) on S^d, defined by density f_{C_k}(x;μ,ρ)=ω_d^{-1}{1+ρ\\tilde C_k^{(d-1)/2}(x^⊤μ)}, as a higher-dimensional, higher-order generalization of the circular cardioid. It establishes tractability properties: rotational symmetry and various shape regimes (Sec. 3.1), closedness under convolution (Prop. 3.1), explicit vectorized moments whose low orders coincide with the uniform moments (Thm 3.1 and Cor. 3.1–3.2), characteristic/moment generating functions (Prop. 3.2), and two simulation algorithms (Algs 1–2). Estimation is treated by the method of moments for k=1,2 (Thms 4.1–4.2), a Gegenbauer-moment estimator for known location (Thm 4.3), and maximum likelihood with asymptotic normality (Thm 4.4), together with asymptotic relative efficiencies (Sec. 4.3). The second half develops a projected-ecdf goodness-of-fit test (Sec. 5): explicit projected density and cdf (Thm 5.1), V-statistic forms for general and Anderson–Darling weights (Thm 5.2, Cor. 5.1), closed-form kernels for the Cramér–von Mises statistic (Thm 5.3), and a parametric bootstrap procedure (Alg. 3). Numerical experiments and an application to long-period comet orbital normals suggest that C_2 provides an adequate model. The principal inferential claim for the goodness-of-fit test is the validity of the parametric bootstrap under estimated parameters, for which the paper provides no theorem or reference.","tokens_in":50129,"tokens_out":12518,"duration_ms":115022,"significance":"If the results hold, the spherical cardioid family is a genuinely useful addition to directional statistics: a simple one-concentration family on S^d that is close to uniformity, has closed-form density, moments, and characteristic function, and is easy to simulate. The moment formulas and the asymptotic distributions of the estimators are derived in detail, with explicit variance expressions; the ARE analysis is careful and informative. The projected-ecdf test extends a prior uniformity-testing framework to a parametric family and provides closed-form test statistics for several important cases. However, the paper's most novel inferential tool — the bootstrap goodness-of-fit test — is presented without a formal validity argument, and one of the closed-form results depends on unshown computer algebra. These gaps currently prevent full confidence in the reported comet-data p-values and in the claim that C_2 is an adequate model for long-period comet orbital normals. If the bootstrap validity is established (or the application is recast as exploratory), and the computational derivations are made verifiable, the paper would be a solid contribution to the field.","major_comments":[{"comment":"The parametric bootstrap is load-bearing for the goodness-of-fit conclusion. The statistic P_n^{W,λ} uses the projected cdf F̂_γ with parameters estimated from the same sample; its null distribution is therefore not the simple-hypothesis distribution. The paper states in Sec. 5.3 that the bootstrap procedure is 'standard' but gives no theorem, proof, or citation ensuring that bootstrap samples from C_k(μ̂,ρ̂) approximate the null distribution of P_n^{W,λ}. This is particularly important because the comet p-values in Table 3 — and the claim that C_2 is adequate for long-period comet orbital normals — rest on this procedure. The empirical size checks in Table 1 are limited to n=100 and a small grid of (ρ,k,d); several entries lie well outside the 95% prediction interval (e.g., k=1,d=2,ρ=0.75 rows, with rejection rates 1.4–2.6% against a 5% nominal level). These checks do not substitute for","section":"Sec. 5.3, Algorithm 3; Table 3"},{"comment":"The closed-form kernels φ and ψ in Theorem 5.3 are a stated contribution and are used in the exact evaluation of P_n^{CvM,Unif}. The proof in Appendix C.2 delegates parts of the integral evaluations to Mathematica (e.g., the expressions for φ̃^(1) and φ̃^(2) in the proof of Theorem 5.3), without showing the symbolic derivations or providing an independently checkable verification. A referee cannot easily confirm that the formulas are correct, and a dependence on unshown computer algebra is especially delicate in a statistics paper where these formulas feed into test implementations and numerical experiments. Please either provide derivations (or at least a clear, verifiable reduction to standard integrals) or make the computer algebra notebook/code available, and state explicitly that the results have been independently checked.","section":"Sec. 5.2, Theorem 5.3; Appendix C.2"},{"comment":"The text states: 'The bias for μ_1 has order 10^{−2}, but for ρ the estimated bias ρ̄̂−ρ is still significant: approximately 0.20 for k=1 and 0.33 for k=2.' As written, this claims a bias of order 0.2–0.3 in ρ̂ at n=1000 with true ρ=0.5, which is inconsistent with the strong consistency and asymptotic normality established in Theorems 4.1–4.4. If these are biases of the standardized statistics √n(ρ̂−ρ), that should be stated explicitly and the values are plausible (they correspond to original-scale biases of order 10^{−2}). If they are not, the statement is erroneous and needs correction. This ambiguity undermines the numerical validation in Section 6.1.","section":"Sec. 6.1"}],"minor_comments":[{"comment":"The caption defines the null order as k_0=(k+1) mod 1, which is always 0. The text above the table says (k,k0) ∈ {(1,2),(2,1)}, so the formula is presumably a typo; please correct it.","section":"Sec. 6.2, Table 2 caption"},{"comment":"The statement that the variable performance in Table 1 'can be explained by the fact that the Monte Carlo samples are shared within each row' is not a convincing explanation for the systematically low rejection rates in the (k=1,d=2,ρ=0.75) row. Shared random numbers explain correlation, not a level shift away from 5%. Please discuss or investigate whether this reflects conservativeness of the bootstrap test at n=100.","section":"Sec. 6.2"},{"comment":"The sentence 'The inverse transformation method is rarely preferable over rejection sampling besides k=1,2' should read '...except for k=1,2'.","section":"Sec. 3.5"},{"comment":"In the paragraph before Eq. (24), the notation for the Anderson–Darling statistic uses U_(i) but the definition of U_(i) is given later as ordered values of U_i^{(γ)}=F̂_γ(γ^⊤X_i). This is clear enough, but a parenthetical reminder would help.","section":"Sec. 5.2"},{"comment":"The y-axis is labeled 'ARE' without indicating which estimator (μ or ρ) is being plotted; the captions clarify, but the axis itself could be more informative.","section":"Figure 3"},{"comment":"The parenthetical '(see gray dashed lines)' in Sec. 6.1 is ambiguous because the gray dashed lines in Figure 5 represent the empirical mean of the standardized statistics, not a bias in the original scale; please align the text with the figure.","section":"Sec. 4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is carefully written and the moment/estimation derivations are detailed and internally consistent. The main risk is the unproven parametric bootstrap. I would be willing to accept after a proper theorem or reference is supplied for the bootstrap validity, and after the Mathematica-dependent derivations are made verifiable. The bias statement in Sec. 6.1 also needs correction. None of these issues appear to be fatal to the central distribution theory, but they are load-bearing for the goodness-of-fit test and the application."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Thanks for the report. I read the paper too. The core distribution-theoretic work is real: the moments, convolution closure, simulation algorithms, and estimator asymptotics are detailed, and I did not find a load-bearing error. The author also deserves credit for clearly citing Baringhaus–Grübel and Borodavka–Ebner; the novelty claim is modest but honest—nobody had treated this as a parametric model. The projected cdf and the closed-form CvM/AD kernels are a genuine service, even if the derivations lean on Mathematica.\n\nThe soft spots are where your report puts them. The parametric bootstrap in Algorithm 3 is the clear gap. The paper calls the procedure standard and gives no theorem or reference establishing that resampling from C_k(μ̂, ρ̂) yields valid p-values under the composite null. Table 1's size checks are only n=100 with B=100, and the conservative behaviour at ρ=0.75 shows the procedure is not obviously safe. I would want either a proof or a citation to a general bootstrap validity result for this kind of statistic before relying on the comet p-values.\n\nThe other issues are smaller. Theorem 5.3's kernel formulas are asserted after \"Mathematica\", with no derivations or code; that is hard to check and should be made reproducible. The Table 2 caption formula k0=(k+1) mod 1 is simply wrong, and the Section 6.1 bias statement about ρ is misworded—the bias in the standardized scale is not the actual bias. The comet analysis picks k=2 after looking at k=1,2,3,4, with no multiplicity correction; that is worth a sentence of caution, though not a reason to reject the application.\n\nNet: the paper is a competent, useful addition to directional statistics. It should go to review. But the reviewer should push for a bootstrap validity argument or reference, and for reproducible derivations of the kernels. If those land, the GOF test becomes trustworthy. With the current gaps, I would treat the comet conclusions as suggestive rather than established.","headline":"Solid, honest development of the spherical cardioid as a working model; the distribution theory is competent, but the bootstrap GOF test lacks a validity proof and the application over-reads k=2.","tokens_in":50582,"tokens_out":2447,"would_cite":true,"duration_ms":24883,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H11","62F03","62F12","33C45"],"pacs":[],"model":"deepseek-v4-flash","headline":"The spherical cardioid distribution is a fully tractable spherical model: closed under convolution, with explicit moments, estimators, and a bootstrap goodness-of-fit test that fits long-period comet orbital normals with k=2.","keywords":["spherical cardioid distribution","directional statistics","Gegenbauer polynomials","Chebyshev polynomials","moment estimation","maximum likelihood","goodness-of-fit","projected cumulative distribution function"],"falsifier":"Simulate iid samples from $C_2(\\mu,0.5)$ on $S^2$ with $n=100$, estimate $(\\mu,\\rho)$, run Algorithm 3 with $B=1000$, and check whether the resulting p-values are $\\mathrm{Uniform}(0,1)$ under the null; substantial miscalibration would make the reported comet p-values uninterpretable.","tokens_in":49767,"feed_emoji":"🌐","tokens_out":4514,"duration_ms":39401,"temperature":0.7,"texified_at":"2026-08-05T20:48:52.509067+00:00","pith_summary":"This paper develops the spherical cardioid distribution, a family of distributions on the sphere defined as a uniform density plus a Gegenbauer-polynomial perturbation, and argues that it is a fully tractable statistical model. The paper establishes closedness under convolution, explicit vectorized moments (with all moments of order below the model order matching the uniform distribution), a closed-form characteristic function, efficient simulation algorithms, and consistent asymptotically normal moment and maximum-likelihood estimators. It then constructs a parametric-bootstrap goodness-of-fit test based on projected empirical distribution functions, with closed-form test statistics in low dimensions. Applied to orbital normals of long-period comets, the order-two spherical cardioid fits the data, while uniformity and orders one, three, and four are rejected. A sympathetic reader would care because this provides a simple, analytically workable alternative to the uniform distribution for mildly non-uniform spherical data.","texify_model":"deepseek-v4-flash","texify_usage":{"total_tokens":7099,"prompt_tokens":743,"completion_tokens":6356,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":743,"completion_tokens_details":{"reasoning_tokens":5679}},"feed_headline":"Spherical cardioid family gains explicit fit and tests","feed_subtitle":"Moments, estimators, and a bootstrap goodness-of-fit test are derived; order-two model fits long-period comet orbits.","key_machinery":"The key object is the normalized Gegenbauer polynomial $\\tilde C_k^{(d-1)/2}(x^T \\mu)$ (with Chebyshev polynomials for $d=1$), used as a perturbation of the uniform density on the sphere; its orthogonality and parity supply the moment-vanishing structure and the convolution formula. The estimators rely on the vectorized moments, and the goodness-of-fit test builds on the closed-form projected cumulative distribution function $F_\\gamma$, which is the uniform projected cdf plus a $\\rho$-times-polynomial term. The bootstrap in Algorithm 3 is the mechanism that turns the test statistic into p-values.","core_discovery":"The central discovery is that the family $C_k(\\mu,\\rho)$ on $S^d$ — densities proportional to $1 + \\rho \\tilde C_k^{(d-1)/2}(x^T \\mu)$, where $\\tilde C$ is the normalized Gegenbauer (Chebyshev when $d=1$) polynomial — is closed under convolution in a simple sense, has vectorized moments computable in closed form, and admits moment and maximum-likelihood estimators with explicit asymptotic variances. A notable structural fact is that for any model of order $k$, all moments of order $m<k$ coincide with the uniform sphere's moments, as do moments of orders $m>k$ with $m-k$ odd; this makes high-order spherical cardioids nearly indistinguishable from uniformity by low-order moment information. The paper also derives the exa","pith_inferences":["A likely use of the moment-vanishing property is to construct high-order spherical cardioids as explicit alternatives for power studies of uniformity tests; the paper notes this as a challenge, but it is a direct consequence of its moment theorem.","The bootstrap test's validity under composite nulls is assumed rather than proved; if one filled that gap (e.g., via a conditional convergence argument or a corrected bootstrap), it would solidify the reported comet p-values.","The closed-form projected cdf could support other goodness-of-fit statistics (e.g., characteristic-function or energy distances) for the same family, an extension the paper discusses as future work.","The convolution closure may allow tractable mixtures or Bayesian hierarchical priors on S^d, since the location parameter stays within the same family."],"forward_implications":["For k=1 and k=2, the method-of-moments and maximum-likelihood estimators are strongly consistent and asymptotically normal, with explicit variances; the asymptotic relative efficiency formulas show the moment estimator of concentration loses efficiency for large |ρ|.","The convolution closure means that mixing the location of one spherical cardioid by another spherical cardioid of the same order yields a spherical cardioid with a product-type concentration divided by a known dimension factor—useful for hierarchical modeling.","All moments of order < k match the uniform sphere, so a large-k spherical cardioid is a near-uniform alternative that is difficult for standard low-order uniformity tests to detect.","The projected cdf is explicit, so the paper's CvM and AD statistics have closed-form V-statistic expressions in d=1,2 for k=1,2, avoiding numerical integration in those cases.","For long-period comet orbital normals, the order-two cardioid is not rejected at the 10% level, whereas uniformity and orders 1,3,4 are rejected; for short-period comets every order is rejected."],"fun_headline_variants":["Spherical cardioid: explicit moments, estimators, and fit tests","Cardioid on sphere: closed convolution and bootstrap tests","Spherical cardioid: moment-based estimators and GOF bootstrap","Generalized spherical cardioid: properties and goodness-of-fit"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The goodness-of-fit test relies on the unproven assumption that the parametric bootstrap in Algorithm 3 yields valid p-values when the null parameters are estimated from the same sample; the paper states the procedure as standard and provides no theorem showing the bootstrap distribution approximates the null distribution.","fun_headline_variants_meta":{"raw":{"variants":["Spherical cardioid: explicit moments, estimators, and fit tests","Cardioid on sphere: closed convolution and bootstrap tests","Spherical cardioid: moment-based estimators and GOF bootstrap","Generalized spherical cardioid: properties and goodness-of-fit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000664,"raw_usage":{"total_tokens":2857,"prompt_tokens":717,"completion_tokens":2140,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":461,"completion_tokens_details":{"reasoning_tokens":2071}},"tokens_in":461,"tokens_out":2140,"duration_ms":13963,"temperature":1.0,"reasoning_tokens":2071,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T08:38:36.008118+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate iid samples from $C_2(\\mu,0.5)$ on $S^2$ with $n=100$, estimate $(\\mu,\\rho)$, run Algorithm 3 with $B=1000$, and check whether the resulting p-values are $\\mathrm{Uniform}(0,1)$ under the null; substantial miscalibration would make the reported comet p-values uninterpretable.","supporting_citations":[],"review_version":1}