{"id":"3451c012-27de-4589-98ac-0cd4a94bc81b","arxiv_id":"2602.11467","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"PRISM learns a neural Gaussian field of shape displacements conditioned on age and uses Fisher information to estimate local temporal uncertainty, but the Cramér–Rao justification confuses latent age with an estimator.","lead":"PRISM models 3D anatomical shapes as a neural probability field whose mean and variance change with age, and uses Fisher information to estimate how precisely each body region reveals a child's developmental stage. The paper claims this gives closed-form, spatially varying uncertainty for tasks like predicting airway growth and detecting narrow airways.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Cramér–Rao bound is applied to the latent variable τ as if it were an estimator computed from d; in the paper's own hierarchical generative model this is invalid, so Eq. (8) and the temporal-uncertainty interpretation do not follow.","rationale":"The paper's central theoretical contribution is the closed-form Fisher Information metric for temporal uncertainty, culminating in σ²τ≈1/I (Eq. 11) and the CR lower bound Var(τ|p,t)≥1/I (Eq. 8). The reader's weakest assumption correctly identifies the load-bearing flaw: the CR theorem requires an unbiased estimator that is a function of the observed data, but τ is a latent variable in the generative model. The proof in Appendix A.2.4 Step 2 computes E[τU]=1 by treating τ as a fixed function of d; when τ is latent, the covariance between τ and the score is not 1, and the bound fails. The simple normal-normal counterexample is decisive: Var(τ|t)=γ² can be smaller than 1/I=γ²+σ², so the claimed inequality is false. This is not a disagreement with consensus or a matter of external validity; it is an internal inconsistency in the theoretical derivation. The Gaussian Fisher-information algebra in A.2.1–A.2.3 is correct, and the noiseless synthetic experiments match a special case where 1/Iμ happens to equal the true variance, which explains why the empirical validation appears successful. However, that special case masks the theoretical error, and the clinical uncertainty maps are not independently validated. I considered whether the discarded IΣ term or the lack of code/error bars might be more central, but the category error in applying CR to a latent variable is the most fundamental: it invalidates the meaning of σ²τ as a bound on population temporal variability. The framework may still be salvageable as a heuristic if reframed as estimation variance of the inverse encoder, but the stated theoretical contribution should be rejected. Hence the reader's REJECT verdict stands unchanged.","tokens_in":19898,"tokens_out":6700,"duration_ms":69782,"concrete_test":"Instantiate the paper's hierarchical generative model in the simplest case: d|τ∼N(τ,σ²), τ|t∼N(t,γ²). Compute (i) Var(τ|t)=γ² and (ii) the Fisher information of the marginal d|t∼N(t,γ²+σ²), I=1/(γ²+σ²). Eq. (8) claims γ²≥γ²+σ², which is false for σ²>0. If desired, repeat on Starman(G) by adding zero-mean noise to the displacement after sampling τ and compare PRISM's σ²τ from Eq. 11 against the true στ²(t); under the paper's bound, σ²τ should not exceed the true variance, but the calculation predicts σ²τ≈στ²+σ²/||∂τμ||².","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 (Eq. 8) and Appendix A.2.4 apply the Cramér–Rao bound to the latent variable τ by assuming E_{p(d|p,t)}[τ]=t and then asserting Cov(τ,U)=1. This step is valid only if τ is a deterministic function of the observed displacement d. In PRISM's own generative story, d is generated from τ (d∼N(μ(p,τ),Σ(p,τ))), so τ is a latent random variable, not an estimator computed from d; p(d|p,t) is the marginal over τ, and the proof moves τ outside the d-integral without justification. Under the hierarchical model d|τ∼N(τ,σ²), τ|t∼N(t,γ²), the true variance of τ is Var(τ|t)=γ², while the Fisher information of the marginal d|t∼N(t,γ²+σ²) is I=1/(γ²+σ²). Thus 1/I=γ²+σ²>γ², so Eq. (8) is false. The paper's synthetic validation is noiseless and nearly linear, where 1/Iμ approaches the true στ², so the experiments do not exercise the assumption. Additionally, even if τ were an estimator, dropping IΣ (Eq. 9→10) means 1/Iμ is not the CR bound for the full model; it is a heuristic upper estimate. The clinical claims of 'interpretable temporal uncertainty' therefore rest on an invalid theorem, not merely an approximation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"PRISM models the conditional distribution of 3D shape displacement d at template point p given covariate t as a heteroscedastic Gaussian N(μ(p,t),Σ(p,t)), with μ and Σ predicted by an MLP. An inverse encoder g maps (p,d) to an estimated intrinsic time τ̂. The paper's central theoretical claim is that the Fisher information of this Gaussian yields a closed-form local temporal uncertainty σ²_τ(p,t)≈1/I(p,t), interpreted as a Cramér–Rao lower bound on the population variance of the latent intrinsic time τ. Experiments on synthetic Starman(G/L), ANNY, and a pediatric airway dataset evaluate mean reconstruction, intrinsic-time estimation, personalized prediction, and OOD detection.","tokens_in":20373,"tokens_out":9594,"duration_ms":96328,"significance":"The paper targets a real gap: spatially resolved, covariate-conditioned uncertainty for statistical shape analysis. The framework is broad, the experiments are extensive, and the closed-form Gaussian Fisher information calculation in Appendix A.2 (Eqs. 68–69) is correct. The synthetic datasets with known ground-truth temporal uncertainty are a strength, as is the idea of an amortized inverse encoder. If the Cramér–Rao-based interpretation were valid, the contribution would be significant. However, the central theoretical step is not valid: the Cramér–Rao bound is applied to a latent variable rather than to an estimator, and the covariance term is discarded in a way that destroys the bound. The reported uncertainty quantities are therefore not the claimed population variances, and the clinical interpretability claims are unsupported. The empirical results may still be useful, but the paper's main contribution as stated does not hold.","major_comments":[{"comment":"The proof of Eq. (8) assumes E_{p(d|p,t)}[τ]=t and differentiates to obtain E[τ U]=1. This is only valid if τ is a function of the observation d. In the paper's model, d is generated from τ (d|p,τ∼N(μ(p,τ),Σ(p,τ))), so τ is a latent random variable, not a statistic computed from d. The differentiation step in Eq. (72) is therefore unjustified. Concretely, in the hierarchical model d|τ∼N(τ,σ²), τ|t∼N(t,γ²), the marginal is d|t∼N(t,σ²+γ²), so I=1/(σ²+γ²) and Var(τ|t)=γ²<1/I, contradicting Eq. (8). The synthetic validations use noiseless, nearly monotone displacements where 1/Iμ approximately equals Var(τ|t), so they cannot detect this failure.","section":"Sec. 4.3, Eq. (8); Appendix A.2.4"},{"comment":"The full Fisher information of the Gaussian is Iμ+IΣ with IΣ≥0. The paper discards IΣ and redefines I(p,t):=Iμ. This is not a Cramér–Rao lower bound: the full-model bound is Var(τ)≥1/(Iμ+IΣ), which does not imply Var(τ)≥1/Iμ because 1/(Iμ+IΣ)≤1/Iμ. Orthogonality of mean and covariance in the Fisher–Rao metric justifies additivity of the two information terms, not omission of one term from the bound. Thus Eq. (11) at best defines a heuristic “temporal discriminability” measure; interpreting 1/Iμ as a lower bound on the population variance of intrinsic time is incorrect.","section":"Sec. 4.3, Eqs. (9)–(11)"},{"comment":"The inverse encoder g is trained on triplets (p,d,τ) with d=μ(p,τ) sampled from the learned forward model f. The clinical calibration plot in Fig. 4 checks whether observed points lie inside intervals derived from the same f. This is a self-consistency check between g and f, not validation against an independent ground truth. The paper acknowledges that no ground-truth intrinsic time exists for the airway dataset (Sec. 5.1.1), but the sentence “validating the calibration of our uncertainty estimates” (Sec. 5.2) overstates the evidence. Independent biological annotations or a held-out longitudinal cohort would be needed.","section":"Sec. 4.2, Eq. (5); Sec. 5.2, Fig. 4"}],"minor_comments":[{"comment":"The notation g(q,p) should be g(q,d_q) (or the point q's observed displacement); as written it suggests the encoder takes p as input.","section":"Sec. 4.4, Eq. (17)"},{"comment":"Table 3 reports only PRISM results; the claim that baselines cannot perform local time estimation should be stated in the main text, and the experimental protocol for baselines in this task should be described.","section":"Table 3"},{"comment":"Eq. (14) defines a global z-score using pointwise σ_τ(t0); clarify how the pointwise uncertainties are aggregated.","section":"Sec. 4.4, Eq. (14)"},{"comment":"The derivation of T3 uses symmetry of A and Σ; state this explicitly for readability.","section":"Appendix A.2.3, Eqs. (62)–(63)"},{"comment":"Figure 7 is referenced in Sec. 5.2 but appears in Appendix B.6; add a cross-reference.","section":"Sec. 5.2 / Appendix B.6"},{"comment":"The statement “The code will be open to public” should include a repository URL or a clear availability statement.","section":"Abstract"}],"recommendation":"reject","confidential_remarks":"To the editor: The empirical framework is well executed and could form the basis of a useful systems paper, but the central theoretical claim is demonstrably incorrect; the Cramér–Rao interpretation of Eq. (11) is not fixable by a local edit without changing the paper's main contribution. I recommend rejection rather than major revision, unless the authors wish to resubmit a substantially revised manuscript in which σ_τ is presented as an empirical uncertainty heuristic with explicit validation on independent outcomes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth reading for the PRISM framework itself: a heteroscedastic Gaussian implicit field for covariate-conditioned shape modeling, plus an amortized inverse encoder for local intrinsic-time estimation. That combination is genuinely new and practically attractive, and the experiments are extensive—three synthetic datasets plus pediatric airway data, with reconstructions, time estimation, forecasting, and OOD detection. PRISM often beats the baselines, and the appendix algebra for the Gaussian Fisher information is correct as far as it goes.\n\nThe soft spot is load-bearing and sits right at the center of the theory. Section 4.3 and Appendix A.2.4 apply the Cramér–Rao bound to the latent variable τ as if τ were an unbiased estimator computed from d. But in the paper's own generative story, d is generated from τ (d ~ N(μ(p,τ), Σ(p,τ))). The proof differentiates E[τ] under p(d|p,t) as though τ were a fixed function of d, independent of t. That only holds in the noiseless, invertible case. The simple counterexample d|τ~N(τ,σ²), τ|t~N(t,γ²) gives Var(τ|t)=γ² while 1/I=γ²+σ², so the claimed inequality is false in general. The synthetic experiments are noiseless and nearly linear, exactly the regime where the bound happens to be tight, so they never exercise the assumption. The paper also drops IΣ without flagging that 1/Iμ is not the CR bound for the full model; that's a further gap.\n\nThe inverse encoder is also trained on synthetic triplets generated from the forward model, so the intrinsic-time evaluations largely measure consistency between g and f, not agreement with independent ground truth. That is a weaker validation than the presentation suggests, though the Starman and ANNY results do compare against known τ during evaluation.\n\nThat said, I don't think this should be a desk reject. The empirical framework is solid and potentially salvageable—if the temporal-uncertainty interpretation is reframed as a heuristic, or replaced with a correctly derived posterior variance (e.g., deconvolution or variational inference), the rest of the paper stands. The clinical OOD results are interesting, and the local intrinsic-time idea has clear value. A good reviewer could push the authors to fix the theory or cut it down to an approximation with appropriate caveats. My own assessment: the current version needs major revision before it makes a sound theoretical claim, but it deserves a serious referee rather than a desk rejection.","headline":"The PRISM framework is useful and the experiments are thorough, but the central CR-bound argument for temporal uncertainty misapplies the bound to a latent variable and does not stand up as stated.","tokens_in":20903,"tokens_out":6129,"would_cite":false,"duration_ms":61916,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PRISM claims that a heteroscedastic Gaussian displacement field, learned by an implicit neural network, can describe population shape evolution and yield pointwise developmental-uncertainty maps in closed form through a Fisher-information i","keywords":["implicit neural representation","statistical shape modeling","Fisher information","intrinsic time","heteroscedastic uncertainty","pediatric airway","anomaly detection","Cramér–Rao bound"],"falsifier":"Simulate a dataset exactly from PRISM's generative assumption: draw τ|t ∼ N(t, γ²), then d|τ ∼ N(μ(τ), Σ(τ)); fit the model and compare the empirical variance of recovered τ at fixed t with 1/I. In the one-dimensional Gaussian case the bound predicts Var(τ|t) ≥ γ² + σ², but the generative truth is Var(τ|t) = γ², so the comparison will refute the bound unless τ is reconstructed deterministically from d.","tokens_in":19728,"feed_emoji":"📐","tokens_out":5554,"duration_ms":59415,"temperature":0.7,"pith_summary":"PRISM is trying to establish that a single probabilistic implicit neural field can give a closed-form, spatially continuous estimate of how much a person's developmental stage varies at every point on the anatomy, conditioned on age. It models the displacement from a shared template as a heteroscedastic Gaussian with learned mean μ(p,t) and covariance Σ(p,t), then defines temporal uncertainty as the reciprocal Fisher information of the mean trajectory, 1/((∂μ/∂t)ᵀΣ⁻¹(∂μ/∂t)). A sympathetic reader would care because this turns the population distribution of shapes into a directly queryable uncertainty map, enabling intrinsic-age estimation, personalized growth prediction, and local anomaly detection without Monte Carlo sampling or per-patient optimization. The paper validates the idea on synthetic data with known ground truth and on pediatric airway CT scans, where the predicted bands align with observed variation.","feed_headline":"Closed-form metric puts time uncertainty on each surface point","feed_subtitle":"PRISM derives pointwise estimates of how much a child's airway development lags or leads chronological age.","key_machinery":"The key object is the reciprocal Fisher information, σ²τ(p,t) = 1/((∂μ/∂t)ᵀ Σ⁻¹ (∂μ/∂t)), a closed-form, pointwise temporal-uncertainty measure obtained from derivatives of the learned mean with respect to time. It works by converting the heteroscedastic Gaussian likelihood into an estimation-theoretic bound on latent-time variance, and it is paired with an amortized inverse encoder g(p,d) that predicts intrinsic time from a local displacement in a single forward pass.","core_discovery":"The central discovery is that the conditional distribution of a 3D displacement at any template point and age can be modeled as a Gaussian with network-predicted mean μ(p,t) and covariance Σ(p,t), and that the variance of a subject's latent developmental stage τ is then expressed, in closed form, as the reciprocal of the mean-trajectory Fisher information: σ²τ(p,t) ≈ 1/((∂μ/∂t)ᵀ Σ⁻¹ (∂μ/∂t)). The full Fisher information also contains a covariance-evolution term, but the paper argues that term measures how population diversity changes with time rather than how an individual localizes along the mean trajectory, so it is dropped. The result is a spatially heteroscedastic uncertainty field that","pith_inferences":["The formula σ²τ ≈ 1/I is a Cramér–Rao lower bound, not an equality; comparing the predicted band to the empirical spread of inferred τ across subjects at a fixed age would show whether the approximation is conservative or optimistic in real data.","By discarding the covariance term IΣ, the uncertainty map reflects only how fast the mean shape changes; regions whose variability changes with age but whose mean is static would be reported as fully certain, which may miss a clinically relevant signal.","The single-covariate setup suggests a natural extension: for covariates such as age plus sex or height, replace the scalar Fisher information with a Fisher information matrix and a multivariate Cramér–Rao bound, and the same closed-form derivation should carry through."],"forward_implications":["Spatially continuous, resolution-independent uncertainty maps on anatomy become available without Monte Carlo resampling.","Per-point intrinsic time can be estimated in one forward pass of the inverse encoder, avoiding test-time optimization.","Personalized longitudinal prediction follows from assuming a subject's temporal z-score stays constant and propagating along the learned mean trajectory.","Anomaly detection can be localized by comparing each region's intrinsic time to the most developmentally advanced region within the same anatomy."],"fun_headline_variants":["Closed-form Fisher info gives pointwise time uncertainty","Each surface point gets its own temporal variance","PRISM: shape evolution with interpretable uncertainty","Closed-form metric localizes growth timing per point"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The temporal-uncertainty formula only holds if a subject's latent developmental stage τ can be treated as an unbiased estimator of chronological age t that is a function of the observed displacement; the generative model does not enforce this, and when τ has its own variance the claimed bound can fail.","fun_headline_variants_meta":{"raw":{"variants":["Closed-form Fisher info gives pointwise time uncertainty","Each surface point gets its own temporal variance","PRISM: shape evolution with interpretable uncertainty","Closed-form metric localizes growth timing per point"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000252,"raw_usage":{"total_tokens":1364,"prompt_tokens":679,"completion_tokens":685,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":423,"completion_tokens_details":{"reasoning_tokens":639}},"tokens_in":423,"tokens_out":685,"duration_ms":6974,"temperature":1.0,"reasoning_tokens":639,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T00:07:59.493073+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a dataset exactly from PRISM's generative assumption: draw τ|t ∼ N(t, γ²), then d|τ ∼ N(μ(τ), Σ(τ)); fit the model and compare the empirical variance of recovered τ at fixed t with 1/I. In the one-dimensional Gaussian case the bound predicts Var(τ|t) ≥ γ² + σ², but the generative truth is Var(τ|t) = γ², so the comparison will refute the bound unless τ is reconstructed deterministically from d.","supporting_citations":[],"review_version":1}