{"id":"e65f29fe-68f7-4da4-83fb-1e2178929215","arxiv_id":"2501.12500","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"CaDRe jointly recovers latent dynamic processes and observed causal graphs from time-series data, with identifiability theory and competitive climate forecasting.","lead":"CaDRe is a new framework that simultaneously learns latent climate drivers and causal links between observed climate variables from time-series data, with theoretical identifiability guarantees. It combines causal representation learning with causal discovery and matches or beats specialized forecasting models on climate benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3's proof assumes J_hs is independent of z_t,l, but hs = m^{-1}∘ĝ_m depends on z_t through both Jacobians; Eq. (A38) omits ∂J_hs/∂z_t,l terms, so A5 alone does not yield the monomial matrix.","rationale":"The reader correctly identified A5 as a fragile, data-dependent assumption, and the paper's own ablation shows performance collapses when A5 is violated. My stress-test goes further: even when A5 is granted, the proof of Theorem 3 contains a chain-rule error. The transformation hs = m^{-1}∘ĝ_m has a Jacobian J_hs = (J_ĝ_m)^{-1} J_gm that depends on z_t through both factors, so the derivative w.r.t. z_t,l in Eq. (A38) cannot simply discard ∂J_hs/∂z_t,l terms. Without those terms, the linear-independence condition on V and U does not force J_hs to be monomial, and the route to supp(J_g)=supp(J_ĝ) breaks. This is a concrete, fixable proof gap rather than a refutation of the framework, so the appropriate editorial disposition remains conditional acceptance pending a corrected proof. The reader's verdict is therefore unchanged, but the revision requirements should explicitly include re-deriving Eq. (A38) and either proving the omitted terms vanish or adding an assumption to Theorem 3.","tokens_in":40654,"tokens_out":10123,"duration_ms":101647,"concrete_test":"Recompute Eq. (A38) by symbolic differentiation of Eq. (A36) without dropping ∂J_hs/∂z_t,l terms, using J_hs = (J_ĝ_m)^{-1} J_gm. If nonzero extra terms survive, construct the minimal example d_x=d_z=1, g_m(z,s)=a(z)s, ĝ_m(ĝ_z,ĝ_s)=b(ĝ_z)ĝ_s, h_z(z)=z; then J_hs=b(z)/a(z) depends on z and the omitted derivative is nonzero. Check whether the monomial conclusion and supp(J_g)=supp(J_ĝ) still hold for such an example satisfying A5. If extra terms vanish by an unstated cancellation, the proof can be repaired; otherwise Theorem 3 requires an additional assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proof of Theorem 3 (Appendix A.7) rests on a chain-rule step that appears invalid. From Eq. (A36), log p(ĝ_s|ĝ_z) = log p(s|z) − log|J_hs(ĝ_s)|. Eq. (A38) differentiates the second derivative of this identity w.r.t. z_t,l and drops all terms involving derivatives of J_hs w.r.t. z_t,l, stating that 'entries of J_hs(ĝ_s) do not depend on z_t,l.' But hs is defined as m^{-1}∘ĝ_m, and x_t = g_m(z_t,s_t) = ĝ_m(ĝ_z,ĝ_s) with ĝ_z = h_z(z_t). Differentiating w.r.t. s_t gives J_hs = (J_ĝ_m(ĝ_s))^{-1} J_gm(s_t); both Jacobians generically depend on z_t (through g_m and through h_z), so ∂J_hs/∂z_t,l is typically nonzero. The omitted terms are not in the span of the V and U functions in A5, so the linear-independence argument that forces J_hs to be monomial is unsupported. Since the monomial step is what upgrades blockwise recovery to ordered component-wise identifiability and hence to supp(J_g)=supp(J_ĝ), Theorem 3 as stated is not established. This is independent of whether A5 is verifiable; even granting A5, the proof step fails unless an additional condition (e.g., mixing Jacobian independent of z_t) is imposed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CaDRe, a time-series generative model that jointly learns latent dynamic processes and causal relations among observed variables. The theoretical part claims, in Theorem 3, that under Assumptions 1, 2, and A1–A5, the causal graph over observed variables is identifiable, i.e., supp(J_g(x_t)) = supp(J_ĝ(ẑ_t, ŝ_t)), building on a latent-space identifiability result (Theorem 1) and a functional equivalence between the SEM and nonlinear ICA (Theorem 2). The method instantiates this via a variational autoencoder with flow-based priors, Jacobian-based structural penalties, and DAG constraints. Experiments on synthetic data show strong recovery of latent components and causal structure, and experiments on climate benchmarks report competitive forecasting and graphs that align with a wind-field surrogate.","tokens_in":41013,"tokens_out":10666,"duration_ms":109314,"significance":"If the identifiability results were fully established, the paper would make a useful advance: it extends causal representation learning from invertible, deterministic mixing to settings with noisy, non-invertible generation and causally related observed variables, and it provides an integrated algorithm with code. Strengths of the submission include explicit assumptions, ablation studies that test assumption violations, extensive comparisons to constraint-based and CRL baselines, and reproducible experimental protocols. However, the central theoretical claim is not currently established: the proof of Theorem 3 contains a false statement about the Jacobian J_{h_s}, and Theorem 1's proof invokes a spectral decomposition without justifying the required operator conditions. These are load-bearing issues, so the paper needs substantial revision before the theoretical claims can be accepted.","major_comments":[{"comment":"The proof of Theorem 3 differentiates Eq. (A37) with respect to z_{t,l} and drops all terms involving derivatives of J_{h_s}(ŝ_t), asserting that 'entries of J_{h_s}(ŝ_t) do not depend on z_{t,l}'. This assertion is false as stated: h_s is defined as m^{-1}∘ĝ_m, and since x_t = g_m(z_t,s_t) = ĝ_m(ẑ_t,ŝ_t) with ẑ_t = h_z(z_t), both J_m and J_ĝ_m, hence J_{h_s}, generically depend on z_t. The omitted ∂J_{h_s}/∂z_{t,l} terms are not controlled by Assumption A5, which only constrains derivatives of A_{t,k} = log p(s_{t,k}|z_t). Consequently Eq. (A38) does not establish that [J_{h_s}]_{k,i}[J_{h_s}]_{k,j} = 0 for i≠j, and the monomial-matrix step, ordered component-wise identifiability, and the final support equality supp(J_g)=supp(J_ĝ) are unsupported. The proof needs either a correct treatment of the dropped terms or an additional explicit condition, such as the mixing Jacobian being independent of z_t.","section":"Appendix A.7, Eq. (A38)"},{"comment":"Theorem 1's proof obtains Eq. (A13) by invoking the uniqueness of a spectral decomposition of L_{x_{t+1}|z_t} D_{x_t|z_t} L^{-1}_{x_{t+1}|z_t}. The stated hypotheses are injectivity and boundedness of the operators, but these do not imply that the operator is normal or that it admits a unique spectral decomposition in the sense of the cited theorems (Conway Ch. VII; Dunford & Schwartz, Theorem XV 4.5). For a general injective bounded operator L, the factorization L D L^{-1} is not unique, so the identification of the eigenfunctions {p(x_{t+1}|z_t)} up to permutation and scaling is not justified. This gap is load-bearing because Theorem 3 uses Theorem 1 to obtain ẑ_t = h_z(z_t) and to justify the change of variables leading to Eq. (A35).","section":"Appendix A.2, Eqs. (A11)–(A13)"},{"comment":"The step that removes the permutation indeterminacy is also not rigorously established. From J_{g_L}(x_t) = P_{d_x} J_g(x_t) P_{d_x}^\top and the relations in Corollary 2.2, direct algebra gives J_g = P^\top J_{g_L} P and hence J_m = P^\top J_{m_L} D_{m_L}^{-1} P D_m, rather than the expression in Eq. (A40), which appears to swap P and P^\top. Moreover, the transition from Eq. (A41) to 'Using Lemma 3, we obtain ... = I' is asserted without showing that the product in question satisfies the hypotheses of Lemma 3. This step is needed to conclude that J_ĝ_m and J_m have the same support, and it is therefore load-bearing for the final causal graph identifiability claim.","section":"Appendix A.7, Eqs. (A39)–(A41)"}],"minor_comments":[{"comment":"The text says 's-encoder ψ and decoder η', but Eq. (6) defines ŝ_t = η(x_{1:T}) and ẑ_t = ψ(ẑ_t, ŝ_t); the roles of ψ and η should be swapped in the prose for consistency.","section":"Section 4, Eq. (6)"},{"comment":"The sentence referring to the 'latent causal process in Theorem A.3' appears to point to a non-existent theorem; the component-wise latent identifiability result in Appendix A.3 is Theorem A1.","section":"Section 3 introduction"},{"comment":"The assumption labels in the bullet list are inconsistent with Table A11: item i is labeled 'Violation of Assumption 1' while the table says 'Violate A2', and item iii is labeled 'Violation of Assumption 3' while the table says 'Violate A5'. These should be aligned.","section":"Appendix D.1, ablation study"},{"comment":"The text should state more explicitly that WSHD and WTPR compare the learned graph to the wind-field surrogate B_ref, which is a physically motivated proxy and not a validated causal ground truth; as written, 'Quantitative Results on CD' could be read as claiming direct validation of causal structure.","section":"Section 5.2 and Appendix C.1.2(vii)"}],"recommendation":"major_revision","confidential_remarks":"The experimental component is substantial and the manuscript is clearly written, but the theory section requires major work. The false statement in Eq. (A38) is not a local typo: it removes exactly the terms that prevent Assumption A5 from implying the monomial-matrix conclusion. If a correct proof cannot be supplied under transparent additional conditions, the paper should be reframed as an empirical method with heuristic identifiability. I do not see a path to acceptance with the current Appendix A.7."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about arXiv:2501.12500. First, the problem it tackles—jointly recovering latent dynamic processes and observed-variable causal graphs from time series—is real, and the proposed CaDRe framework is a plausible way to approach it. Second, the proof of the paper's central claim, Theorem 3, has a gap that I don't think is cosmetic.\n\nThe stress-test note is right: Appendix A.7, Eq. (A38), drops all terms involving ∂J_hs/∂z_t,l when differentiating the log-density identity. But hs = m^{-1}∘ĝ_m depends on z_t through both Jacobians, so those derivatives are generically nonzero. The linear-independence argument in assumption A5 therefore cannot, as written, force J_hs to be a monomial matrix. That monomial step is what upgrades blockwise identifiability to ordered component-wise identifiability and gives supp(J_g)=supp(J_ĝ). Without it, Theorem 3 is not established. The paper's own ablation shows performance drops when A5 is violated, but that doesn't repair the proof step.\n\nWhat's genuinely good. The SEM-to-ICA equivalence in Lemma 2/Theorem 2 is a nice conceptual bridge, even though the proof of Theorem 2 in Appendix A.5 contains a false statement about off-diagonal entries of J_m: it claims [Jm]_{k,i}=0 for k≠i, which is generally wrong because s_t,i can reach x_t,k through indirect paths. The experimental section is thorough: synthetic sweeps over dx, dz, sample size, and sparsity, plus real-world climate benchmarks with a wind-pattern surrogate. Code is available. The framing of the problem is new relative to Monti et al. and Reizinger et al., and the nonparametric latent recovery via spectral decomposition, if it can be made rigorous, would be a real contribution.\n\nThe soft spots. Besides the proof gap, Theorem 1's spectral decomposition is invoked without sufficient justification for nonnormal operators; the regularity conditions are not stated carefully enough. There's also an internal theorem-numbering mismatch and the component-wise identifiability of latent variables is imported from the authors' own prior work (Li et al. 2025), which is fine as a reference but means the paper's contribution is narrower than the abstract suggests. The real-world causal graph evaluation uses a self-defined wind surrogate, so it's evidence of plausibility, not validation.\n\nWho this is for and whether it deserves review. If you work on causal representation learning or causal discovery for time series, this paper is worth reading for its problem setup and the CaDRe framework, but you shouldn't take the identifiability guarantees at face value. The gap in Theorem 3 is load-bearing; I don't see how to fix it without additional assumptions on the mixing Jacobian. That said, it deserves a serious referee: the ideas are important enough and the experimental evidence is substantial enough that a careful review could either find a repair or produce a counterexample. I would not cite the theoretical claims in their current form.\n\nSend to peer review, but be prepared for major revision or rejection depending on whether the gap can be closed. If I were the editor, I'd want a referee who can check the analysis in Appendix A.7 carefully.","headline":"Promising framework for joint latent-process and observed-graph identification, but the central theorem's proof has a gap that is likely not patchable without extra assumptions.","tokens_in":41559,"tokens_out":3529,"would_cite":false,"duration_ms":34261,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62M10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that a time-series generative model can recover both the hidden dynamic process and the causal graph among observed climate variables, and implements this as CaDRe.","keywords":["causal discovery","causal representation learning","latent dynamic processes","time series","climate analysis","identifiability","nonlinear independent component analysis","variational autoencoder"],"falsifier":"Construct a simulated system that satisfies the other assumptions but violates A5 by giving every noise source the same linear dependence on $z_t$ (identical second and third derivatives of $\\log p(s_{t,k}|z_t)$); run CaDRe and compare the estimated support of $J_g$ with ground truth. The paper already reports that violating A5 drops source-recovery MCC from about 0.96 to 0.71, so a clean pass/fail test is whether the causal-graph SHD degrades correspondingly.","tokens_in":40452,"feed_emoji":"🌍","tokens_out":9042,"duration_ms":82804,"temperature":0.7,"pith_summary":"The paper sets out to show that a single time-series model can simultaneously recover two things that are usually treated separately: the hidden dynamic process driving a system (the latent variables and their causal interactions over time) and the causal graph among the measured variables themselves. This matters for climate analysis because temperature and precipitation records are driven by unmeasured processes such as pressure and solar radiation while also influencing each other across nearby regions; knowing both kinds of structure turns forecasting into something interpretable. The paper establishes identifiability conditions under which both structures are recoverable from observational data alone, without interventions, and instantiates the theory in CaDRe, a variational autoencoder with flow-based priors and structural penalties. If the central claim is right, climate scientists could read real causal structure off historical records rather than only chase predictive accuracy.","feed_headline":"Hidden climate drivers made visible with recovered causal graphs","feed_subtitle":"A single generative model recovers latent drivers and observed causal graphs from time-series data alone.","key_machinery":"The load-bearing objects are Jacobian matrices of two equivalent views of the same generation process: $J_g(x_t)$, whose support encodes direct causal edges among observed variables, and $J_m(s_t)$, the ICA mixing Jacobian from noise sources to observations, with $D_m$ its diagonal part. The identity $J_g J_m = J_m - D_m$ converts the mixing structure into a causal graph in one step, $J_g = I - D_m J_m^{-1}$. On the estimation side, CaDRe uses two encoders, flow-based prior networks whose inverse transition functions encode latent causal structure, and sparsity plus DAG penalties applied to the recovered Jacobians; CaDRe is the name of the resulting variational-autoencoder method.","core_discovery":"On the paper's own terms, the central result is Theorem 3: under the Markov/faithfulness assumption on the full graph, functional faithfulness of the observed Jacobian, and a generation-variability condition (A5), the support of the Jacobian matrix of the learned observation map is identical to the support of the true observation-level causal graph, $\\mathrm{supp}(J_g(x_t))=\\mathrm{supp}(J_{\\hat g}(\\hat x_t))$. The proof chains three identifiability results: latent space is recoverable up to an invertible differentiable map from three consecutive observations (Theorem 1); the structural equation model is equivalent to a nonlinear ICA mixing model (Lemma 2), which yields the functional identity $J_g J_m = J_m - D_m$ (Theorem 2); and from that identity the observation-level causal graph is computed as $J_g = I - D_m J_m^{-1}$. The same machinery also gives ordered component-wise identifiability of the latent process, so individual latent components and their time-lagged and instantaneous dependencies are recoverable, not just the subspace.","pith_inferences":["A practical consequence the paper does not spell out: before trusting a recovered climate graph, a user should test whether the generation-variability condition (A5) plausibly holds, since heterogeneity across time is precisely what makes latent sources separable; a homogeneity diagnostic would be a natural companion tool.","The Jacobian identity suggests a transferable principle: in any spatiotemporal system where hidden drivers modulate noise, the observed causal graph is one matrix inversion away from the mixing Jacobian, so a similar decoder-plus-Jacobian architecture could be adapted to neuroscience or economics.","The paper's wind-based evaluation is a correlation surrogate, not a causal gold standard; a stronger test would perturb one grid variable in a climate simulation and check whether the recovered graph predicts the resulting response pattern."],"forward_implications":["A generative model trained only to reconstruct time-series observations can output the true causal graph among measured variables, even when hidden drivers exist and the mapping from latent to observed variables is non-invertible and noisy.","Latent drivers identified by the model are component-wise aligned with physical quantities, so learned factors such as solar radiation or cloud cover can be read as interpretable climate variables.","The recovered graph and latent dynamics are produced in one forward pass, giving causal structure learning without the repeated conditional-independence tests of constraint-based methods.","Because the framework carries identifiability guarantees into the nonparametric regime, it applies to forecasting benchmarks beyond climate, including health, electricity, and traffic series."],"supporting_citations":[{"why":"Supplies Assumption 1 (Markov and faithfulness to a DAG) and the DAG semantics that define the causal graph the paper recovers.","marker":"Spirtes et al., 2001"},{"why":"Provides the nonparametric spectral-decomposition identification strategy that Theorem 1 extends to value-level latent recovery without knowing the function form.","marker":"Hu & Schennach, 2008"},{"why":"Supplies the cross-derivative condition (Lemma 4) that forces the estimated latent-to-latent Jacobian to be monomial, eliminating mixing within latent components.","marker":"Lin, 1997"},{"why":"Contributes the lower-triangular permutation argument used to eliminate permutation ambiguity and equate supports of estimated and true causal Jacobians.","marker":"Shimizu et al., 2006"},{"why":"Provides the differentiable DAG penalty used in the loss to keep the learned causal structures acyclic.","marker":"Yu et al., 2019"}],"fun_headline_variants":["CaDRe: one model recovers climate causal graph and hidden drivers","Provable joint identification of observed and latent climate causes","Climate causality from time series: hidden drivers and observed links","Unified model learns latent drivers and causal structure for climate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the generation-variability condition (A5): the second- and third-order derivatives of each noise source's log-density with respect to the latent variables must be sufficiently varied across time steps, otherwise the method cannot tell the noise components apart and the recovered observation graph loses its guarantee.","fun_headline_variants_meta":{"raw":{"variants":["CaDRe: one model recovers climate causal graph and hidden drivers","Provable joint identification of observed and latent climate causes","Climate causality from time series: hidden drivers and observed links","Unified model learns latent drivers and causal structure for climate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001007,"raw_usage":{"total_tokens":4274,"prompt_tokens":982,"completion_tokens":3292,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":3223}},"tokens_in":598,"tokens_out":3292,"duration_ms":24255,"temperature":1.0,"reasoning_tokens":3223,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:09:24.018853+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a simulated system that satisfies the other assumptions but violates A5 by giving every noise source the same linear dependence on $z_t$ (identical second and third derivatives of $\\log p(s_{t,k}|z_t)$); run CaDRe and compare the estimated support of $J_g$ with ground truth. The paper already reports that violating A5 drops source-recovery MCC from about 0.96 to 0.71, so a clean pass/fail test is whether the causal-graph SHD degrades correspondingly.","supporting_citations":[{"cited_title":"Factorizing multivariate function classes","cited_arxiv_id":null,"evidence_quote":"Supplies the cross-derivative condition (Lemma 4) that forces the estimated latent-to-latent Jacobian to be monomial, eliminating mixing within latent components."}],"review_version":1}