{"id":"f791da4c-2b28-4c52-aa73-5931ec0a416e","arxiv_id":"2502.05122","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Bivariate causal direction can be inferred by comparing how well velocity-based models, fitted via estimated score functions, explain the data in each direction.","lead":"This paper introduces causal velocity models, a way to represent cause-effect relationships by treating the cause as time in a dynamical system, and uses them to infer causal direction from data. The approach covers a wider class of mechanisms than standard additive or location-scale noise models, which could help in fields where those assumptions fail.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The decision rule lacks a population-level separation guarantee: Theorem 4.1 characterizes a single direction, and Theorem 5.1 does not establish that comparing L(v) across directions selects the true cause, since every smooth density is bijectively representable in both directions.","rationale":"The reader's weakest_assumption focuses on score estimation quality, which is indeed load-bearing and is honestly documented in Section 7.3 and the abstract. My pass identifies a more fundamental gap: even with perfect scores, nothing in the paper guarantees that the objective comparison identifies the correct direction. Theorem 4.1 is a correct-looking characterization of representability along one direction, but every bivariate distribution admits a bijective SCM representation in both directions when the velocity is unrestricted. The method's practical success therefore depends on the parametric velocity families being restrictive enough to create an asymmetry. The paper's own Section 5.2 admits that no general identifiability analysis is available, and its consistency theorem covers only pointwise loss estimation, not argmin behavior or direction separation. This is not an internal inconsistency; the claims are carefully qualified. But it means the central causal-discovery assertion is supported mainly by selected simulations rather than by a theorem, and Table 2 already shows imperfect accuracy in misspecified settings. My proposed test directly measures population separation with oracle scores, removing the score-estimation confound. If the test confirms a positive gap, the concern is discharged and the method's empirical support is stronger; if not, the conditional verdict should require either a separation theorem or a decision rule with an explicit identifiability check. Since the reader already issued a CONDITIONAL verdict, I keep that verdict unchanged while shifting attention to the direction-separation assumption.","tokens_in":33474,"tokens_out":4695,"duration_ms":54775,"concrete_test":"Run an oracle-score version of the paper's 100 synthetic Velocity datasets. Generate each dataset from a known B-LIN or B-QUAD velocity, compute the population causal loss L(v_true) by numerical integration or very large Monte Carlo, and fit the anti-causal velocity in the same class by globally minimizing equation (17) with exact scores, using dense parameter grids plus multiple Adam restarts. If a nontrivial fraction of datasets satisfy L_anticausal(hat v) <= L_causal(v_true), the direction rule has no population separation and the central claim fails independently of score estimation. To stress the result further, repeat across a range of monotone marginal densities p(x) and check whether the gap inf_anticausal L - L_causal is bounded away from zero uniformly; this isolates exactly the missing identifiability condition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5 proposes estimating a velocity in both candidate directions by minimizing L(v) and choosing the smaller empirical objective. The mathematical support for this decision rule is incomplete. Theorem 4.1 states that equation (14) holds iff PXY can be represented by a velocity-v SCM, but causal discovery additionally requires a population-level asymmetry: for the chosen velocity class V and marginal p(x), inf_v in V of the causal loss must be strictly smaller than inf_v in V of the anti-causal loss, with a margin that survives finite-sample score error. The paper proves no such result. Because every smooth full-support bivariate density is generated by some bijective SCM in both directions, all of the work must come from restricting V; Section 5.2 explicitly stops short of a uniform identifiability analysis, stating that a general analysis 'may require new techniques.' Proposition 5.2 only gives a PDE criterion for simultaneous representability, not a lower bound or a measure-zero argument for the classes actually used (B-LIN, B-QUAD). Theorem 5.1 bounds |hat L_n(v) - L(v)| for a fixed v; it does not control the argmin over v, nor the gap between the two directional losses, so it cannot make direction selection consistent. The empirical evidence in Figure 4 shows one well-specified ANM example where oracle scores separate the directions, but this does not cover the Velocity and Sigmoid benchmarks where Table 2 shows accuracy well below 100 percent even at n=5000. Thus the load-bearing assumption is not only score quality; it is the unverified existence of a strict separation between causal and anti-causal fits within the velocity model class.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a parametrization of bivariate structural causal models (SCMs) through a “causal velocity” v(y,x), viewing the cause variable as time in a dynamical system. The central theoretical result (Theorem 4.1) is an identity relating the joint score, marginal score, and velocity via the continuity equation: sx(x,y) = -∂y v(y,x) - v(y,x) sy(x,y) + sx(x), holding if and only if the joint distribution is representable by an SCM with that velocity. The paper proposes a goodness-of-fit objective L(v) based on this identity, estimates scores nonparametrically, fits velocities in both candidate directions, and selects the direction with smaller objective. Theorem 5.1 states a consistency result for the empirical objective at a fixed v, and Proposition 5.2 derives a PDE condition for non-identifiability. Experiments on synthetic and benchmark data show the method succeeding on Velocity and Sigmoid benchmarks where ANM/LSNM methods struggle, with ablation studies examining the dependence on score estimation quality.","tokens_in":33782,"tokens_out":5097,"duration_ms":49054,"significance":"The velocity-based parametrization is a conceptually new and elegant way to connect SCMs, counterfactuals, and score functions, extending functional causal discovery beyond the ANM and LSNM classes without specifying noise distributions. The paper is honest about the central role of score estimation, provides code, and includes oracle-score experiments that cleanly isolate the behavior of the objective. If the method is reliable, it offers a practical tool for bivariate causal discovery in settings where existing methods are misspecified. However, the theoretical support for the direction-selection rule is incomplete, and the empirical results show non-negligible error rates even in well-specified settings, which tempers the practical significance until the decision rule is better understood.","major_comments":[{"comment":"The proposed causal discovery rule selects the direction with the smaller empirical objective, but the paper provides no population-level separation guarantee for the velocity classes actually used (B-LIN, B-QUAD, V-NN). Theorem 4.1 characterizes representability in a single direction, and Theorem 5.1 only bounds |That L_n(v) - L(v)| for a fixed v; neither controls the gap between the two directional infima. Since every smooth full-support bivariate density is generated by some bijective SCM in both directions, all identifiability must come from restricting the velocity class, and Section 5.2 explicitly stops short of a uniform identifiability analysis, stating that a general analysis “may require new techniques.” This is load-bearing for the central claim that the method distinguishes cause from effect.","section":"Section 5, Eq. (15)–(18)"},{"comment":"The consistency proof bounds the empirical-process term by Rademacher averages over the class {f(s,·) : s ∈ B(H,M)}. In the expansion (36), product classes such as B(H1)⊗B(H2) and B(H1)⊗B(H3)⊗{h2} are treated as closed balls in tensor-product RKHSs. The set of products of two functions with bounded RKHS norm is not itself a ball in the tensor-product RKHS, and the assertion that “every term in (36) is the Rademacher average of a closed ball in an RKHS of bounded functions” is not justified. Thus Theorem 5.1, which is presented as establishing consistency of the estimation procedure, is not fully supported by the supplied proof.","section":"Appendix B.3 (proof of Theorem 5.1)"},{"comment":"Proposition 5.2 is stated as an “if and only if” characterization of non-identifiability, but the proof in Appendix C only equates the two mixed partial derivatives (50) and (51). Sufficiency—that equality of these derivatives implies existence of a joint density satisfying both directional continuity equations—is not shown. Moreover, the proposition does not yield a measure-zero or lower-bound argument for the specific basis classes (B-LIN, B-QUAD) used in the experiments, so the identifying power of those classes is only empirically demonstrated, not theoretically established.","section":"Section 5.2, Proposition 5.2"}],"minor_comments":[{"comment":"The sentence “It is easy to see that Equation (9) satisfies the axioms of a flow Theorem 2.1” is missing a word; it should read “of a flow in Theorem 2.1.”","section":"Section 3, Definition 3.1"},{"comment":"The phrase “which we suspect is due to the score being well-estimated” is informal; a more precise statement or a reference to the score-MSE results would strengthen the explanation.","section":"Section 7.2"},{"comment":"For the well-specified LSNM experiment, the success rate is only 73% at n=5000 and 80% at n=10000 even when the true score is not used; the text should explicitly discuss why a well-specified model with improving score estimates still does not approach perfect accuracy.","section":"Section 7.3, Table 8"},{"comment":"The Sigmoid benchmark is described verbally as “a variation on LSNMs with additional post-nonlinear and affine transformations,” but the mechanism in Table 4 is sufficiently complex that a brief interpretation or a reference to the table in the main text would improve reproducibility.","section":"Section 7.1 / Table 4"},{"comment":"The caption states “Upper and lower limits indicate 1st– 3rd quartile over 100 datasets,” but the figure appears to show shaded bands; specifying whether the limits are bands or error bars would help the reader.","section":"Figure 5 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the core identity is a solid contribution. The main concern is that the decision rule's theoretical basis is weaker than the title and abstract suggest; the authors should either supply a population-level separation result for the simple basis families, or substantially temper the methodological claims while keeping the score-identity contribution. The citation of prior cocycle work (Dance & Bloem-Reddy 2024) is appropriate and not excessive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time, but go in knowing the strongest part and the weakest part. The strongest part is Theorem 4.1: the score identity that ties a bivariate SCM to a causal velocity via the continuity equation. That piece is derived cleanly, it generalizes score-based ANM results to arbitrary bijective mechanisms, and it gives a genuinely different way to think about counterfactuals as flow trajectories. The velocity parametrization itself is a nice contribution, and the B-LIN/B-QUAD families are new model classes that go beyond ANM/LSNM. I believe the core identity is correct and useful.\n\nWhat the paper does well beyond the math: the experiments are honestly reported. They include ablations on score-estimation quality, show where the method degrades, release code, and they explicitly say in Section 5.2 that a general identifiability analysis 'may require new techniques.' That last admission matters because the decision rule is the soft spot. The paper chooses causal direction by comparing the minimized objective in both directions, but there is no proof that the causal loss is strictly smaller than the anti-causal loss for the velocity classes they actually use. Theorem 5.1 bounds the empirical loss for a fixed velocity; it does not control the argmin or the gap between directions. Since every nice bivariate density is compatible with some bijective SCM in both directions, all identifiability has to come from restricting the velocity class, and the paper does not establish that B-LIN/B-QUAD have the needed separation. This is not a fatal flaw—the oracle-score experiments in Figure 4 show separation in a well-specified case—but it is a real gap between the claim of 'we can extend causal discovery' and what is proven. The paper is transparent about this, which I credit.\n\nThe other soft spot is the acknowledged dependence on score estimation. When scores are bad, the method fails; benchmarks with integer-valued data are excluded, disclosed, but it is still a post-hoc choice that flatters results. The RKHS consistency proof also has some sketching gaps in the tensor-product details, though the result is plausible.\n\nWho should read it: anyone working on functional causal discovery or on SCMs as flows. It deserves a serious referee. I'd send it to review, but I'd expect the reviewers to ask for either a separation result under stated conditions or a clearer acknowledgment that direction selection is a heuristic with empirical support rather than a theorem.","headline":"The velocity–score identity is the real contribution; the direction-selection rule is a heuristic that needs either a separation proof or a humbler framing.","tokens_in":34336,"tokens_out":3360,"would_cite":true,"duration_ms":30611,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper establishes a noise-agnostic score identity for bivariate causal direction: a joint distribution is generated by an SCM with causal velocity $v$ if and only if $s_x(x,y)=-\\partial_y v(y,x)-v(y,x)s_y(x,y)+s_x(x)$, and comparing…","keywords":["causal discovery","bivariate structural causal models","score function","dynamical systems","measure transport","continuity equation","noise-agnostic inference","counterfactual curves"],"falsifier":"Compute the true joint and marginal scores for a generated SCM outside the ANM and LSNM classes and fit the velocity in the anti-causal direction using those exact scores; if the anti-causal fit reaches zero or matches the causal fit, the direction-selection claim fails. A concrete variant is to choose a velocity whose flow has closed-form counterfactuals, generate the exact joint density, and then check whether any alternative velocity in the same family satisfies the Theorem 4.1 identity in the reverse direction.","tokens_in":33277,"feed_emoji":"🧭","tokens_out":5960,"duration_ms":53289,"temperature":0.7,"pith_summary":"The paper tries to establish that, for bivariate systems, the causal direction can be read from the score function of the observed data if the generative mechanism is re-expressed as a dynamical system. The central move is to treat the cause variable as time and describe the SCM's counterfactual curves through a causal velocity. A score identity connects that velocity to the joint and marginal scores without ever specifying the noise distribution. The paper then compares a goodness-of-fit objective built from the identity in the $X\\to Y$ and $Y\\to X$ directions, and whichever direction fits is inferred to be causal. This matters because the velocity formulation covers mechanisms that additive and location-scale noise models cannot, while still allowing simple parametric velocity families to be estimated directly from data.","feed_headline":"A score identity decides which way causality runs","feed_subtitle":"Pairwise fit of a causal velocity to the score separates cause from effect, with no noise-distribution assumption.","key_machinery":"The central object is the causal velocity $v(y,x)$, the derivative at the factual point of the counterfactual curve that maps an observed $y$ to its value had the cause $x$ been changed infinitesimally. It plays the role of the velocity field in an ODE whose time direction is the cause variable, generated by the flow $\\varphi_{x,x'}(y)=f_{x'}(f_x^{-1}(y))$. The paper pairs this velocity with the continuity equation $\\partial_x \\log p(y\\mid x) = -\\partial_y v(y,x) - v(y,x)\\partial_y \\log p(y\\mid x)$, converting it into the score identity of Theorem 4.1; this conversion is the mechanism that turns observational scores into a noise-agnostic regression target.","core_discovery":"The paper's central claim is an if-and-only-if identity. Under regularity conditions, a joint distribution with full support and differentiable log-densities can be represented by a structural causal model with velocity $v$ exactly when $$s_x(x,y) = -\\partial_y v(y,x) - v(y,x)\\,s_y(x,y) + s_x(x)$$ holds for all $(x,y)$, where $s_x(x,y)$ and $s_y(x,y)$ are partial derivatives of the joint log-density and $s_x(x)$ is the marginal score of the cause. The identity is the log form of the continuity equation applied to the SCM flow, and it makes the noise distribution disappear from the estimation problem. The paper fits $v$ by minimizing $L(v)=\\mathbb{E}[(s_x(X)-\\partial_y v(Y,X)-s_v(X,Y))^2]$, where $s_v(x,y)=s_x(x,y)+v(y,x)s_y(x,y)$ is the directional derivative of the log-density along the causal curve, and chooses the direction with the smaller fitted objective. Because nothing in the objective requires an additive or location-scale mechanism, the same procedure applies to a wider class of models and needs no assumption on the noise.","pith_inferences":["Applied to multivariate data, the pairwise score comparison could in principle be combined with graph-structure search, since the identity is local and does not depend on the noise distribution; the paper itself leaves that extension unexplored.","The identifiability equation obtained by writing the continuity equation in both directions suggests that similar uniform non-identifiability criteria might be derivable for other mechanism classes beyond ANM and LSNM, not just case-by-case checks.","Because the method needs a well-defined log-density, extending it to discrete or integer-valued variables would require smoothed densities or surrogate scores; the paper explicitly drops such benchmark instances rather than handling them.","The velocity parametrization also gives practitioners a way to encode mechanistic knowledge directly as infinitesimal-intervention responses, so a domain-specified velocity could be tested against data before making any parametric noise assumption."],"forward_implications":["For any mechanism whose velocity can be represented, causal direction can be decided without fitting an additive-noise or location-scale model and without Gaussian assumptions on the noise.","The same goodness-of-fit value can flag model misspecification or non-identifiability when the score is estimated well, because the identity holds only when the velocity matches the data-generating SCM.","Integrating the estimated velocity recovers counterfactual curves, so velocity estimation yields a counterfactual model from observational data as a by-product.","If the score estimators converge, the fitted objective converges at a rate no worse than the score estimators, so improvements in nonparametric score estimation directly improve the reliability of the causal discovery procedure."],"supporting_citations":[{"why":"Supplies the continuity-equation and measure-transport facts that turn SCM flows into the score identity.","marker":"[Santambrogio, 2015]"},{"why":"Provides the nonparametric score estimators whose convergence rates set the consistency rate of the fitted objective.","marker":"[Zhou et al., 2020]"},{"why":"Defines the ANM identifiability problem whose differential-equation characterization the paper recovers as a special case.","marker":"[Hoyer et al., 2008]"},{"why":"Develops the location-scale noise model that is the main prior class the velocity method extends and the main comparison baseline.","marker":"[Immer et al., 2023]"},{"why":"Shows scores can be deterministic functions of noise in additive models; the paper recovers that identity as an ANM special case.","marker":"[Montagna et al., 2023a]"},{"why":"Connects simulation-free velocity regression from generative modeling to the paper's fitting objective.","marker":"[Lipman et al., 2023]"},{"why":"Establishes the cocycle and dynamical view of causal models that the paper specializes to bivariate flows.","marker":"[Dance & Bloem-Reddy, 2024]"},{"why":"Provides the benchmark cause-effect pairs and evaluation protocol used to test the discovery method.","marker":"[Mooij et al., 2016]"}],"fun_headline_variants":["Score identity pinpoints the causal direction","Causal velocity matched to score reveals cause","No noise assumptions: velocity picks cause from effect","Cause or effect? Fit velocity to the score"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the joint and marginal score functions of the observed distribution can be estimated accurately enough from finite data; the paper's own ablations show discovery quality degrades when score estimates are poor, and discrete or integer-valued benchmark instances cannot be used because their scores are not defined.","fun_headline_variants_meta":{"raw":{"variants":["Score identity pinpoints the causal direction","Causal velocity matched to score reveals cause","No noise assumptions: velocity picks cause from effect","Cause or effect? Fit velocity to the score"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000559,"raw_usage":{"total_tokens":2685,"prompt_tokens":999,"completion_tokens":1686,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":1630}},"tokens_in":615,"tokens_out":1686,"duration_ms":12958,"temperature":1.0,"reasoning_tokens":1630,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T20:09:39.825323+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the true joint and marginal scores for a generated SCM outside the ANM and LSNM classes and fit the velocity in the anti-causal direction using those exact scores; if the anti-causal fit reaches zero or matches the causal fit, the direction-selection claim fails. A concrete variant is to choose a velocity whose flow has closed-form counterfactuals, generate the exact joint density, and then check whether any alternative velocity in the same family satisfies the Theorem 4.1 identity in the reverse direction.","supporting_citations":[{"cited_title":"Optimal Transport for Applied Mathematicians","cited_arxiv_id":null,"evidence_quote":"Supplies the continuity-equation and measure-transport facts that turn SCM flows into the score identity."},{"cited_title":"Nonparametric score estimators","cited_arxiv_id":null,"evidence_quote":"Provides the nonparametric score estimators whose convergence rates set the consistency rate of the fitted objective."},{"cited_title":"M., Peters, J., and Sch \\\"o lkopf, B","cited_arxiv_id":null,"evidence_quote":"Defines the ANM identifiability problem whose differential-equation characterization the paper recovers as a special case."},{"cited_title":"o lkopf, B., B \\","cited_arxiv_id":null,"evidence_quote":"Develops the location-scale noise model that is the main prior class the velocity method extends and the main comparison baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Connects simulation-free velocity regression from generative modeling to the paper's fitting objective."},{"cited_title":"M., Peters, J., Janzing, D., Zscheischler, J., and Sch \\\"o lkopf, B","cited_arxiv_id":null,"evidence_quote":"Provides the benchmark cause-effect pairs and evaluation protocol used to test the discovery method."}],"review_version":1}