{"id":"8c7189cd-b2ce-438e-9ce3-839b3b0bab08","arxiv_id":"2507.23495","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Bayesian model averaging over the two possible causal directions in a bivariate system is decision-optimal under well-specified models and improves decisions in simulations when structural uncertainty is genuine.","lead":"This paper studies whether decision-makers should average over possible causal directions (X causes Y or Y causes X) instead of committing to one structure. It shows that averaging helps when uncertainty is moderate, the two structures imply different actions, and the loss function is sensitive to these differences.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The bootstrap-derived structural probabilities P(G|D) are never calibrated against the true structure; since Section 7.1 concedes they are not a Bayesian posterior, Section 6.6.1's Delta L = 0.103 cannot directly support Theorems 1 and 2.","rationale":"I read the paper in good faith: it tackles a real gap, the theoretical optimality results are correct under Assumption 2, and the simulations are extensive. The single most load-bearing point, however, is the validity of P(G|D) as the posterior weight. Theorems 1 and 2 prove that model averaging is optimal when the analyst's structural posterior matches the true generative process; they say nothing about arbitrary bootstrap frequencies. The only evidence that the practical approximation is adequate is the simulation, but the simulation never checks calibration. This is not a disagreement with consensus; it is an internal mismatch between an assumption of the theorems and the estimator used in the empirical test. Section 5.2 itself identifies the failure mode: if P(G|D) is poorly calibrated, model selection can beat averaging. The author's own Section 7.1 candidly states that bootstrapping is a heuristic, not a Bayesian posterior. A calibration check would settle the matter and is fully within reach of the existing simulation code. Because the reader already flagged this same concern and assigned CONDITIONAL, I do not change the verdict; I agree with the reader's weakest-assumption analysis. Other issues, such as the unproved Proposition 3 and the mismatch between the kappa-sensitivity definition and quadratic loss, are real but secondary: they affect the 'when does structural uncertainty matter' characterization, whereas the calibration gap undermines the empirical validation of the central optimality claim.","tokens_in":16711,"tokens_out":5487,"duration_ms":65020,"concrete_test":"Decisive check: replicate the 6,400 simulation runs, recording for each run the bootstrap value P(G1|D) and the true structure indicator. Bin P(G1|D) into deciles and plot the mean predicted probability per bin against the empirical frequency of G1; compute the expected calibration error. If any decile deviates by more than 0.1, or ECE exceeds 0.05, the bootstrap frequencies are not posterior probabilities and the Section 6.6.1 claim of supporting Theorems 1 and 2 fails. A supporting comparison: replace the bootstrap P with oracle posterior weights under the known DGP and confirm the Delta L pattern persists; if it does not, the reported benefit is an artifact of the particular heuristic weighting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The empirical support for the central claim rests on treating the bootstrap frequency P(G1|D), computed as the proportion of 100 bootstrap samples in which an ANM or polynomial-regression score favors X->Y (Section 6.2), as a usable posterior over causal structures. Theorems 1 and 2 are conditional on Assumption 2, where the analyst's structural weights coincide with the true generative posterior. Section 7.1 explicitly concedes that bootstrapping 'is not a direct Bayesian posterior.' No calibration check, reliability diagram, or proper-scoring-rule evaluation is reported anywhere in Section 6. If the bootstrap frequencies are miscalibrated, then model averaging is just a deterministic weighting of a heuristic score, and its superiority over model selection is not guaranteed; Section 5.2's Proposition 1 states that with poorly calibrated P(G|D), model selection can outperform averaging. Therefore the observed Delta L = 0.103 over 6,400 runs does not 'directly support' the optimality theorems: it only demonstrates that, for these specific DGPs, one weighting heuristic beats a hard threshold. Because the paper's practical contribution is precisely that modern discovery methods can supply the required quantification, the missing validation of P(G|D) is load-bearing for the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether and when uncertainty about the direction of a bivariate causal relation (X → Y versus Y → X) should be incorporated into decisions, and it proposes Bayesian model averaging over the two candidate structures as an alternative to model selection. It formalizes a decision-theoretic setup, states conditions under which averaging helps (moderate structural uncertainty, differing optimal actions, sensitive loss functions), proves Bayesian and frequentist optimality of averaging under a well-specified hierarchical model, and reports simulations in which bootstrap-derived structural probabilities are used to weight the two structures. The headline simulation result is an average loss reduction of ΔL = 0.103 across 6,400 runs, which the authors interpret as direct empirical support for Theorems 1 and 2. The paper closes with a discussion of limitations and extensions to multivariate settings.","tokens_in":17003,"tokens_out":9075,"duration_ms":89030,"significance":"The topic is relevant: structural uncertainty is often ignored in applied causal inference, and a practical recipe that connects causal discovery output to decision-making would be a useful addition. The paper is clearly written, and it is honest about several limitations, especially in Sections 7 and 8 where it concedes that the bootstrap is not a Bayesian posterior and that calibration of P(G|D) remains an open problem. If the results were fully established, the paper would provide a simple justification for averaging over causal structures. However, the theoretically novel content is thin: Theorem 1 restates the definition of the proposed estimator, Theorem 2 is a standard Bayes-risk-minimization argument, and Proposition 3, which is the main characterization of when averaging helps, is stated without proof and is applied to a loss function that does not satisfy its key condition. The empirical section does not calibrate the bootstrap structural probabilities and uses plug-in effect estimates, so the headline ΔL does not directly validate the optimality theorems. The framework is salvageable, but the main practical claims need substantial additional support.","major_comments":[{"comment":"The claim that ΔL = 0.103 'directly supporting Theorems 1 and 2' is not warranted. Theorems 1 and 2 are conditional on Assumption 2, under which P(G_i|D) are posterior probabilities from a well-specified hierarchical model and the action minimizes posterior expected loss over parameters as well as structures. In the simulations, P(G1|D) is the proportion of bootstrap samples in which a heuristic score favors X → Y (Section 6.2), and Section 7.1 concedes that bootstrapping 'is not a direct Bayesian posterior.' Moreover, Section 6.3 uses point estimates of the causal effect rather than posterior parameter expectations. No calibration check, reliability diagram, or proper-scoring-rule evaluation for P(G1|D) is reported anywhere in Section 6. The observed positive ΔL therefore demonstrates only that this particular bootstrap-weighting scheme beats hard thresholding on these specific DGPs; it does not provide empirical validation of the optimality theorems.","section":"Section 6.6.1; Sections 6.2, 6.3, 7.1"},{"comment":"Theorem 1 is true by construction: Eq. (15) defines a_MA as the minimizer of the posterior expected loss, so Eq. (18) is a restatement of that definition. Theorem 2 then follows immediately from the standard Bayes-risk argument in Eqs. (21)–(22), and the proof given is correct. Presenting these as the paper's 'optimality results' overstates their content. The authors should either reframe them as formal observations that connect the proposed estimator to standard decision theory, or add a substantive result such as finite-sample regret bounds or conditions under which the two-stage model selection rule is approximately optimal.","section":"Section 5.1, Eqs. (15) and (18)"},{"comment":"Proposition 3 is stated without proof, and its application to the simulation loss is invalid. Definition 1 requires |L(a1, Gi, θi) − L(a2, Gi, θi)| ≥ κδ for all pairs with |a1 − a2| ≥ δ, but the quadratic loss L(a) = 0.5(E − a)^2 + λa^2 is not κ-sensitive for any κ > 0: for E = 0, a1 = −ε, and a2 = ε, the loss difference is zero while |a1 − a2| = 2ε. The quantity 1 + 2λ is the second derivative of the loss, not the sensitivity constant of Definition 1. Consequently inequality (17) cannot be invoked for the main example used in the simulation study, and the paper's characterization of loss sensitivity as a driver of averaging benefits is not established for its own primary loss function.","section":"Section 4.3, Proposition 3; Section 6.5"},{"comment":"Proposition 2 is stated without proof and is not generally true for arbitrary loss functions satisfying only Assumption 1. The minimizer of a weighted average of two nonconvex loss functions need not lie in the interval between the individual minimizers, so |a*_1 − a*_2| ≤ ε does not by itself imply |a_MA − a_MS| ≤ ε. A proof, or an explicit convexity assumption on the loss, is needed before this result can support the paper's tripartite characterization of when structural uncertainty matters.","section":"Section 4.2, Proposition 2"}],"minor_comments":[{"comment":"Two different results are both labeled 'Proposition 1': the extreme-certainty result in Section 4.1 and the suboptimality-under-misspecification result in Section 5.2. These should be renumbered to avoid confusion.","section":"Sections 4.1 and 5.2"},{"comment":"The text refers to 'Proposition 4.1' and 'Proposition 4.2' when discussing predictions about sample size and effect size effects; these should be the proposition numbers from Section 4, not section numbers.","section":"Section 6.5"},{"comment":"There is a grammatical error in the sentence 'the causal relationships is:' which should read 'the causal relationships are:'.","section":"Section 3.1.3"},{"comment":"The column header 'n' is used for the number of simulation runs, which conflicts with the sample size n used in Table 2 and throughout the paper; renaming this column 'N' would improve clarity.","section":"Table 3 and Table 4"},{"comment":"The description of causal effect estimation is ambiguous for model averaging runs in which the bootstrap favors G2: the MA action formula requires an estimate of the effect under G1, but the text only describes fitting G1-based models when the estimated direction is G1. Please clarify how Ä E(x) is obtained in those runs.","section":"Section 6.3"},{"comment":"In the derivation of the optimal action, the line after taking the derivative has a sign inconsistency: the equation should read −(E_true − a) + 2λa = 0, which then gives a(1 + 2λ) = E_true.","section":"Section 6.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about its limitations, and the underlying framework is coherent, but the central empirical validation needs a calibration analysis and the theoretical results need reframing or additional content. I recommend major revision rather than rejection because the paper can be repaired by adding a proof or suitable assumptions for Proposition 3, a calibration check for the bootstrap-derived structural probabilities, and a more careful interpretation of the simulation results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper gives a clean and readable account of when structural uncertainty should matter for a bivariate causal decision problem, and it is honest about the gap between its theory and its heuristics. If you work on applied causal inference, the three-factor characterization (uncertainty magnitude, action gap, loss sensitivity) is worth remembering. But the empirical section claims more support than the design can deliver. The bootstrap frequencies used for P(G|D) are never calibrated against true structure, and Section 7.1 admits they are not a Bayesian posterior. So the simulation demonstrates that, for these DGPs, a weighting heuristic beats a hard threshold; it does not directly validate Theorems 1 and 2, which are conditional on a well-specified posterior. The paper's own Proposition (in 5.2) says model selection can beat averaging when P(G|D) is poorly calibrated, so the missing calibration check is the load-bearing gap.\n\nWhat is genuinely new: the formal characterization of when structural uncertainty is decision-relevant, especially the sensitivity condition, and a practical demonstration that modern bivariate discovery methods can be plugged into this averaging recipe. The related-work discussion is fair, and the author explicitly flags many limitations. That is refreshing.\n\nThe math has soft spots. Theorem 1 is a tautology: the MA action is defined as the minimizer of posterior expected loss, so asserting it minimizes posterior expected loss is true by construction. Theorem 2 is standard decision theory. Proposition 3 is stated without proof and its bound does not apply to the quadratic loss used in the simulations: that loss does not satisfy the global kappa-sensitivity condition of Definition 1, since symmetric points around the optimum have equal loss. So the bound is not operational there. Also, the simulations only report average Delta-L, not a calibration check, and the n=10 samples will make bootstrap probabilities noisy. These are addressable issues.\n\nWho is this for? A practitioner who wants a simple rule for when to worry about structural uncertainty in small bivariate studies, and a methods researcher thinking about calibration of discovery outputs. It deserves a serious referee. I would ask for a proof or fix of Proposition 3, a calibration analysis of the bootstrap proportions, and a softened empirical claim. The core direction is sound and the paper is honest enough to build on.","headline":"Clean decision-theoretic framing of when structural uncertainty matters, but the empirical support overreaches because the bootstrap structural probabilities are never calibrated and Proposition 3 doesn't apply to the simulated loss.","tokens_in":17471,"tokens_out":3614,"would_cite":true,"duration_ms":36532,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62P25","91B06"],"pacs":[],"model":"deepseek-v4-flash","headline":"Uncertainty about which way causation runs should be averaged over, not selected away, in causal decision-making.","keywords":["Bayesian model averaging","causal discovery","structural uncertainty","decision theory","bivariate causality","additive noise models","model selection","causal decision-making"],"falsifier":"Run repeated simulations with known true direction and, for each, record the bootstrap $P(G_1\\mid D)$; check calibration by seeing whether the empirical frequency of $G_1$ among cases with $P(G_1\\mid D)\\approx p$ equals $p$. A calibration failure there, or deliberately overconfident structure probabilities plugged into the decision rule, would predict that the positive $\\Delta L$ found in Section 6.6.1 shrinks or reverses.","tokens_in":16486,"feed_emoji":"🎯","tokens_out":8294,"duration_ms":79275,"temperature":0.7,"pith_summary":"The paper confronts a gap in applied causal inference: analysts usually assume the causal structure—which way the arrow points between two variables—is known, even though it often is not. It argues that for bivariate decisions ($X \\to Y$ versus $Y \\to X$), the right response is Bayesian model averaging over the two structures rather than selecting one and proceeding as if it were true. Theoretically, it proves that averaging minimizes both posterior expected loss and frequentist risk under a well-specified hierarchical model, and it characterizes when the gain is largest: moderate structural uncertainty, large differences in implied optimal actions, and loss functions sensitive to action deviations. Simulations with two causal discovery methods support the claim, with an average loss reduction of 0.103 over 6,400 runs relative to model selection.","feed_headline":"In causal decisions, weighting both directions beats picking one","feed_subtitle":"Under structural uncertainty, Bayesian model averaging lowers expected loss in theory and in 6,400 simulations.","key_machinery":"The load-bearing objects are the posterior structure probabilities $P(G_i\\mid D)$ and the averaging rule of Equation (15) that weights each structure's conditional expected loss by those probabilities. A third piece is Proposition 3's lower bound, which decomposes the benefit of averaging into the action gap $\\Delta=|a_1^*-a_2^*|$, the loss sensitivity $\\kappa$, the minimum structural probability, and the model-selection error rate. The bound does the work of saying when structural uncertainty is worth taking seriously.","core_discovery":"On the paper's own terms, the central claim is that the action rule $a^{MA}=\\arg\\min_a \\sum_i \\mathbb{E}[L(a,G_i,\\theta_i)\\mid G_i,D]\\,P(G_i\\mid D)$, which weights each structure's expected loss by its posterior probability, is the decision-theoretically correct response to structural uncertainty. Theorems 1 and 2 state that under a well-specified hierarchical model this rule minimizes posterior expected loss and frequentist risk over all measurable decision rules. Proposition 3 quantifies the advantage over model selection: the expected loss saving is at least $\\kappa\\,\\Delta\\,\\min_i P(G_i\\mid D)\\,P_{\\mathrm{err}}$, where $\\Delta$ is the gap between optimal actions, $\\kappa$ is loss sensitivity, and $P_{\\mathrm{err}}$ is the probability that selection picks the wrong structure. The simulation study then shows the rule in action: averaging outperforms selection on average across heteroskedastic and nonlinear DGPs, with the paper noting that the practical bottleneck is the quality of $P(G\\mid D)$.","pith_inferences":["A practical next step would be to recalibrate bootstrap-derived structure probabilities on data with known direction before using them as decision weights; the paper's own caveat suggests this is where the method would stand or fall.","The same logic implies a cheap diagnostic for practitioners: estimate the gap between optimal actions under candidate structures; if it is tiny, skip structural uncertainty quantification entirely.","For multivariate problems, exact averaging over all DAGs is infeasible, and an approximate posterior over a pruned set would inherit the same calibration risk; the optimality theorems would need re-examination rather than automatic extension."],"forward_implications":["A decision maker facing two plausible causal directions should use the posterior-weighted action rather than the action of the most probable structure, unless the probabilities are known to be miscalibrated.","The biggest gains from averaging should appear in small samples, where discovery is uncertain, and in settings where the two structures recommend very different interventions.","If the loss function is insensitive to action deviations or the optimal actions under the two structures nearly coincide, structural uncertainty does not matter and averaging buys little.","Because the simulation evidence depends on causal discovery methods exploiting nonlinear or heteroskedastic data, averaging will not deliver its advertised benefit on data where the direction is not identifiable."],"supporting_citations":[{"why":"Defines the structural causal model and do-calculus used to formalize causal effects under each bivariate structure.","marker":"[34]"},{"why":"Supplies the Bayesian model averaging framework that the paper adapts from prediction to causal-structure decisions.","marker":"[16]"},{"why":"Closest prior work averaging over candidate causal graphs for effect estimation; the paper extends it with decision-theoretic results.","marker":"[28]"},{"why":"Provides the additive noise model principle used as the ANM causal discovery method in the simulations.","marker":"[12]"},{"why":"Provides the regression-based conditional independence approach used as the second causal discovery method.","marker":"[35]"},{"why":"Provides the bootstrap procedure used to convert discovery outputs into approximate structure probabilities P(G|D).","marker":"[36]"}],"fun_headline_variants":["Structural uncertainty? Average causal models, not choose one","Weight both causal directions to cut decision loss","Bayesian averaging beats model selection in causal choices","Uncertain causality: average structures for lower expected loss","Causal decisions: mix structures, don't pick a side"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The practical gain depends on $P(G\\mid D)$ being a well-calibrated posterior probability over causal structures; the paper computes it as the frequency of bootstrap samples favoring a direction, offers no calibration check, and concedes that bootstrapping is only a heuristic, so miscalibration can flip the conclusion and make model selection win.","fun_headline_variants_meta":{"raw":{"variants":["Structural uncertainty? Average causal models, not choose one","Weight both causal directions to cut decision loss","Bayesian averaging beats model selection in causal choices","Uncertain causality: average structures for lower expected loss","Causal decisions: mix structures, don't pick a side"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1556,"prompt_tokens":883,"completion_tokens":673,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":598}},"tokens_in":499,"tokens_out":673,"duration_ms":8159,"temperature":1.0,"reasoning_tokens":598,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:41:14.933940+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run repeated simulations with known true direction and, for each, record the bootstrap $P(G_1\\mid D)$; check calibration by seeing whether the empirical frequency of $G_1$ among cases with $P(G_1\\mid D)\\approx p$ equals $p$. A calibration failure there, or deliberately overconfident structure probabilities plugged into the decision rule, would predict that the positive $\\Delta L$ found in Section 6.6.1 shrinks or reverses.","supporting_citations":[{"cited_title":"Bayesian model averaging: A tutorial","cited_arxiv_id":null,"evidence_quote":"Supplies the Bayesian model averaging framework that the paper adapts from prediction to causal-structure decisions."},{"cited_title":"Causal effect estimation under multi- ple candidate causal graphs","cited_arxiv_id":null,"evidence_quote":"Closest prior work averaging over candidate causal graphs for effect estimation; the paper extends it with decision-theoretic results."},{"cited_title":"Causal infer- ence in statistics: A primer","cited_arxiv_id":null,"evidence_quote":"Provides the additive noise model principle used as the ANM causal discovery method in the simulations."},{"cited_title":"Causal discovery using regression-based conditional independence tests","cited_arxiv_id":null,"evidence_quote":"Provides the regression-based conditional independence approach used as the second causal discovery method."},{"cited_title":"An introduction to the boot- strap","cited_arxiv_id":null,"evidence_quote":"Provides the bootstrap procedure used to convert discovery outputs into approximate structure probabilities P(G|D)."}],"review_version":1}