{"id":"4664ac8b-7ebd-45c6-9edd-b9b7d1533ce2","arxiv_id":"2607.21667","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Equilibrium causal games give conditions for validating and transporting counterfactual predictions in feedback systems, plus an impossibility theorem for finite experimental designs.","lead":"This paper builds a mathematical framework for deciding when a digital twin of a system with feedback can be trusted after interventions and after the system changes. It shows that enough experiments can validate some counterfactual predictions, but finite experiments can never rule out all cross-world errors without extra structural assumptions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 31's reshuffled twin can change the equilibrium law when the selection rule reads the actual noise rank rather than the reshuffled rank; the impossibility proof is therefore not established for multi-equilibrium systems.","rationale":"The reader's weakest assumption was the exogenous-independence/copula exclusion in Section 2.4. That exclusion is explicit, stated as a standing assumption, and the paper openly declines to relax it; it is a scope limitation rather than an unproved step. The more load-bearing issue is internal to the proof of Theorem 31, which underpins the paper's headline claim that finite experiments cannot validate cross-world predictions without structural assumptions. The construction of S^- relies on a measure-preserving reshuffle, but the paper does not account for the selection rule T3: Sel is a function of the actual noise rank vector and the solution set, and after the reshuffle these two arguments are generated by different rank vectors. A measure-preserving transformation of one noise coordinate need not preserve the law of the selected equilibrium, so S^- is not automatically observationally equivalent to S+ on every design law. The concrete two-node nonlinear example above shows the mechanism by which the laws can separate. This does not mean the paper's overall conditional stance is wrong; the positive theorems may still be sound, and a suitably scoped version of Theorem 31 (e.g., restricted to unique-equilibrium or linearly solved systems) would preserve the main message. But the theorem as currently stated is not fully proved, and the claimed impossibility is the paper's most important negative result. The verdict label remains CONDITIONAL because the paper is already conditional; however, the conditions should now include a repaired or explicitly scoped proof of Theorem 31.","tokens_in":33698,"tokens_out":20738,"duration_ms":215167,"concrete_test":"Instantiate the two-node counterexample. Let pi ~ N(0,1), U_r, U_q independent standard normal, V_r = -V_q^2 + pi + U_r, V_q = V_r + U_q, so V_q solves V_q^2 + V_q - (pi + U_r + U_q) = 0. Let Sel choose the positive root if F_r(U_r) > 1/2 and the negative root otherwise. Apply the parent-dependent rank-band reflection of Theorem 31 to U_r, with band [a, a + softplus(pi)] chosen so that a < 1/2 and a + softplus(pi) > 1/2 on a set of positive pi-probability. Simulate a large sample from S+ and S^- in a regime that does not probe r, and compare the empirical law of V_q (for instance by Kolmogorov-Smirnov distance). If the laws differ, Theorem 31's 'agrees on every law' assertion fails for this admissible Sel; if they match for many band choices and selection rules, the concern is refuted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central impossibility result (Theorem 31) constructs a twin S^- by replacing the noise mechanism at r with a parent-dependent, measure-preserving rank-band reflection, and asserts that S^- agrees with S+ on every law of every regime in the finite design. This assertion is not verified against the standing selection rule T3. T3 defines Sel as a measurable function of (i) the actual private-noise rank vector and (ii) the solution set. After the reshuffle, a draw with rank vector tau has solution set B(rho_pi(tau)), because the equations are those of S+ evaluated at the reflected rank rho_pi(tau), but Sel still receives tau as the noise rank. A measure-preserving marginal transformation on tau does not ensure that the law of Sel(tau, B(rho_pi(tau))) equals the law of Sel(tau, B(tau)). Concretely, in a two-node cycle V_r = -V_q^2 + pi + U_r, V_q = V_r + U_q, equilibria are the two roots of V_q^2 + V_q - (pi + U_r + U_q) = 0. If Sel chooses the positive root when the rank of U_r exceeds 1/2 and the negative root otherwise, a rank-band reflection that flips the threshold for a positive-measure set of tau changes the law of V_q in a non-probing regime. Thus the claim that S^- is observationally equivalent to S+ on every design law holds only when equilibria are unique (as in the linear special case) or when Sel is invariant under the reshuffle-induced map on solution sets. The theorem as stated needs an additional argument or an additional condition.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a theory of validation, transport, and identification for counterfactual queries in feedback systems modeled as equilibrium causal games (ECGs). It separates four conditions -- monotone noise mechanisms, a complete legal fibre condition, stable equilibrium selection, and query-sufficient design -- and shows how, under these conditions, population-level agreement with staged interventional distributions identifies per-unit equilibrium counterfactuals. It introduces cyclic selection diagrams for transport, gives linear-model intervention requirements, characterizes partial identification when point identification fails, and provides statistical procedures for mean and covariance comparison. It also states an impossibility theorem asserting that no finite experimental design can validate cross-world counterfactual predictions without structural class assumptions.","tokens_in":34140,"tokens_out":10959,"duration_ms":112802,"significance":"If the main claims hold, the paper is a substantive contribution to causal inference in cyclic systems, where acyclic transport and validation methods are known to fail. Its strengths are the explicit separation of abduction, selection, and design support; the necessity constructions in Theorem 8; the query-specific design framework; and the sharp partial-identification bounds in Section 9. The main qualification is that the central impossibility result, Theorem 31, is not fully proved as stated for multi-equilibrium systems.","major_comments":[{"comment":"The assertion that the reshuffled system S^- agrees with S^+ on every law of every regime in the finite design is not established for multi-equilibrium systems. After the parent-dependent rank reset rho_pi, the selected outcome under S^- is Sel(tau, B(rho_pi(tau))), whereas under S^+ it is Sel(tau, B(tau)). The map rho_pi being measure-preserving for each parent value implies only that tau and rho_pi(tau) have the same marginal law; it does not imply that the laws of Sel(tau, B(rho_pi(tau))) and Sel(tau, B(tau)) coincide when Sel depends on the actual rank vector tau, which T3 explicitly allows. The proof's statement that preserving conditional mechanism laws preserves every regime law is therefore insufficient under feedback. The theorem needs an additional argument or an additional condition, such as uniqueness of equilibrium in every relevant regime, or invariance of the selection rule under the reshuffle-induced map on solution sets. The same gap affects Corollary 32, which applies Theorem 31 in the switch-augmented model.","section":"Section 8, Theorem 31; Appendix A.3"},{"comment":"Several claims are explicitly deferred to companion preprints: payoff and curvature interventions are covered only after the reward-layer conditions of the companion reward paper hold, and the nonlinear source-block positive result cited in Corollary 7 is attributed to a companion representation paper. Because these results are used in the main text, the manuscript should either include the needed arguments or clearly mark those statements as conditional on unpublished work, so that the journal version can be assessed as a standalone contribution.","section":"Section 2.2 and Corollary 7"}],"minor_comments":[{"comment":"The notation T^+ is used before it is formally defined; please define the class T^+ (and T^+(D)) at its first occurrence.","section":"Section 3, Lemma 4"},{"comment":"There are formatting artifacts such as inline line breaks in displayed formulas and the phrase \"the y+x 2 counterexample\"; the latter should presumably read \"the y+x^2 counterexample.\" Please clean up these artifacts.","section":"Section 4, Proposition 6 and Theorem 5"},{"comment":"The standing exogenous-independence assumption excludes shared-marginal copula shifts, and the paper notes this exclusion. Because this materially limits the transport and identification statements, the limitation should be highlighted in the abstract or introduction, not only in Section 2.4.","section":"Section 2.4"},{"comment":"The statement should specify explicitly whether S^- inherits the same selection rule Sel as S^+; the proof currently leaves this implicit, which contributes to the gap described in the first major comment.","section":"Section 8, Theorem 31"},{"comment":"The rank-degenerate covariance branch is clearly scoped, but the exposition is extremely dense; a short intuitive summary of when hard-clamp versus soft-probe rank behavior differs would help readers apply the result correctly.","section":"Section 11, Proposition 50"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is unusually careful about stating conditions and limitations, and the positive results are largely well motivated. The central issue is Theorem 31: the impossibility proof needs a repair or a strengthened condition before the result can support the paper's main conclusions. The companion-preprint dependencies also deserve attention for a journal submission. I recommend major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The paper has a genuine centerpiece—Theorem 31's finite-design impossibility—and that centerpiece has a proof gap that matters. The rest is solid.\n\nWhat's new: cyclic selection diagrams, a validation/transport theory for equilibrium causal games, and linear intervention counts. The paper is unusually careful: each positive result states its conditions, Theorem 8 shows the four conditions are necessary by construction, and the limitations section is honest. The distinction between validation, transport, and identification is well drawn.\n\nThe problem: Theorem 31's construction of S^- uses a parent-dependent rank-band reflection at r. The proof claims this preserves 'every conditional mechanism law and hence every regime in the finite design.' That inference fails when the selection rule Sel depends on both the rank vector and the solution set. After the reshuffle, the equations change, so the solution set changes for some rank draws. Sel(tau, B'(tau)) can then pick a different equilibrium than Sel(tau, B(tau)), changing the law in a design regime. The two-node cycle example in the stress-test note makes this concrete. So the impossibility result is not established for multi-equilibrium systems; it needs an added condition (unique equilibria, or Sel invariant under the reshuffle-induced map on solution sets). This is not a minor caveat—it's the load-bearing result. The theorem may be repairable, but as stated it's too strong.\n\nOther soft spots: key support is deferred to two companion preprints (Section 2.2, Corollary 7), which slows verification. And a lot of the guarantees are existential—no numerical constants or code. That's acceptable for a theory paper, but worth knowing.\n\nWho gets value: anyone working on counterfactuals in feedback systems, transportability, or digital twin validation. It deserves a serious referee, but the referee should press hard on Theorem 31. I'd send it to peer review with a request for major revision, and specifically require the authors to fix the selection-rule issue.","headline":"The paper's central impossibility result has a proof gap involving the selection rule's dependence on the solution set, but the surrounding theory is careful and worth engaging.","tokens_in":34530,"tokens_out":5530,"would_cite":false,"duration_ms":49698,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62F03"],"pacs":[],"model":"deepseek-v4-flash","headline":"For feedback systems, experimental agreement alone cannot certify a digital twin's counterfactuals.","keywords":["equilibrium causal game","digital twin","counterfactual validation","transportability","cyclic structural causal model","partial identification","causal inference"],"falsifier":"Enumerate all two- and three-node linear equilibrium causal games with strictly monotone additive noises and a fixed selection rule; for a query-sufficient design, test whether any two models agree on every staged interventional distribution yet differ on a per-unit counterfactual, since Theorem 5 predicts no such pair exists, and any such pair found would refute the positive identification claim.","tokens_in":33438,"feed_emoji":"🔁","tokens_out":6552,"duration_ms":57932,"temperature":0.7,"pith_summary":"This paper asks when a digital twin of a feedback-driven system can be trusted to answer intervention questions, and when that trust survives a change of domain. It claims that, for equilibrium causal games, population-level agreement with a sufficient set of staged interventional distributions does identify per-unit counterfactuals, provided the mechanisms are strictly monotone, the equilibrium selection rule is stable, and the design is query-sufficient (Theorem 5). Without those structural assumptions, a finite collection of experiments cannot certify a cross-world prediction: two systems can match every experimental law in the design and still disagree on the counterfactual (Theorem 31). The paper therefore draws the boundary between what can be validated from data alone and what requires an a priori class assumption, and it gives transport rules for reusing a source twin in a target domain, including a hybrid that re-estimates only the changed mechanisms.","feed_headline":"Feedback twins can't be validated by experiments alone","feed_subtitle":"A new theory shows when counterfactuals are identified and when they are not.","key_machinery":"The central object is the equilibrium causal game (ECG): a cyclic structural equation model $V_i = f_i(V_{\\mathrm{pa}(i)}, U_i)$ with independent noises and a declared measurable equilibrium selection rule Sel. The argument is carried by three devices: Lemma 4 (quantile abduction under feedback), which rewrites each monotone mechanism through its interventional quantile function so that factual observations fix noise ranks; the complete legal fibre, the set of all models matching the retained observational and staged laws, on which query identification is exactly constancy; and the cyclic selection diagram $\\mathcal{D} = G \\cup \\{S_i \\to i : i \\in D\\}$, which locates the changed mechanisms. Theorem 31's witness-admissible reshuffle, a parent-dependent, measure-preserving rank reflection at a node the design never probes, is the engine of the impossibility result.","core_discovery":"The central discovery is that validation of equilibrium counterfactuals is query-specific and class-conditional. Under strict noise monotonicity (T1), a factual observation recovers each unit's private-noise rank; with a stable selection rule (T3) and a design whose experiments pin the query-relevant interventional kernels (T4), agreement on complete staged interventional distributions forces equality of full-factual per-unit counterfactuals almost everywhere (Theorem 5). Partial factual information requires a stronger singleton condition on the complete legal fibre (T2). Conversely, Theorem 31 constructs, for any finite design, a witness-admissible node whose mechanism can be reshuffled by a parent-dependent, measure-preserving transformation that is invisible to every law in the design yet changes the counterfactual by a positive gap; hence cross-world validation is possible only conditional on class membership. Transport results include cyclic selection diagrams, direct reuse under ancestral separation, and hybrid transport that replaces only the changed mechanism–noise pairs on the post-surgery ancestral closure of the query.","pith_inferences":["A testable extension: any practical validation pipeline for feedback twins should report the width of the counterfactual range implied by the complete legal fibre whenever the monotone boundary cannot be certified, rather than a single predicted value.","The impossibility result implies that certification standards relying on behavioral equivalence under a finite test suite are sound only if the structural class itself is enforced by design, for example when monotone mechanisms are physically guaranteed; this consequence for engineering practice is left implicit in the paper.","The hybrid transport rule suggests a model-agnostic recipe for domain adaptation under feedback: freeze invariant mechanisms estimated in the source, re-estimate the target rows, and re-solve the full equilibrium; the paper does not connect this to the broader domain-adaptation literature.","The separation between interventional-layer testability and cross-world non-testability could be probed empirically by fitting two models to a real feedback system that agree on all available experiments yet differ in the paper's witness construction, then measuring the counterfactual divergence under a hidden intervention."],"forward_implications":["Validation protocols for feedback twins must test complete interventional kernels, not just means and covariances; the paper shows that moment agreement does not identify distributional queries, even in linear models (a Gaussian and a shifted exponential can share every additive moment yet differ in tail probabilities).","Direct reuse of a source twin in a target is valid when the post-intervention ancestral set of the query contains no changed mechanism and the retained mechanisms satisfy cross-domain invariance; otherwise a hybrid that re-identifies only the changed mechanism–noise pairs on that ancestral closure suffices.","A finite experiment design cannot certify cross-world counterfactuals without class assumptions; Theorem 31 gives explicit witness-admissibility conditions (an unprobed node, a non-descendant parent moved by the intervention, and a nonzero resolvent coefficient) under which indistinguishable twins differ.","In linear models, identifying changed mechanisms requires an intervention count that depends on the observation model and graph support: all discrepancy rows in the ambient unknown-support branch, $m-1$ anchors for a full rotational source block, and potentially fewer for query-specific designs when the query is constant on the remaining fibre.","When point identification fails, the sharp identified set is the image of the complete legal fibre, which under semialgebraic conditions is a finite union of points and intervals with explicit endpoints."],"supporting_citations":[{"why":"Supplies the SCC/loop-solvability hypothesis and causal calculus for cyclic SCMs used in Definition 3 and the graphical transport rule.","marker":"[Forré and Mooij, 2019]"},{"why":"Provides the foundations of structural causal models with cycles and latent variables, used for interventional semantics and unique solvability.","marker":"[Bongers et al., 2021]"},{"why":"Gives the selection-diagram and transportability framework that the paper extends to cyclic equilibrium settings.","marker":"[Pearl and Bareinboim, 2014]"},{"why":"Establishes the data-fusion and transportability background for direct reuse across domains.","marker":"[Bareinboim and Pearl, 2016]"},{"why":"Defines which counterfactuals are testable in acyclic models, the baseline the finite-design impossibility result contrasts with.","marker":"[Shpitser and Pearl, 2007]"},{"why":"Provides the monotone-determinism argument for bijective causal models that Lemma 4 adapts to equilibrium structural mechanisms.","marker":"[Nasr-Esfahany et al., 2023]"},{"why":"Offers the acyclic counterfactual transportability framework that Theorem 15 extends to per-unit counterfactuals under feedback.","marker":"[Correa et al., 2022]"},{"why":"Companion paper cited for the nonlinear source-block positive identification result in the equilibrium causal game class.","marker":"[Dadgostari and Nazemi, 2026]"}],"fun_headline_variants":["Feedback twins: finite experiments can't prove counterfactual validity","Cyclic selection diagrams show when causal twins can be reused","Hybrid transport: swap changed mechanisms, keep invariant ones","Structural assumptions needed: means and covariances insufficient"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The positive results assume the exogenous noises are mutually independent across nodes in both domains (diagonal covariance), so a change of domain can only alter per-node mechanisms and marginals; if the joint noise distribution can change while every node marginal stays fixed, the validation, transport, and identification criteria do not apply.","fun_headline_variants_meta":{"raw":{"variants":["Feedback twins: finite experiments can't prove counterfactual validity","Cyclic selection diagrams show when causal twins can be reused","Hybrid transport: swap changed mechanisms, keep invariant ones","Structural assumptions needed: means and covariances insufficient"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1439,"prompt_tokens":932,"completion_tokens":507,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":440}},"tokens_in":548,"tokens_out":507,"duration_ms":5218,"temperature":1.0,"reasoning_tokens":440,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:30:33.692143+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Enumerate all two- and three-node linear equilibrium causal games with strictly monotone additive noises and a fixed selection rule; for a query-sufficient design, test whether any two models agree on every staged interventional distribution yet differ on a per-unit counterfactual, since Theorem 5 predicts no such pair exists, and any such pair found would refute the positive identification claim.","supporting_citations":[],"review_version":2}