{"id":"d408fffa-fefc-43fd-a3e7-885daa3c7016","arxiv_id":"2608.11954","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Under latent group contamination of structured outcomes, only orbit-invariant targets are identifiable, quotient reduction is lossless exactly when a parameter-free conditional law exists, and an exact quotient randomization test was constructed and applied to RxRx1 images.","lead":"This paper studies causal inference when the observed outcome is a structured object, such as an image, recorded after an unknown, possibly treatment-dependent group transformation. It characterizes which targets of the underlying outcome are recoverable, when quotienting by the transformation loses no statistical information, and builds an exact randomization test for a lattice-rigid-motion model applied to microscopy images.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Primary RxRx1 p-value rests on unverifiable conditional-uniform allocation; the theoretical core (Theorems 2.1, 3.2, 6.1) is sound.","rationale":"The central theoretical claims hold up under scrutiny. Theorem 3.2 is a clean, correct characterization of Blackwell equivalence via a common regular conditional distribution, and the quotient-faithfulness clause is essential to avoid the counterexample in Example 3.3. Theorem 2.1 is a correct observability boundary, and Theorem 6.1's lattice canonicalizer is a genuine maximal invariant for the stated finite-support image space. The exact paired-swap test is valid for the sharp null provided the conditional-uniform assignment assumption holds. The paper is transparent about this assumption and about the unverified independence of the public embeddings, and it flags the Lipschitz limitations of the approximate-contamination theorem. The reader's weakest assumption identifies the same empirical Achilles heel: the original RxRx1 allocation uniformity cannot be proven from the metadata alone, so the reported p-value's validity is contingent on an assumption that cannot currently be checked. The additional concerns (code not yet public, primary pair selection, coarse p-value resolution) are secondary. The reader's CONDITIONAL verdict is appropriate: the theoretical contribution is sound, but the headline empirical result should not be taken at face value until the allocation assumption is either verified through original randomization logs or shown to be robust through a sensitivity analysis.","tokens_in":14098,"tokens_out":32687,"duration_ms":355747,"concrete_test":"Obtain the original RxRx1 plate randomization algorithm or logs (or the authors' audit trail) and verify that, conditional on each matched confirmation pair and the stated design information, the probability of the observed labeling is exactly 1/2 per pair, independently across blocks. If logs are unavailable, perform a sensitivity analysis: enumerate the maximum paired-swap p-value for the primary contrast over all within-plate allocation mechanisms consistent with the metadata audit (e.g., all distributions on {0,1}^8 satisfying the observed plate/batch constraints); if any plausible mechanism yields p>0.05, the primary conclusion is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is in the empirical demonstration, not the theoretical core. Theorem 7.1's exact paired-swap p-value requires that, conditional on the unordered eligible wells, their complete potential outcomes and pretreatment design information, the assignment vector Z is uniform on {0,1}^B. The paper itself states (Section 7, repeated in Section 10) that the metadata audit can establish that paired siRNAs occur within the same plate randomization set but cannot prove the original allocation algorithm was uniform. If the original RxRx1 within-plate randomization was not exactly uniform over the 2^8 swap patterns—for example, because the eight confirmation pairs were formed post hoc from a plate-level or batch-level randomization structure—then Theorem 7.1 does not apply to the reported contrast, and the primary p-value of 0.0078 is not a valid randomization p-value for the real experiment. This is load-bearing for the confirmatory causal conclusion, though it does not undermine Theorems 2.1, 3.2, or 6.1, which appear correct as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a theory of causal inference for structured outcomes observed as X = Γ·Y(A), where Γ is an unobserved group element whose distribution may depend arbitrarily on treatment, covariates and intrinsic outcomes. It proves an observability theorem (measurable targets are uniformly recoverable iff they are G-invariant; a maximal invariant factorizes all invariant targets), a losslessness theorem (quotient reduction is Blackwell-equivalent to the full experiment exactly when a parameter-free, quotient-faithful reconstruction kernel exists, with conditional-Haar contamination as a special case), a product-versus-diagonal action dichotomy, and an approximate-contamination Wasserstein/MMD stability bound. The paper then constructs a maximal invariant for multichannel lattice images under integer translations and quarter turns, defines a characteristic Gaussian kernel on the quotient, and combines these with a complete paired-swap randomization test under the sharp null. The methods are applied to simulations and to an RxRx1 HUVEC confirmation experiment, with a reported primary paired-swap p-value of 0.0078. The text is unusually explicit about its assumptions and limitations.","tokens_in":14199,"tokens_out":22068,"duration_ms":224599,"significance":"If the main theorems hold (and they appear correct), the paper makes a valuable conceptual contribution by separating observability from statistical sufficiency for group-contaminated outcomes: Theorem 3.2 gives a precise condition under which quotienting is decision-theoretically lossless, and the lattice maximal invariant of Theorem 6.1 provides an explicit noncompact construction with a characteristic kernel. The paired-swap randomization test in Theorem 7.1 is clean, and the simulation evidence (0.052 rejection at the null, 0.992 power at effect strength 1.0) is supportive. The authors are also to be credited for flagging the unverifiable assumptions behind the RxRx1 p-value and for refusing to present Theorem 5.1 as certifying the approximate-action stress test. The main uncertainty is about the empirical demonstration, where the headline p-value depends on assumptions that metadata cannot verify.","major_comments":[{"comment":"Theorem 7.1's exactness requires that, conditional on the unordered eligible wells, their complete potential outcomes and pretreatment design information, Z is uniform on {0,1}^B. The manuscript states (Section 7, repeated in Section 10) that the metadata audit can establish that paired siRNAs occur within the same plate randomization set but cannot prove the original allocation algorithm was uniform. Since the abstract's headline p-value of 0.0078 is a randomization p-value for the real RxRx1 experiment only under this assumption, the confirmatory language in the abstract and Section 9.2 overstates the evidence. Please either carry this caveat into the abstract and the results section, explicitly labeling the RxRx1 finding as conditional on an unverifiable uniform-within-plate allocation, or add a sensitivity analysis over plausible non-uniform allocation mechanisms (for example, plate-level or batch-level randomization) to quantify how the p-value would change.","section":"Section 7 and Abstract"},{"comment":"The selection of the ten confirmation pairs uses public 128-dimensional embeddings from discovery rows only, but the paper concedes that 'the row split does not prove that the public embedding generator was trained independently of every confirmation image.' If the embedding model was trained on confirmation images, the selected pairs are not outcome-independent, and the 'prespecified' characterization of the primary and secondary contrasts is weakened. Please provide evidence of the embedding generator's training provenance, or run a sensitivity analysis with alternative selection rules (for example, random eligible pairs or embeddings from a model trained exclusively on discovery images), and otherwise describe the RxRx1 contrasts as exploratory rather than confirmatory.","section":"Section 9.1"}],"minor_comments":[{"comment":"Q_1(z) and Q_0(z) are written with set braces, but the MMD statistic uses them as indexed lists; use multisets or tuples to avoid ambiguity when quotient outcomes coincide.","section":"Section 7, Eq. (7.2)"},{"comment":"The key vector in Theorem 6.1 is described as containing bytes in 'site–channel–row–column order'; for a single image there is no site index, so the description should be 'channel–row–column order'.","section":"Section 6.1"},{"comment":"The header arrangement makes it unclear which rows correspond to which method; add an explicit Method column and state that |Δp| is the change in the paired-swap p-value, not a p-value itself.","section":"Table 5"},{"comment":"The phrase 'retain the original simulations and RxRx1 HUVEC study' is vague; specify that these analyses are preserved from the prior workflow and that the RxRx1 results are conditional on the assumptions listed in Section 10.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The theoretical core (Theorems 2.1, 3.2, 6.1 and the randomization result) appears correct and publishable. The main risk is that the headline RxRx1 p-value is presented prominently while depending on two unverifiable data-pipeline assumptions: conditional-uniform within-plate allocation and algorithmic independence of the embedding generator from confirmation images. The authors acknowledge both assumptions in the text, but the abstract and results sections do not carry the caveat. If the authors temper the confirmatory language or add sensitivity analyses, the paper would be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The part worth reading is the theory. The paper cleanly separates three questions people usually blur: what is observable from a single transformed outcome (only orbit-invariant targets), what the interventional quotient law is identified by (standard adjustment on the quotient), and when discarding the within-orbit coordinate is statistically lossless (only when a parameter-free conditional law given the quotient exists, equivalently a quotient-faithful kernel). Theorem 3.2 is the real contribution, and it is correct. Corollary 3.4 is a nice special case, not the main model. The product-versus-diagonal action result is simple but practically useful, and the lattice maximal invariant construction is explicit and checkable. The proofs are short but valid. The paper also gets credit for honesty: it repeatedly flags that maximal invariance is not sufficiency, that the approximate-contamination bound is conditional on Lipschitz regularity it does not verify, and that the metadata audit cannot prove the original RxRx1 allocation was uniform.\n\nThe soft spots are empirical, not theoretical. The primary p-value of 0.0078 is an exact randomization p-value only if the original within-plate allocation was conditionally uniform over the 2^8 swap patterns. The paper itself says this cannot be proven from metadata. That is a load-bearing assumption for the confirmatory conclusion, and the stress-test note lands. On top of that, the primary pair was selected from ten candidates by a score, and the reported p-value is not adjusted for that selection. The paper does adjust the nine secondary contrasts, but the primary contrast was chosen precisely because its split-half score was highest, so the raw 0.0078 overstates the evidence against the null. Also, the reproducibility statement says the code archive is not yet released; that is a concrete gap, though an addressable one.\n\nNone of this damages Theorems 2.1, 3.2, or 6.1. The simulation is sensible and the type I error is fine. The RxRx1 analysis is better described as a controlled unknown-acquisition demonstration on real images than as a confirmatory biological finding.\n\nThis paper deserves a serious referee. The theoretical contribution is worth engaging even if the empirical section needs revision. The authors should be asked to release code, to account for selection in the primary p-value or clearly label it as a discovery result, and to state plainly that the causal conclusion is conditional on an unverifiable allocation mechanism. I would cite the sufficiency-boundary theorem in my own work on structured outcomes.","headline":"A sound theoretical core on observability vs. sufficiency for group-contaminated outcomes, with an empirical demonstration whose confirmatory p-value rests on an unverifiable allocation assumption and unadjusted pair selection.","tokens_in":14795,"tokens_out":956,"would_cite":true,"duration_ms":11502,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62B15","62D20","62B05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Quotienting group-contaminated outcomes is statistically lossless exactly when the conditional law of the raw observation given the quotient is parameter-free.","keywords":["causal inference","group action","maximal invariant","Blackwell sufficiency","microscopy image","quotient space","randomization test","structured outcome"],"falsifier":"Re-run the paired-swap analysis on the eight RxRx1 confirmation blocks under a documented non-uniform within-plate assignment (for example, $P(Z_b = 1) = 0.7$ in some blocks) and check whether the sharp-null rejection rate exceeds the nominal level; if it does, the conditional-uniformity assumption behind Theorem 7.1 and the reported $p = 0.0078$ has been violated. Independently, the equivalence in Theorem 3.2 can be attacked directly: on a small finite group example, exhaustively search for a parameter-free quotient-faithful kernel $K$ with $P^{\\mathrm{quot}}_\\theta K = P^{\\mathrm{full}}_\\theta$ while $\\mathrm{Law}_\\theta(X \\mid C, A, M(X))$ still depends on $\\theta$; the theorem asserts that no such family exists.","tokens_in":13794,"feed_emoji":"🔬","tokens_out":19522,"duration_ms":169075,"temperature":0.7,"pith_summary":"This paper studies experiments in which structured outcomes, such as microscopy images, are recorded after an unknown unit-specific transformation $\\Gamma$ that may depend on treatment, covariates, and the outcome itself, so the observed object is $X = \\Gamma \\cdot Y(A)$ rather than the intrinsic potential outcome $Y(A)$. It separates three questions: a target is uniformly recoverable exactly when it is constant on group orbits; the interventional law of the quotient is identified by the ordinary covariate-adjustment formula; and quotient reduction is statistically lossless exactly when the conditional law of the raw observation given treatment, covariates, and the quotient has a parameter-free version, making the quotient experiment Blackwell-equivalent to the full experiment. This third boundary is the paper's central discovery because it distinguishes observability (invariance) from statistical sufficiency. The framework is carried through to an exact paired-swap randomization test for lattice images, which under the sharp null rejects in $0.052$ of simulation replicates at the nominal level, reaches power $0.992$ at unit effect strength, and reports $p = 0.0078$ for the primary RxRx1 HUVEC contrast.","feed_headline":"One condition decides when quotienting transformed data is lossless","feed_subtitle":"Maximal invariance shows what is observable; this theorem pins down when the quotient experiment equals the full one.","key_machinery":"The machinery is the comparison of two statistical experiments, the full experiment $(C, A, X)$ and the quotient experiment $(C, A, M(X))$, in the Blackwell sense. The object carrying the argument is the quotient-faithful kernel: a Markov kernel $K$ from the quotient sample space to the full sample space that is supported on the fibre of $T$ over $s$ and reproduces every full-experiment law through $P^{\\mathrm{quot}}_\\theta K = P^{\\mathrm{full}}_\\theta$. Theorem 3.2 shows that such a parameter-free kernel exists precisely when the conditional law of the raw observation given the quotient is parameter-free, with the proof running through the disintegration identity for a common regular conditional distribution. Supporting constructions include the maximal invariant $M$, whose equality classes are exactly the group orbits, together with a Borel factorization theorem for invariant targets; the orbit kernel formed by integrating a Borel section against normalized Haar measure, which yields the conditional-Haar equivalence corollary; and, at the application level, the lattice canonicalizer that removes the translation normalizing the support bounding box and then selects the lexicographically smallest of the four quarter-turn variants, producing a Borel maximal invariant for $\\mathbb{Z}^2 \\rtimes C_4$ without placing a probability law on the noncompact translation group.","core_discovery":"The paper's central claim is Theorem 3.2: under the unrestricted contamination model $X = \\Gamma \\cdot Y(A)$, the quotient experiment based on a maximal invariant $M$ is sufficient for, and Blackwell-equivalent to, the full experiment if and only if there exists a parameter-free quotient-faithful Markov kernel $K$ satisfying $P^{\\mathrm{quot}}_\\theta K = P^{\\mathrm{full}}_\\theta$ for every $\\theta$, equivalently if and only if the conditional law $\\mathrm{Law}_\\theta(X \\mid C, A, M(X))$ admits a version independent of $\\theta$. Conditional Haar contamination on a compact group is derived as the special case that guarantees this condition, but it is not imposed in the main model. The paper also proves that invariance is exactly what makes a target observable (Theorem 2.1), that every measurable invariant target factors through a Borel maximal invariant (Theorem 2.4), that interventional quotient laws are identified by adjustment (Theorem 2.5), that componentwise canonicalization is maximal for independent sitewise actions but erases relative configuration under a shared diagonal action (Proposition 4.1), and that for finite-support multichannel lattice images a lexicographic canonicalizer is a Borel maximal invariant under integer translations and quarter turns (Theorem 6.1). Together with a characteristic Gaussian kernel on the quotient and exhaustive enumeration of the paired assignment distribution, these results ground an exact finite-sample randomization test (Theorem 7.1).","pith_inferences":["The observability-sufficiency gap yields a diagnostic for any proposed invariant summary: test whether the conditional law of the raw outcome given the summary varies with the treatment parameter; if it does, the summary is not statistically lossless even when it is maximal.","The product-versus-diagonal split points to an intermediate model the paper does not pursue: a partially shared action, in which some sites share one motion while others do not, would require a quotient retaining relative pose only within the sharing structure, and the testable question is whether such relative configuration predicts the outcome.","The lattice canonicalizer embodies a transferable recipe, an explicit cross-section for the noncompact part of the group followed by finite lexicographic minimization over the compact residual, which should extend to other semidirect products with countable orbit structure, such as lattice actions with reflections or scaling.","Because the exact test enumerates only $256$ assignments, its $p$-values are coarse; combining the quotient kernel with more blocks or a continuous statistic would sharpen resolution while preserving the design-based validity guarantee."],"forward_implications":["If the quotient experiment is Blackwell-equivalent to the full experiment, every bounded decision problem has the same attainable risk set under both, so analyses of invariant targets can be conducted on the quotient without sacrificing any experiment-relevant information.","The paired-swap test is exactly valid under the Fisher sharp null: rejection when the enumerated $p$-value is at most $\\alpha$ has conditional probability at most $\\alpha$, and in the simulations the test rejected in $0.052$ of replicates at the nominal $0.05$ level with power $0.992$ at unit effect strength.","Sitewise canonicalization is the correct quotient when sites undergo independent motions, but under one shared diagonal motion it erases relative configuration such as inter-site displacement; a diagonal quotient is required to preserve that information.","Under explicit Lipschitz regularity, approximate contamination propagates linearly: orbit-metric error $\\varepsilon_a$ implies quotient-law Wasserstein error at most $L_M \\varepsilon_a$ and population-MMD perturbation at most $L_k L_M (\\varepsilon_1 + \\varepsilon_0)$; the paper leaves the lexicographic canonicalizer's Lipschitz status open, so its perturbation table remains an empirical sensitivit","The primary RxRx1 HUVEC contrast is significant at $p = 0.0078$ with exact invariance holding across all imposed acquisition seeds, while no secondary contrast survives Holm adjustment; the eight-block design gives the enumeration a coarse resolution of $2^{-8}$."],"supporting_citations":[{"why":"supplies the comparison-of-experiments framework in which the losslessness theorem is stated.","marker":"[2]"},{"why":"provides the standard decision-theoretic theory of experiment comparison used for Blackwell dominance and equivalence.","marker":"[12]"},{"why":"serves as the reference for comparing statistical experiments on which the equivalence argument leans.","marker":"[23]"},{"why":"supplies the Lusin separation and measurable factorization tools behind the Borel factorization theorem for maximal invariants.","marker":"[10]"},{"why":"is the design-based randomization-inference principle that the paired-swap test instantiates.","marker":"[6]"},{"why":"provides the finite-sample randomization-test validity theory used in the proof of the exact test.","marker":"[13]"},{"why":"defines the kernel two-sample MMD statistic that the exact paired-swap test uses as its discrepancy measure.","marker":"[7]"},{"why":"establishes the characteristic property of the Gaussian kernel on Euclidean space used to show the quotient kernel separates all laws.","marker":"[21]"},{"why":"is the public RxRx1 dataset source for the confirmation experiment analyzed in Section 9.","marker":"[22]"}],"fun_headline_variants":["When is quotient reduction lossless? A single iff condition","Exact randomization test for group-contaminated outcomes","No loss in reducing transformed data iff conditional law is parameter-free","One theorem characterizes lossless quotient reduction of contaminated data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The real-data conclusion rests on the assumption that the original within-plate allocation of the two siRNAs per block was conditionally uniform: the metadata audit can establish that the paired siRNAs belonged to the same randomization set, but, as the paper states in Section 9.2, it cannot itself prove that the allocation algorithm was uniform, and if it was not, the reported paired-swap $p$-values are not valid randomization $p$-values for the actual experiment.","fun_headline_variants_meta":{"raw":{"variants":["When is quotient reduction lossless? A single iff condition","Exact randomization test for group-contaminated outcomes","No loss in reducing transformed data iff conditional law is parameter-free","One theorem characterizes lossless quotient reduction of contaminated data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000409,"raw_usage":{"total_tokens":2223,"prompt_tokens":1148,"completion_tokens":1075,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":764,"completion_tokens_details":{"reasoning_tokens":1010}},"tokens_in":764,"tokens_out":1075,"duration_ms":10934,"temperature":1.0,"reasoning_tokens":1010,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:22:22.768801+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the paired-swap analysis on the eight RxRx1 confirmation blocks under a documented non-uniform within-plate assignment (for example, $P(Z_b = 1) = 0.7$ in some blocks) and check whether the sharp-null rejection rate exceeds the nominal level; if it does, the conditional-uniformity assumption behind Theorem 7.1 and the reported $p = 0.0078$ has been violated. Independently, the equivalence in Theorem 3.2 can be attacked directly: on a small finite group example, exhaustively search for a parameter-free quotient-faithful kernel $K$ with $P^{\\mathrm{quot}}_\\theta K = P^{\\mathrm{full}}_\\theta$ while $\\mathrm{Law}_\\theta(X \\mid C, A, M(X))$ still depends on $\\theta$; the theorem asserts that no such family exists.","supporting_citations":[{"cited_title":"(1953) Equivalent comparisons of experiments.Ann","cited_arxiv_id":null,"evidence_quote":"supplies the comparison-of-experiments framework in which the losslessness theorem is stated."},{"cited_title":"(1986)Asymptotic Methods in Statistical Decision Theory","cited_arxiv_id":null,"evidence_quote":"provides the standard decision-theoretic theory of experiment comparison used for Blackwell dominance and equivalence."},{"cited_title":"(1991)Comparison of Statistical Experiments","cited_arxiv_id":null,"evidence_quote":"serves as the reference for comparing statistical experiments on which the equivalence argument leans."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the Lusin separation and measurable factorization tools behind the Borel factorization theorem for maximal invariants."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"is the design-based randomization-inference principle that the paired-swap test instantiates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the finite-sample randomization-test validity theory used in the proof of the exact test."},{"cited_title":"M., Rasch, M","cited_arxiv_id":null,"evidence_quote":"defines the kernel two-sample MMD statistic that the exact paired-swap test uses as its discrepancy measure."},{"cited_title":"K., Gretton, A., Fukumizu, K., Sch¨ olkopf, B","cited_arxiv_id":null,"evidence_quote":"establishes the characteristic property of the Gaussian kernel on Euclidean space used to show the quotient kernel separates all laws."}],"review_version":1}