{"id":"61cfcbbf-fa1c-4d9b-85e4-0ac2506a9150","arxiv_id":"2510.17303","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Averaging a model over a symmetry group with a data-dependent kernel shrinks the KL term in PAC-Bayes bounds, giving tighter guarantees for non-compact groups and non-invariant data.","lead":"This paper extends PAC-Bayes generalization guarantees to models with symmetries such as translation and rotation, even when the data distribution is not invariant under those symmetries. It gives a theoretical argument for why symmetric models can be preferable, and demonstrates tighter guarantees on rotated MNIST.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1.2 is false as stated for unbounded losses; Thm 2.6/3.7 inherit the defect and the cross-entropy experiments are not covered.","rationale":"The reader's weakest assumption is the free-action condition. That is a real scope limitation, but the single most load-bearing issue for the central claim is the unbounded-loss gap in the baseline PAC-Bayes theorem, which makes the main bounds false in the stated generality. This is an internal inconsistency: Theorem 1.2's claim for 'any measurable loss' contradicts the standard McAllester bound and fails on a simple two-point heavy-tailed distribution. The symmetry construction may be salvageable by assuming bounded losses and adjusting the experiments, but the manuscript as written overclaims. Since the reader already returned CONDITIONAL, I keep the same categorical verdict but ground it in a more severe correctness concern; hence verdict_should_be=UNCHANGED. I partly agree with the reader's free-action concern, but I regard the loss-boundedness flaw as more load-bearing.","tokens_in":14046,"tokens_out":24180,"duration_ms":211191,"concrete_test":"Run the counterexample numerically: let H={f≡0}, prior=posterior=δ_f, squared loss, Y=0 with probability 0.99 and Y=100 with probability 0.01, n=100, δ=0.05. For the all-zero sample (probability ≈0.366), compute the RHS of Theorem 1.2: sqrt((0 + log(1/δ) + log n + 2)/(2n−1)) ≈ 0.22, while Q[R]=100. If 100 > 0.22, the theorem is disproved. Then add the bounded-loss assumption ℓ∈[0,1] and re-derive Theorems 2.6 and 3.7; if the proofs require the same boundedness, the paper must state it and the experiments must be rerun with a bounded loss.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Load-bearing flaw: Theorem 1.2 is stated for any measurable ℓ:Y×Y→[0,∞), but the McAllester square-root bound with constant (log(1/δ)+logn+2)/(2n−1) requires a loss bounded in [0,1] (or a sub-Gaussian assumption). Without boundedness the theorem is false. Counterexample: H={f≡0}, squared loss ℓ(a,b)=(a−b)^2, Y=0 with probability 0.99 and Y=100 with probability 0.01. Then E[ℓ]=100 and KL(Q∥P_H)=0. For n=100, δ=0.05, the claimed RHS is at most about 0.22 whenever all samples have Y=0, which occurs with probability 0.99^100≈0.366>δ. Thus the high-probability inequality fails. Theorem 2.6 and 3.7 are corollaries of Theorem 1.2, so the main guarantees inherit the false 'any measurable loss' claim. The experiments use cross-entropy, an unbounded loss, so the reported bounds are not justified by the stated theory. This is more central than the free-action restriction: it breaks the theorems even in the free, compact-group setting.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript develops a PAC-Bayesian framework for equivariant models under general (possibly non-compact) group actions and non-invariant data distributions. The main technical tool is an averaging operator that maps arbitrary hypotheses to equivariant ones via a disintegration of the data distribution over orbit representatives. Theorem 2.5 proves a KL decomposition for this operator; Theorem 2.6 derives a PAC-Bayes bound with a reduced KL term; Proposition 3.4 states that, under symmetric data, averaging does not increase the true risk; Theorem 3.7 gives a bound using empirical risk on orbit representatives. The paper validates the approach on rotated and translated MNIST with cross-entropy loss.","tokens_in":14384,"tokens_out":8424,"duration_ms":69009,"significance":"If the results were correct, the paper would extend symmetry-based PAC-Bayes analysis beyond compact groups and invariant distributions, providing theoretical support for the practical success of equivariant models. The KL-decomposition lemma (Thm 2.5) is conceptually clean and general, and the appendix contains detailed measure-theoretic proofs and a fully worked Gaussian toy example. However, the central theorems rely on a McAllester-type bound that is false for unbounded losses, and the free-action assumption excludes the rotation symmetries used in the experiments. These issues must be resolved before the contribution can be considered sound.","major_comments":[{"comment":"Theorem 1.2 is stated for 'any measurable loss function ℓ:Y×Y→[0,∞)' but uses the square-root McAllester bound with denominator 2n−1. This bound requires a loss bounded in [0,1] (or a sub-Gaussian tail condition). Without boundedness the theorem is false: take H={f≡0}, squared loss ℓ(a,b)=(a−b)^2, Y=0 w.p. 0.99 and Y=100 w.p. 0.01, n=100, δ=0.05, and Q=P_H. Then KL=0 and the RHS is below 0.22 whenever all samples have Y_i=0, an event of probability 0.99^100≈0.366>δ, while the true risk is 100. Thus the claimed high-probability inequality fails. Since Theorems 2.6 and 3.7 are direct consequences of Theorem 1.2, they inherit the defect. The experiments use cross-entropy, which is unbounded, so the reported bounds are not justified by the stated theory. The authors should either restrict all statements to bounded losses (e.g., [0,1]) and re-run experiments with a bounded surrogate, or repla","section":"§1.2, Theorem 1.2 (and Theorems 2.6, 3.7)"},{"comment":"The entire construction requires the G-action on X to be free. This is not satisfied for the rotation group acting on images: any rotationally symmetric image — e.g., a blank image or a symmetric digit — has a non-trivial stabilizer. The paper's experiments use rotated MNIST, so Assumption 2.1.2 is violated in exactly the empirical setting that is claimed to validate the theory. The assumption is introduced only 'for simplicity' and no relaxation is discussed. To make the experimental claims covered, the authors must either handle actions with stabilizers (e.g., by passing to the quotient by stabilizers) or restrict the data support to an open free subset, which is not done.","section":"§2, Assumption 2.1.2"}],"minor_comments":[{"comment":"Typo: 'the cAllester’s PAC-Bayesian boun' should be 'McAllester's PAC-Bayesian bound'.","section":"§3.1"},{"comment":"The theorem states the bound in terms of S_n^φ but the final sentence defines only S_n. Define S_n^φ explicitly as n i.i.d. copies of (X_φ, Y_φ).","section":"Theorem 3.7"},{"comment":"The pushforward covariance Σ is singular; the KL computation is only valid after restricting the measures to the subspace U. This should be stated explicitly to avoid confusion about Gaussian KL on R^2.","section":"Appendix C, Example C.1"},{"comment":"The statement that the kernel equals the Haar measure 'PX-almost surely' is ambiguous. Clarify that this holds for the conditional kernel under the invariant-data assumption.","section":"Remark 2.3"}],"recommendation":"major_revision","confidential_remarks":"The unbounded-loss counterexample is decisive and should be addressed before further review. The free-action issue is also significant because the experimental setting (rotated MNIST) has stabilizers. If the authors can re-scope the theory to bounded losses and handle or avoid stabilizers, the KL-decomposition idea may be publishable; as it stands, the main theorems are false in the stated generality."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has a genuinely interesting idea: replace the Haar-measure averaging in symmetry-based PAC-Bayes with a kernel from the disintegration theorem, so you can handle non-compact groups and non-invariant distributions. The construction of Q and the KL decomposition (Thm 2.5) are sound under the free-action assumption, and the representative-set bound (Thm 3.7) is a nice addition. This is a real contribution to the equivariant learning theory.\n\nThat said, there is a load-bearing flaw. Theorem 1.2 is stated for \"any measurable loss function\", but the McAllester bound with that constant requires the loss to be bounded in [0,1] or a sub-Gaussian condition. The claim as written is false. A simple counterexample: H contains only the zero function, squared loss, and Y that is 0 with prob 0.99 and 100 with prob 0.01. Then the expectation is 100, KL=0, and the claimed bound fails with non-negligible probability. Since Theorem 2.6 and 3.7 are corollaries of Theorem 1.2, they inherit the defect. The experiments use cross-entropy, which is unbounded, so the reported bounds are not justified by the stated theory.\n\nThe free-action assumption is also restrictive. Many natural group actions have nontrivial stabilizers (e.g., rotations on symmetric images), and the entire construction relies on the bijection G×Xφ→X. The paper says \"for simplicity\" but doesn't discuss how to relax it. That's a limitation, not an error.\n\nFinally, the abstract claims \"several datasets\" but the experiments are two variants of MNIST with restricted rotations and translations. The innovation is more modest than advertised.\n\nIf the authors fix the loss-boundedness issue—state the theorems for bounded losses or add a sub-Gaussian assumption—the core results would hold, and the representative-set bound would still be interesting. The experiments should either use a bounded loss or the theory must be extended. This is worth a serious referee; the idea is solid enough to engage, but the paper needs revision before publication. I would not cite it in its current form.","headline":"Original idea undermined by a false base bound for unbounded losses; needs revision.","tokens_in":14795,"tokens_out":3055,"would_cite":false,"duration_ms":26137,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Averaging hypotheses over a symmetry group shrinks the KL-divergence term in PAC-Bayes bounds and cannot increase true risk, extending symmetry-based generalization guarantees to non-compact groups and non-invariant data.","keywords":["PAC-Bayes","generalization bounds","equivariance","non-compact symmetry","non-invariant data","KL divergence","averaging operator","translation symmetry"],"falsifier":"Rerun the paper's rotated-MNIST experiment with rotation angles spanning the full circle, including 180° where digits like '0' and '8' are fixed by the rotation (a non-free action), and check whether the equivariant model's PAC-Bayes bound is still tighter and whether the averaging operator is well-defined for inputs with nontrivial stabilizers. If the bound degrades or Q depends on the chosen orbit representative, the extension to general non-compact symmetries is limited to free actions.","tokens_in":13975,"feed_emoji":"🔄","tokens_out":11429,"duration_ms":95602,"temperature":0.7,"pith_summary":"The paper tries to show that symmetry can be converted directly into tighter generalization guarantees: take any hypothesis distribution, average each hypothesis over the symmetry group to make it equivariant, and the KL-divergence term in a standard PAC-Bayes bound can only decrease. In addition, if the data share the symmetry—an equivariant target function and a group-invariant convex loss—the averaged hypothesis has true risk no larger than the original. If correct, this moves the theory of symmetric models beyond compact groups and perfectly invariant distributions, covering practically important cases such as translations and partially symmetric data. The authors demonstrate the mechanism on a standard PAC-Bayes bound and confirm on rotated and translated MNIST that equivariant networks attain both lower test error and a tighter generalization bound than a non-symmetric baseline.","feed_headline":"Averaging over symmetries shrinks KL, tightening PAC-Bayes bounds","feed_subtitle":"Projecting hypotheses onto equivariant functions tightens generalization bounds for translations and non-invariant data.","key_machinery":"The averaging operator Q (Definition 2.2): it sends a measurable hypothesis f to πG(x)·∫ g^{-1}·f(g·πXφ(x)) κ(πXφ(x), dg), where πXφ selects the representative of the orbit of x, πG gives the group element taking that representative to x, and κ is the disintegration kernel of the input distribution over orbit representatives. Q is a measurable projection onto the subspace of equivariant functions. Its role is to feed into the KL-decomposition lemma for pushforward measures: pushing two measures through a measurable map cannot increase their KL divergence, and Q isolates the equivariant component, leaving a nonnegative remainder that describes the non-equivariant difference. This is the mecha","core_discovery":"The central claim is a KL-divergence decomposition for the averaging operator Q: for any two distributions on a hypothesis class closed under Q, DKL(μ∥ν)=DKL(Q∗μ∥Q∗ν)+∫ log((dμ/dν)/(dQ∗μ/dQ∗ν)) dμ, so pushing both measures through the equivariant projection can only reduce the divergence. Inserting this into a standard PAC-Bayes bound gives a bound whose complexity term is no larger for the equivariant posterior (Theorem 2.6). When the data come from an equivariant target function plus independent noise and the loss is group-invariant and convex in its first argument, averaging a hypothesis also cannot increase its true risk (Proposition 3.4); and for equivariant hypotheses the risk and empi","pith_inferences":["Beyond the paper: if the free-action assumption were relaxed by quotienting out stabilizers, the construction would likely extend to symmetries with fixed points, such as rotations acting on digits like '8' or '0'; the abstract KL inequality only needs measurability of Q, not freeness, so the obstruction is the representative-based definition rather than the bound itself.","Beyond the paper: the representative-set reduction suggests a concrete training recipe—train an equivariant model on one sample per orbit and evaluate the PAC-Bayes bound there; this could be tested by comparing generalization gaps on quotient datasets.","Beyond the paper: because the improvement comes only from the KL term, the empirical-risk term is untouched; a testable prediction is that the bound improvement shrinks continuously as the data distribution becomes less symmetric, disappearing when the remainder term vanishes.","Beyond the paper: the same pushforward-KL trick could be applied to other complexity measures in PAC-Bayes bounds, but the remainder term would have to be re-derived for each."],"forward_implications":["Convolutional (translation-equivariant) models fall under PAC-Bayes guarantees, since translations are non-compact; the paper claims these are the first such bounds.","When data are symmetric, equivariant models are preferable to non-equivariant ones: their generalization bound is no worse and their true risk is no larger.","For equivariant hypotheses, the empirical risk can be estimated on one representative per orbit, so training and evaluation can be restricted to a reduced input set without weakening the guarantee.","The KL-decomposition argument is not tied to the particular baseline bound; it transfers to other PAC-Bayes formulations that use a KL complexity term.","In the special case of compact groups and invariant data, the averaging operator reduces to classical Haar-measure group averaging, so prior results are contained as a special case."],"fun_headline_variants":["Symmetry averaging tightens PAC-Bayes bounds beyond compact groups","Projecting to equivariant functions shrinks KL, sharpens bounds","PAC-Bayes guarantees for non-invariant data via symmetry","Averaging hypotheses reduces KL, tightens generalization bounds"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The construction assumes the symmetry group acts freely on the input space—no non-identity transformation fixes any input—so every input decomposes uniquely into a group element and an orbit representative; natural symmetries such as rotations on images with rotational symmetry (or the all-zero image) violate this.","fun_headline_variants_meta":{"raw":{"variants":["Symmetry averaging tightens PAC-Bayes bounds beyond compact groups","Projecting to equivariant functions shrinks KL, sharpens bounds","PAC-Bayes guarantees for non-invariant data via symmetry","Averaging hypotheses reduces KL, tightens generalization bounds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000145,"raw_usage":{"total_tokens":1006,"prompt_tokens":722,"completion_tokens":284,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":212}},"tokens_in":466,"tokens_out":284,"duration_ms":3157,"temperature":1.0,"reasoning_tokens":212,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T09:04:09.168175+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the paper's rotated-MNIST experiment with rotation angles spanning the full circle, including 180° where digits like '0' and '8' are fixed by the rotation (a non-free action), and check whether the equivariant model's PAC-Bayes bound is still tighter and whether the averaging operator is well-defined for inputs with nontrivial stabilizers. If the bound degrades or Q depends on the chosen orbit representative, the extension to general non-compact symmetries is limited to free actions.","supporting_citations":[],"review_version":1}