{"id":"2c16f2e3-ed23-4b6f-8de9-63eaac590881","arxiv_id":"2411.15328","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Optimal features for a broad family of losses, including cross entropy and hinge loss, factor as a loss-dependent function composed with the HGR maximal correlation functions, making them dependence induced and minimal sufficient.","lead":"This paper proves that, for two random variables, any representation that depends only on how they are statistically related must be a function of the HGR maximal correlation functions, and it identifies a family of loss functions, including cross entropy and hinge loss, whose optimal features take exactly this form.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No load-bearing flaw in the central D-loss argument; however, Section V-B's claim that -H_nest is a regular D-loss is false.","rationale":"The reader's conditional verdict centered on unproved Fact 1 and Property 3. We supply proofs for both, so the central theorem stands. We did find a separate false assertion in Section V-B, but it is peripheral to the paper's main contribution. Therefore, we agree with the reader's overall conditional stance: the paper needs minor corrections and proof sketches before full acceptance, but the central claim is sound.","tokens_in":865,"tokens_out":8911,"duration_ms":566935,"concrete_test":"For the binary joint distribution with P(0,0)=P(1,1)=0.4 and P(0,1)=P(1,0)=0.1, set f(X)=1-2X, g(Y)=1-2Y, and s(X)=t(Y)=constant. Compute H(f,g)=0.1 and H(f|s,g|t)=0; then -H violates the D-loss condition, disproving the Section V-B claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, Theorem 3, is supported. Fact 1 holds: by conditional Jensen on the log-sum-exp functional, first conditioning on the outer X and then on the inner Y, one obtains Γ(f|s,g|t) ≤ Γ(f,g) for the Markov chain X-s(X)-t(Y)-Y. Property 3 also holds because equality in the regularized inequality forces the L2 norms to be equal, which by strict convexity gives f = f|s and g = g|t. Thus the reader's main concerns are resolved. However, Section V-B states that -H_nest is a regular D-loss; this is false. For example, take X,Y with P(0,0)=P(1,1)=0.4 and P(0,1)=P(1,0)=0.1, let f(X)=1-2X and g(Y)=1-2Y, and choose constant s,t. Then H(f|s,g|t)=0 and H(f,g)=0.1, so -H(f|s,g|t)=0 > -H(f,g)=-0.1, violating the D-loss inequality (9). This error is peripheral: it does not affect Theorem 3 or the cross-entropy/hinge examples, but it should be corrected.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies representation learning from paired variables (X,Y), defining 'dependence induced' representations as those invariant to dependence preserving transformations. It proves that the HGR maximal correlation functions are dependence induced and minimal sufficient (Theorem 1, Proposition 2), characterizes dependence learning algorithms (Theorem 2), and introduces a family of 'D-losses' whose minimizers factor through the maximal correlation functions (Theorem 3). The paper then shows that cross entropy and hinge loss (with appropriate calibrations) are D-losses, interprets neural collapse as a consequence of strict dependence, and proposes a feature-adapter framework with the nested H-score for learning the maximal correlation functions.","tokens_in":14992,"tokens_out":6012,"duration_ms":50817,"significance":"The central result, Theorem 3, is a clean, parameter-free characterization: for every regular D-loss, the optimal feature pair is a composition of a loss-dependent map with the HGR functions. This unifies seemingly different losses and connects representation learning to minimal sufficient statistics. The derivation is self-contained for the main theorem, and the cross-entropy and hinge examples are plausible. However, the paper contains a false claim in Section V-B that the negative nested H-score is a regular D-loss; this affects the proposed method for learning f* and g* but does not undermine Theorem 3. The paper also relies on several unproved assertions (Fact 1, Property 3, Proposition 1) that should be proved or cited. The overall contribution is significant and worth publishing after revision, but the incorrect nested H-score claim must be corrected.","major_comments":[{"comment":"The statement 'It can be verified that -H_nest(f,g) is a regular D-loss' is false. For k=1, take X,Y with P(0,0)=P(1,1)=0.4 and P(0,1)=P(1,0)=0.1, f(X)=1-2X, g(Y)=1-2Y, and constant s,t. Then f|s=0 and g|t=0, so H(f|s,g|t)=0, while H(f,g)=0.1; hence -H(f|s,g|t)=0 > -H(f,g)=-0.1, violating the D-loss inequality (9). Since this claim is used to justify learning f*,g* by maximizing the nested H-score, the section needs a corrected construction or a different justification.","section":"Section V-B, Eq. (21)-(22)"},{"comment":"Fact 1, which asserts that E_X[log E_Y[exp(f(X)^T g(Y))]] is a D-loss, is stated without proof. This fact is load-bearing because it places cross entropy in the D-loss family and supports the neural collapse interpretation. The statement is true (it can be proved by two applications of conditional Jensen), but the paper should include the proof or a reference. The same applies to Property 3 in Section IV-A, which is also stated without proof and is used to regularize D-losses.","section":"Section IV-B.1, Fact 1"},{"comment":"Proposition 1, which states that the CDK of transformed variables equals the original CDK at decoded values, is used in the proof of Theorem 1 but its proof is omitted. Since Theorem 1 is foundational for the argument, this lemma should be proved in the paper or given a precise reference.","section":"Section III-A, Proposition 1"}],"minor_comments":[{"comment":"The denominator in the definition of the CDK is written as P_X(y)P_Y(y); it should be P_X(x)P_Y(y).","section":"Section II-B, Eq. (3)"},{"comment":"The notation xi^{-1}, eta^{-1} for the decoding maps is abusive because xi and eta are defined on X x Z and Y x W, respectively; please clarify that these are selected right inverses.","section":"Section III-A, Definition 2"},{"comment":"The condition that gamma is 'convex with respect to both arguments' is ambiguous; please specify whether joint convexity is required for the Jensen argument.","section":"Section IV-A, Property 4"},{"comment":"The footnote states that the weights before each H-score can be arbitrary positive numbers, but Eq. (21) uses unit weights; please reconcile the notation or clarify that this is a special case.","section":"Section V-B"}],"recommendation":"major_revision","confidential_remarks":"The main theorem is sound and the paper has merit, but the false claim about the nested H-score in Section V-B is a genuine error that affects the proposed method for learning maximal correlation functions. The many 'proof is omitted' statements, especially Fact 1, should also be addressed. I recommend major revision rather than rejection because the central contribution is independent of the flawed nested H-score claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on Xu & Zheng's \"Dependence Induced Representations.\" The central result is real: Theorem 3 shows any regular D-loss has minima that factor through the HGR maximal correlation functions, and the optimal value is dependence induced. That provides a common formal explanation for why cross entropy, hinge, and divergence-based objectives end up with structurally similar representations. The composition theorem is new and clean, and its proof is supplied. I think the reader's main worry about Fact 1 is misplaced: the bound Γ(f|s,g|t) ≤ Γ(f,g) for the log-sum-exp functional does hold by Jensen, so cross entropy is indeed in the D-loss family. Property 3 also holds, via strict convexity of the norm terms.\n\nBut the paper has a real error in Section V-B: it claims -Hnest is a regular D-loss. It is not. With P(0,0)=P(1,1)=0.4, P(0,1)=P(1,0)=0.1, take f(X)=1-2X, g(Y)=1-2Y, and constant s,t. Then H(f|s,g|t)=0 while H(f,g)=0.1, so -H(f|s,g|t) > -H(f,g), violating the D-loss inequality. This is peripheral; it does not affect Theorem 3 or the cross-entropy/hinge examples, but it should be corrected. The bigger weakness is the number of results stated without proof: Properties 1-3, Proposition 1, Corollaries 2-3, Proposition 4, and Fact 1 are all asserted with \"proof omitted\" or \"can be verified.\" For a theory paper that is too much. Proposition 4 carries the neural collapse interpretation, and its proof is missing. The adapter proposal is untested, which is fine for a theory preprint, but it is speculative. The citation pattern looks fine; leaning on the authors' own prior work on CDK and H-score is legitimate, since those are formal derivations reused as tools.\n\nWho should read this: anyone working on representation learning theory or on explaining neural collapse. It deserves a serious referee, but the referee should demand full proofs for the unproved statements and a correction to the -Hnest claim.","headline":"Solid central theorem on dependence-induced representations, but one false claim about -Hnest and too many omitted proofs; worth refereeing with revision.","tokens_in":15422,"tokens_out":2900,"would_cite":true,"duration_ms":27416,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62B10","94A17","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"For any regular D-loss, optimal features are composed from the HGR maximal correlation functions and a loss-dependent adapter.","keywords":["dependence induced representations","HGR maximal correlation","D-loss","minimal sufficient statistics","representation learning","neural collapse","cross entropy","feature adapters"],"falsifier":"Take a small finite joint distribution, such as binary $X$ and binary $Y$, choose features $f,g$ and nontrivial coarsenings $s(X),t(Y)$, and numerically compare $\\Gamma(f,g)=E_X\\log E_Y\\exp(f(X)^T g(Y))$ with $\\Gamma(f|s,g|t)$; if the latter is larger for any choice, Fact 1 is false, and the cross-entropy and neural-collapse conclusions built on it would not follow.","tokens_in":14504,"feed_emoji":"🔗","tokens_out":9129,"duration_ms":80775,"temperature":0.7,"pith_summary":"The paper aims to show that a broad family of loss functions, called D-losses, all produce optimal feature representations with one shared structure: the minimizer can always be written as $(f,g)=(\\varphi\\circ f^*, \\psi\\circ g^*)$, where $f^*,g^*$ are the Hirschfeld–Gebelein–Rényi (HGR) maximal correlation functions of the pair $(X,Y)$ and $\\varphi,\\psi$ are loss-dependent adapters. If this is true, representations learned by cross entropy, hinge loss, their regularized variants, and variational divergence objectives are not arbitrary: they are functions of the same dependence-induced backbone and differ only through small adapters. The paper also derives necessary and sufficient conditions for a representation to be dependence induced, connects these to minimal sufficient statistics, and presents a feature-adapter design that lets hyperparameters be tuned during inference. A reader should care because the result offers a common theoretical explanation for why very different losses often learn similar features, and it gives a statistical reading of neural collapse in classifiers.","feed_headline":"Many losses provably share one feature backbone","feed_subtitle":"Cross entropy, hinge, and regularized variants all learn the same dependence features; neural collapse follows.","key_machinery":"The carrying object is the D-loss axiomatization together with the modal decomposition of the canonical dependence kernel. The canonical dependence kernel is $i_{X;Y}(x,y)=P_{X,Y}(x,y)/(P_X(x)P_Y(y))-1$, and its singular value decomposition in function space yields singular functions $(f_i^*, g_i^*)$ with singular values $\\sigma_i$; these are the HGR maximal correlation functions. A functional $\\Gamma(f,g)$ is a D-loss when it is invariant under composing features with dependence-preserving transformations and satisfies $\\Gamma(f|s,g|t)\\le \\Gamma(f,g)$ for every Markov chain $X-s(X)-t(Y)-Y$, where $f|s$ is the conditional expectation of $f$ given $s(X)$. These two axioms force the data-processing direction that lets a minimizer be replaced by its conditional expectation onto the minimal sufficient statistics $f^*(X),g^*(Y)$ without increasing the loss, which is the step that produces the composition structure.","core_discovery":"On the paper's own terms, the central discovery is Theorem 3: for any regular D-loss $L$, the minimum value $v_k(X,Y)$ is dependence induced, and there exist dependence-induced mappings $\\varphi,\\psi$ such that some optimal feature pair is exactly $(\\varphi\\circ f^*, \\psi\\circ g^*)$. Consequently the map from data to the argmin of a regular D-loss is a dependence learning algorithm: the learned features depend only on the $X$-$Y$ dependence structure, not on marginal noise. The paper characterizes D-losses axiomatically by invariance to dependence-preserving transformations and a data-processing inequality, and then verifies that cross entropy (after bias calibration), hinge loss (for balanced binary labels), regularized variants, and variational forms of $\\phi$-divergences belong to the family. For strictly dependent pairs, including classification data with deterministic labels, this yields the neural-collapse phenomenon $f(X)=\\varphi(y(X))$ as a consequence of dependence structure rather than a special property of the softmax loss.","pith_inferences":["If Fact 1 is verified, then a practical prediction follows: softmax classifiers trained with weight decay on noisy labels are, at their optima, dependence learning algorithms, and the calibrated bias $\\beta(Y)$, not the raw bias, is the quantity that should match across runs.","The adapter separation suggests a testable extension: for image or text data, training only small adapters on features from a model that approximates $f^*$ should recover most of a full-network retrain for any D-loss; this could be checked on existing benchmarks.","The axiomatic D-loss definition is not restricted to the losses named in the paper; contrastive and adversarial objectives that satisfy the same data-processing inequality would inherit the composition theorem, though the paper does not analyze them.","Because all D-loss optima factor through the same minimal sufficient statistics, the theory predicts that representations learned from different D-losses are informationally equivalent after calibration; mutual information between feature pairs across losses should be equal, up to estimation error."],"forward_implications":["For every regular D-loss, the optimal features are abstractions of the HGR maximal correlation functions, which are minimal sufficient statistics; so the learned representation discards all non-dependence information.","Cross entropy and hinge loss, with the stated calibrations and regularization, are included in the family; their optimal features share the same $f^*,g^*$ and differ only in $\\varphi,\\psi$.","On any training set where the label is a deterministic function of the input, every regular D-loss collapses all examples with the same label to one representation, which reproduces the NC1 component of neural collapse.","The composition structure permits feature adapters: train $f^*,g^*$ once, then optimize $\\varphi,\\psi$ for a particular loss, and even index $\\varphi,\\psi$ by a hyperparameter so that regularization can be tuned at inference time.","Convex feature constraints and weight-decay-style regularization preserve the D-loss property, so constrained and regularized variants remain within the same framework."],"supporting_citations":[{"why":"supplies the modal decomposition in function space and the nested H-score used to learn the maximal correlation functions $f^*,g^*$.","marker":"[3]"},{"why":"defines the canonical dependence kernel and records that the HGR functions are sufficient statistics.","marker":"[8]"},{"why":"introduces the Hirschfeld correlation coefficient that grounds the HGR maximal correlation functions.","marker":"[9]"},{"why":"develops the eigenproblem formulation of maximal correlation used in the modal decomposition.","marker":"[10]"},{"why":"establishes the measure-of-dependence framework that makes HGR functions canonical.","marker":"[11]"},{"why":"provides the definitions of sufficiency and minimal sufficiency on which Proposition 2 and Corollary 2 rest.","marker":"[12]"},{"why":"documents the neural collapse phenomenon that the paper reinterprets as strict-dependence structure.","marker":"[5]"},{"why":"supplies the hinge loss whose calibrated feature form is shown to be a D-loss.","marker":"[15]"},{"why":"gives the variational forms of $\\phi$-divergences used as another D-loss example.","marker":"[16]"}],"fun_headline_variants":["Dependence alone fixes your feature representation","All D-losses share one feature backbone","Neural collapse is dependence, not softmax","Cross entropy and hinge learn same dependence features","Your loss shapes, but dependence decides features"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is the unproved Fact 1 in Section IV-B.1: the log-sum-exp functional $E_X\\log E_Y\\exp(f(X)^T g(Y))$ satisfies the D-loss data-processing inequality $\\Gamma(f|s,g|t)\\le \\Gamma(f,g)$ for every Markov chain $X-s(X)-t(Y)-Y$; if that inequality fails, cross entropy is not a D-loss and the paper's headline classifier example loses its support.","fun_headline_variants_meta":{"raw":{"variants":["Dependence alone fixes your feature representation","All D-losses share one feature backbone","Neural collapse is dependence, not softmax","Cross entropy and hinge learn same dependence features","Your loss shapes, but dependence decides features"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000437,"raw_usage":{"total_tokens":2195,"prompt_tokens":890,"completion_tokens":1305,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":1239}},"tokens_in":506,"tokens_out":1305,"duration_ms":11065,"temperature":1.0,"reasoning_tokens":1239,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:28:40.807363+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a small finite joint distribution, such as binary $X$ and binary $Y$, choose features $f,g$ and nontrivial coarsenings $s(X),t(Y)$, and numerically compare $\\Gamma(f,g)=E_X\\log E_Y\\exp(f(X)^T g(Y))$ with $\\Gamma(f|s,g|t)$; if the latter is larger for any choice, Fact 1 is false, and the cross-entropy and neural-collapse conclusions built on it would not follow.","supporting_citations":[{"cited_title":"A connection between correlation and contingency,","cited_arxiv_id":null,"evidence_quote":"introduces the Hirschfeld correlation coefficient that grounds the HGR maximal correlation functions."},{"cited_title":"Das statistische problem der korrelatio n als variations- und eigenwertproblem und sein zusammenhang mit der ausglei ch- srechnung,","cited_arxiv_id":null,"evidence_quote":"develops the eigenproblem formulation of maximal correlation used in the modal decomposition."},{"cited_title":"On measures of dependence,","cited_arxiv_id":null,"evidence_quote":"establishes the measure-of-dependence framework that makes HGR functions canonical."},{"cited_title":"Elements of information th eory (wiley series in telecommunications and signal processing),","cited_arxiv_id":null,"evidence_quote":"provides the definitions of sufficiency and minimal sufficiency on which Proposition 2 and Corollary 2 rest."},{"cited_title":"Support-vector networks,","cited_arxiv_id":null,"evidence_quote":"supplies the hinge loss whose calibrated feature form is shown to be a D-loss."},{"cited_title":"Estimati ng diver- gence functionals and the likelihood ratio by convex risk mi nimiza- tion,","cited_arxiv_id":null,"evidence_quote":"gives the variational forms of $\\phi$-divergences used as another D-loss example."}],"review_version":1}