{"id":"bc6dc9b6-5ede-4ac1-af38-2272772c0b83","arxiv_id":"2507.14986","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper reviews i.i.d. identifiability in unlinked linear regression and proves new identifiability bounds and sufficient conditions for non-i.i.d. covariates, connecting the problem to Independent Component Analysis.","lead":"Unlinked Linear Regression deals with two data sets, predictors and outcomes, that are not linked to each other. This paper proves when the regression coefficients can still be uniquely recovered from the two distributions, giving new identifiability results and explicit failures in general settings.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Convolution invariance claim in §4.3 is false: replacing X_j by an equal-in-distribution linear combination changes the distribution unless the other coefficients vanish.","rationale":"The reader's weakest_assumption concerns the subset-sum uniqueness condition in Theorem 5. That condition is restrictive and ad hoc, but it is an explicit assumption and the proof of Theorem 5 appears valid. My review identifies a different, more concrete correctness issue: the convolution invariance construction in §4.3 is stated without the necessary condition that β0i = 0 for all i ∈ I. The included example with X3 d= X1+X2 and β0 = (0,0,1) is correct, but the general claim that B0 always contains β0+δ is false, as shown by the explicit counterexample β0 = (1,0,1) and β = (2,1,0) with Exp/Gamma variables. This is a genuine mathematical error in a supporting section. However, it does not invalidate the main positive results (Theorems 5 and 6), and the negative message about the absence of general identifiability theorems retains support from the corrected version of the convolution example and from spherical symmetry (Theorem 4). Since the error is localized and fixable, the appropriate verdict remains conditional acceptance pending revision, matching the reader's original CONDITIONAL verdict. I therefore mark UNCHANGED, although for a different reason than the reader's abstract-wording concern.","tokens_in":19331,"tokens_out":28462,"duration_ms":265846,"concrete_test":"Compute the moment-generating functions for the counterexample in §4.3: X1, X2 ∼ Exp(1), X3 ∼ Gamma(2,1), independent. Verify that with β0 = (1,0,1) and β = (2,1,0), we have M_{X1+X3}(t) = (1−t)^−3, while M_{2X1+X2}(t) = (1−2t)^−1(1−t)^−1. If the two are not equal, the claimed B0 inclusion in §4.3 is refuted in this instance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In §4.3, the paper claims that if X_j d= Σ_{i∈I} α_i X_i for I ⊆ {1,…,d}\\{j}, then for any β0 the vector β0+δ with δ_j = −β0j, δ_i = α_i β0j for i∈I is in B0. This is not true in general. The error is that after substituting X_j with Σ α_i X_i, the added linear combination is not independent of the remaining part R = Σ_{k≠j} β0k X_k when R contains any X_i with i∈I. Equality in distribution requires more than the marginal replacement; the joint dependence matters.\n\nA concrete counterexample: take X1, X2 ∼ Exp(1), X3 ∼ Gamma(2,1), independent, so X3 d= X1+X2. Let β0 = (1,0,1) and β = (2,1,0), which is β0+δ with j=3, α1=α2=1. Then β0^T X = X1+X3 ∼ Gamma(3,1), whose MGF is (1−t)^−3. But β^T X = 2X1+X2 has MGF (1−2t)^−1(1−t)^−1, which is not equal (e.g., t^2 coefficient 7 vs 6). Hence β ∉ B0. The claim in §4.3 is therefore false as stated. It only holds when β0i = 0 for all i ∈ I, so that R is independent of both X_j and Σ α_i X_i. The specific illustrative example with β0 = (0,0,1) is valid, but the general construction is not.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies identifiability of the regression coefficient vector in unlinked linear regression, where Y is distributed as β⊤X + ε but the data on X and Y are not paired. It reviews classical results for i.i.d. components (Marcinkiewicz, Linnik, Ghurye-Olkin), then turns to the non-i.i.d. case. The paper gives counterexamples showing that general analogues of the i.i.d. results fail (spherical symmetry, convolution invariance), proves a positive strong-identifiability result for independent Gaussian-plus-Gamma covariates under a subset-sum uniqueness condition (Theorem 5), and derives a fourth-moment condition for two-dimensional covariates that bounds the identifiability set (Theorem 6). It also discusses connections to independent component analysis and states an informal conjecture.","tokens_in":19635,"tokens_out":15906,"duration_ms":171385,"significance":"If the positive results hold, the paper provides a useful map of identifiability phenomena in unlinked linear regression and extends earlier work of Azadkia and Balabdaoui. The fourth-moment result for d=2 is elegant and checkable from the first four moments of the covariates, and the review of the i.i.d. case is well organized. The ICA connection is illuminating. However, the paper's negative message is weakened by a false convolution-invariance claim with a concrete counterexample, and the abstract overstates the scope of the negative findings by claiming impossibility rather than presenting counterexamples.","major_comments":[{"comment":"The abstract states that when the covariate components have different distributions, 'we show that it is not possible to prove similar theorems in the general case.' This is a meta-mathematical impossibility claim, but the paper only provides counterexamples (spherical symmetry, convolutions) showing that certain natural generalizations fail. Counterexamples do not establish that no theorem of a similar kind can be proved. The abstract, and the parallel sentence in §1, should be rephrased to say that the paper demonstrates obstacles to direct generalization by exhibiting counterexamples, rather than claiming a proof of impossibility.","section":"Abstract"},{"comment":"The convolution-invariance claim is false as stated. If Xj is equal in distribution to Σ_{i∈I} α_i X_i, the paper asserts that B0 contains β0+δ with δj = -β0j and δi = αiβ0j for i∈I. This does not hold in general: after the substitution, the remainder R = Σ_{k≠j} β0k X_k is generally not independent of the replacement combination when R contains any Xi with i∈I. For a concrete counterexample, take X1,X2∼Exp(1) independent and X3∼Gamma(2,1) independent, so X3 d= X1+X2. Let β0=(1,0,1). The construction yields β=(2,1,0). But β0^T X = X1+X3 has moment-generating function (1−t)^{-3}, whereas β^T X = 2X1+X2 has MGF ((1−2t)(1−t))^{-1}; the t^2 coefficients are 6 and 7, respectively, so these distributions are not equal. The invariance is valid only when β0i=0 for all i∈I, so that the remainder is independent of both Xj and the combination. The specific illustrative example with β0=(0,0,1) works, but the general construction and the surrounding discussion must be corrected.","section":"§4.3"},{"comment":"The paragraph 'Extension to more than two random variables' asserts, without proof, that applying Theorem 6 to subsets yields an upper bound of 3!·2^3 = 48 solutions for d=3 and, recursively, d!·2^d for general d. This is not established: Theorem 6 concerns the identifiability set for a fixed two-dimensional response, and the reasoning does not show that a candidate solution in higher dimension must be controlled by the two-dimensional marginal conditions in a way that yields the claimed global bound. The moment equations couple all coefficients through the full projection β^T X. Either a rigorous proof should be supplied, or the paragraph should be explicitly labeled as a heuristic sketch or open problem.","section":"§4.5"}],"minor_comments":[{"comment":"In the definition of B0 within Theorem 6, 'aX1 + bX1' should read 'aX1 + bX2'; the subsequent proof uses the correct form.","section":"Theorem 6 statement"},{"comment":"The theorem states that there exists ρ>0 with B0 = Sρ, but if the true regression vector β0 equals zero, the sphere has radius zero. The statement should allow ρ≥0 or explicitly exclude the degenerate case β0=0, which is consistent with the proof (ρ = ||β0||).","section":"Theorem 4"},{"comment":"In the appendix proof of Lemma 1, the displayed equivalence is missing a factor of Φε on the right-hand side inside the statement; it should read Φ_{β⊤0X} Φε = Φ_{β⊤1X} Φε, so that cancellation of Φε is justified.","section":"Lemma 1 proof"},{"comment":"The pole-matching argument in the proof of Theorem 5 does not explicitly handle the case where one side has no positive coefficients (c+ = 0) or no negative coefficients (c− = 0), even though such cases can occur. The contradiction can be obtained by taking the limit at the smallest pole of the other side, but this case should be stated, since the current text assumes that both R+(1) and R−(1) exist.","section":"Theorem 5 proof"},{"comment":"The note at the end of Theorem 3 that the assertion 'continues to hold if X is replaced by M X' for an invertible matrix M is unclear: if M is not diagonal, the components of M X are generally not independent, so the proof does not apply. Please clarify the intended statement or remove the note.","section":"Theorem 3 note"},{"comment":"The footnote contains the placeholder 'Supported by XXX'; this should be filled in before publication.","section":"Footnote"}],"recommendation":"major_revision","confidential_remarks":"The paper is a mix of review, counterexamples, and partial positive results. The core positive theorems (Theorems 3, 4, 5, 6) appear sound, though Theorem 5 has a small gap in a corner case and the extension in §4.5 is not rigorous. The false convolution-invariance claim in §4.3 is a substantive error, but it is localized and fixable by adding the independence condition or by demoting the claim to an example. The abstract's 'impossibility' wording should be corrected. I would be willing to see a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Main take: this is a useful identifiability note, but it is not as clean as it presents itself. The abstract says impossibility is shown \"in the general case\" when the components are not i.i.d.; what the paper actually offers is a set of counterexamples, not a general impossibility theorem. Better to read the negative message as evidence of obstacles, and the wording should say that.\n\nWhat is genuinely new: Theorem 3 for scale families, Theorem 4 and Corollary 1 for spherical/elliptical symmetry, Theorem 5 for Gamma-plus-normal coordinates, and Theorem 6 for fourth-moment identifiability in d=2. The MGF pole-matching proof of Theorem 5 is rigorous, and the subset-sum uniqueness assumption is restrictive but clearly stated. Theorem 6's condition is simple and checkable, and the three examples are genuinely useful. The review of the i.i.d. case via Marcinkiewicz and Linnik is accurate, and the ICA connection is well placed.\n\nThe soft spots are real and one is a genuine error. In Section 4.3 the paper claims that if X_j is equal in distribution to a linear combination of other coordinates, then certain shifts of beta0 are still in B0. That is false unless beta0 vanishes on the coordinates used in the combination. Replacing X_j by an equal-in-distribution linear combination changes the joint dependence with the rest. Concrete counterexample: X1,X2 ~ Exp(1), X3 ~ Gamma(2,1), independent, with beta0=(1,0,1). The claimed beta=(2,1,0) gives 2X1+X2 with MGF (1-2t)^-1(1-t)^-1, not X1+X3's (1-t)^-3. The special case beta0=(0,0,1) works, but the general statement needs an extra condition. This is a local flaw, not a load-bearing one for the main positive theorems, but it should be corrected. Also, Theorem 6 has a typo in the display defining B0: \"bX1\" should be \"bX2.\"\n\nWho gets value: people working on unlinked data, identifiability, and ICA. The paper deserves refereeing, but a referee should ask for a revised abstract and a corrected Section 4.3.","headline":"Solid identifiability note with several new results, but it overclaims impossibility in the abstract and contains a false general claim in the convolutions section that needs fixing.","tokens_in":20208,"tokens_out":3472,"would_cite":true,"duration_ms":38966,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62E10","62J05"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that unlinked linear regression is strongly identifiable for a Gamma-plus-nonzero-mean-normal covariate vector, and shows that outside such parametric cases general identifiability fails.","keywords":["unlinked linear regression","identifiability","data integration","independent component analysis","characterization problems","Gamma-normal mixtures","fourth moments","model identifiability"],"falsifier":"A direct check for $d=3$ with $\\alpha_1=1$, $\\alpha_2=\\sqrt{2}$, $\\lambda_1=\\lambda_2=1$, and $X_3 \\sim N(1,1)$: compare the moment-generating functions of $\\beta^\\top X$ and $\\tilde{\\beta}^\\top X$ over a fine grid of $t$; since the subset-sum condition holds, the theorem predicts $\\beta=\\tilde{\\beta}$, so any distinct pair with matching moment-generating functions would refute Theorem 5.","tokens_in":19073,"feed_emoji":"📊","tokens_out":11528,"duration_ms":119126,"temperature":0.7,"pith_summary":"Unlinked linear regression asks whether the coefficient vector $\\beta_0$ can be recovered when the response $Y$ and the covariates $X$ are observed in separate files, with no record-level link, and $Y$ is only assumed equal in distribution to $\\beta_0^\\top X + \\epsilon$. The paper establishes that in the independent-but-not-identically-distributed case, no general identifiability theorem is possible: spherical symmetry makes every vector on a sphere a valid coefficient, and convolution relations among covariates create shifts. It then proves positive results: strong identifiability ($B_0 = \\{\\beta_0\\}$) when one covariate is normal with nonzero mean and the others are Gamma with shape parameters whose subset sums are unique, and for $d=2$ a fourth-moment condition that bounds $B_0$ to at most eight vectors or, often, to sign flips alone. These results tell practitioners when unlinked-data regression is trustworthy and when it is inherently underdetermined.","feed_headline":"One normal plus Gammas makes unlinked regression fully identifiable","feed_subtitle":"The paper also shows why general unlinked regression fails: symmetries and convolutions hide the true coefficient.","key_machinery":"The set $B_0$ defined above is the object that carries the argument; levels of identifiability are literally the cardinality of $B_0$ being infinite, finite and larger than one, or a singleton. For the Gamma-normal theorem the load-bearing identity is the moment-generating-function equality $\\prod_i M_{X_i}(\\beta_i t) = \\prod_i M_{X_i}(\\tilde{\\beta}_i t)$, which after dividing out the normal factor becomes a product of terms $(1 - t/R_i^{\\pm})^{-\\alpha_i}$; matching the smallest pole $R_{(1)}^+$ and then successively higher poles forces the coefficient sets and shape-exponent sums to agree, and the subset-sum uniqueness assumption converts exponent-sum equality into index-set equality. In the i.i.d. review, the classical Marcinkiewicz and Linnik theorems on identically distributed linear forms play the same role, and the $d=2$ theorem reduces identifiability to a quadratic equation in $a^4/r^4$ derived from second and fourth moments.","core_discovery":"On the paper's own terms, the central object is the set $B_0 = \\{\\beta \\in \\mathbb{R}^d : \\beta^\\top X + \\epsilon \\stackrel{d}{=} \\beta_0^\\top X + \\epsilon\\}$, and identifiability is its cardinality. The main positive theorem (Theorem 5) says: if $X_d \\sim N(\\mu,\\sigma^2)$ with $\\mu \\neq 0$ and $X_i \\sim \\mathrm{Gamma}(\\alpha_i,\\lambda_i)$ for $i<d$, all independent, and if equal sums of the shape parameters over subsets of $\\{1,\\dots,d-1\\}$ force the subsets to be equal, then $B_0 = \\{\\beta_0\\}$, so the regression vector is strongly identifiable. The proof matches poles of the moment-generating functions at the ratios $\\lambda_i/|\\beta_i|$, using the subset-sum condition to conclude coefficient-by-coefficient equality. Equally central is the negative message that outside such parametric settings identifiability fails in structured ways: spherical or elliptical symmetry yields the whole sphere or an ellipsoid section as $B_0$, and a covariate that is a convolution of others shifts the apparent coefficient. For $d=2$, the fourth-moment theorem shows that unless both centered standardized covariates have fourth moment $3$, $B_0$ is finite and often consists only of sign flips.","pith_inferences":["Because randomly drawn continuous positive shape parameters have no equal subset sums with probability one, Theorem 5 covers 'almost every' Gamma-plus-nonzero-normal design, so its restrictive-looking condition is generic rather than exceptional.","The $d=2$ fourth-moment criterion could be converted into a finite-sample pre-test: estimate $m_1$, $m_2$, and $c$ from the observed data, compute $w_1$ and $w_2$, and declare identifiability up to sign when $\\min\\{w_1,w_2\\}<1$; the paper proves only the population version.","A natural next step is to test the paper's closing conjecture — weak identifiability for independent unit-variance components with at most one Gaussian and a minimal representation — by searching for counterexamples among discrete or mixture distributions, where the moment-generating-function and moment arguments used here do not directly apply."],"forward_implications":["Under the Theorem 5 conditions, an unlinked-data regression estimate is uniquely anchored: any consistent estimator of $\\beta_0$ is estimating the true coefficient, not an equivalent one.","For independent covariates with different distributions, unlinked regression is identifiable only when the design prevents symmetries: if $X$ is spherically or elliptically symmetric, every $\\beta$ on the corresponding sphere or ellipse is indistinguishable, so no method can single out $\\beta_0$.","If any covariate is a convolution of other covariates (for example, $X_3 \\stackrel{d}{=} X_1 + X_2$ for Gammas), the coefficient vector is identified only up to shifts, so one must either exclude such redundancy or accept set-valued inference.","In the $d=2$ case, computing the fourth moments of the centered standardized covariates gives a practical criterion: if not both kurtoses equal $3$, the solution set is finite and usually just sign flips, so sign is the only ambiguity.","The multi-response variant with $m \\ge 2$ response variables inherits a general weak identifiability result from overcomplete independent component analysis: with non-Gaussian independent sources and pairwise linearly independent columns, $B_0$ consists only of permutations and sign flips of the true matrix."],"supporting_citations":[{"why":"Defines unlinked linear regression, introduces the deconvolution least-squares estimator, and supplies the normal-plus-exponential example that Theorem 5 generalizes.","marker":"[2]"},{"why":"Marcinkiewicz's theorem on linear forms of i.i.d. variables is the backbone of the scale-family identifiability result (Theorem 3).","marker":"[26]"},{"why":"Linnik's theorem on identical distributions of linear forms, stated here as Theorem 2, gives the coefficient-level conditions for weak identifiability in the i.i.d. case.","marker":"[24]"},{"why":"The Ghurye-Olkin survey provides the statement and sufficient conditions for Linnik's theorem used in the i.i.d. review.","marker":"[16]"},{"why":"The representation of elliptically symmetric distributions as affine maps of spherically symmetric ones is used to extend non-identifiability to the elliptical case.","marker":"[5]"},{"why":"Theorem 10.3.5 in Kagan-Linnik-Rao is quoted as Theorem 7 and supplies the general identifiability result for multi-response ULR via overcomplete independent component analysis.","marker":"[20]"},{"why":"Establishes identifiability of overcomplete ICA models, connecting the multi-response ULR variant to that literature.","marker":"[13]"}],"fun_headline_variants":["Normal+Gamma: unlinked regression identifiable","Unlinked regression: only special distributions give identifiability","General unlinked regression: not identifiable, but special cases work","Symmetry hides coefficients in unlinked regression","Identifiability in unlinked regression: a normal plus gammas suffices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise of the positive theorem is the subset-sum uniqueness condition on the Gamma shape parameters: matching pole orders in the moment-generating-function proof concludes $\\beta_i = \\tilde{\\beta}_i$ only because equal sums of $\\alpha$'s over subsets are assumed to force equal index sets, and if two different subsets sum to the same value the argument degrades to identification up to permutation.","fun_headline_variants_meta":{"raw":{"variants":["Normal+Gamma: unlinked regression identifiable","Unlinked regression: only special distributions give identifiability","General unlinked regression: not identifiable, but special cases work","Symmetry hides coefficients in unlinked regression","Identifiability in unlinked regression: a normal plus gammas suffices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000995,"raw_usage":{"total_tokens":4244,"prompt_tokens":1002,"completion_tokens":3242,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":3161}},"tokens_in":618,"tokens_out":3242,"duration_ms":24369,"temperature":1.0,"reasoning_tokens":3161,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:44:02.404698+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct check for $d=3$ with $\\alpha_1=1$, $\\alpha_2=\\sqrt{2}$, $\\lambda_1=\\lambda_2=1$, and $X_3 \\sim N(1,1)$: compare the moment-generating functions of $\\beta^\\top X$ and $\\tilde{\\beta}^\\top X$ over a fine grid of $t$; since the subset-sum condition holds, the theorem predicts $\\beta=\\tilde{\\beta}$, so any distinct pair with matching moment-generating functions would refute Theorem 5.","supporting_citations":[{"cited_title":"Azadkia and F","cited_arxiv_id":null,"evidence_quote":"Defines unlinked linear regression, introduces the deconvolution least-squares estimator, and supplies the normal-plus-exponential example that Theorem 5 generalizes."},{"cited_title":"Marcinkiewicz","cited_arxiv_id":null,"evidence_quote":"Marcinkiewicz's theorem on linear forms of i.i.d. variables is the backbone of the scale-family identifiability result (Theorem 3)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Linnik's theorem on identical distributions of linear forms, stated here as Theorem 2, gives the coefficient-level conditions for weak identifiability in the i.i.d. case."},{"cited_title":"Ghurye and I","cited_arxiv_id":null,"evidence_quote":"The Ghurye-Olkin survey provides the statement and sufficient conditions for Linnik's theorem used in the i.i.d. review."},{"cited_title":"Cambanis, S","cited_arxiv_id":null,"evidence_quote":"The representation of elliptically symmetric distributions as affine maps of spherically symmetric ones is used to extend non-identifiability to the elliptical case."},{"cited_title":"Kagan, Y","cited_arxiv_id":null,"evidence_quote":"Theorem 10.3.5 in Kagan-Linnik-Rao is quoted as Theorem 7 and supplies the general identifiability result for multi-response ULR via overcomplete independent component analysis."},{"cited_title":"Eriksson and V","cited_arxiv_id":null,"evidence_quote":"Establishes identifiability of overcomplete ICA models, connecting the multi-response ULR variant to that literature."}],"review_version":1}