REVIEW 3 major objections 6 minor 45 references
Identifiability in Unlinked Linear Regression: Some Results and Open Problems
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proves that unlinked linear regression is strongly identifiable for a Gamma-plus-nonzero-mean-normal covariate vector, and shows that outside such parametric cases general identifiability fails.
desk verdict Solid identifiability note with several new results, but it overclaims impossibility in the abstract and contains a false general claim in the convolutions section that needs fixing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The set $B_0$ defined above is the object that carries the argument; levels of identifiability are literally the cardinality of $B_0$ being infinite, finite and larger than one, or a singleton. For the Gamma-normal theorem the load-bearing identity is the moment-generating-function equality $\prod_i M_{X_i}(\beta_i t) = \prod_i M_{X_i}(\tilde{\beta}_i t)$, which after dividing out the normal factor becomes a product of terms $(1 - t/R_i^{\pm})^{-\alpha_i}$; matching the smallest pole $R_{(1)}^+$ and then successively higher poles forces the coefficient sets and shape-exponent sums to agree, and the subset-sum uniqueness assumption converts exponent-sum equality into index-set equality. In the i.i.d. review, the classical Marcinkiewicz and Linnik theorems on identically distributed linear forms play the same role, and the $d=2$ theorem reduces identifiability to a quadratic equation in $a^4/r^4$ derived from second and fourth moments.
What would settle it
A direct check for $d=3$ with $\alpha_1=1$, $\alpha_2=\sqrt{2}$, $\lambda_1=\lambda_2=1$, and $X_3 \sim N(1,1)$: compare the moment-generating functions of $\beta^\top X$ and $\tilde{\beta}^\top X$ over a fine grid of $t$; since the subset-sum condition holds, the theorem predicts $\beta=\tilde{\beta}$, so any distinct pair with matching moment-generating functions would refute Theorem 5.
Extended reading notes
Core claim
On the paper's own terms, the central object is the set $B_0 = \{\beta \in \mathbb{R}^d : \beta^\top X + \epsilon \stackrel{d}{=} \beta_0^\top X + \epsilon\}$, and identifiability is its cardinality. The main positive theorem (Theorem 5) says: if $X_d \sim N(\mu,\sigma^2)$ with $\mu \neq 0$ and $X_i \sim \mathrm{Gamma}(\alpha_i,\lambda_i)$ for $i<d$, all independent, and if equal sums of the shape parameters over subsets of $\{1,\dots,d-1\}$ force the subsets to be equal, then $B_0 = \{\beta_0\}$, so the regression vector is strongly identifiable. The proof matches poles of the moment-generating functions at the ratios $\lambda_i/|\beta_i|$, using the subset-sum condition to conclude coefficient-by-coefficient equality. Equally central is the negative message that outside such parametric settings identifiability fails in structured ways: spherical or elliptical symmetry yields the whole sphere or an ellipsoid section as $B_0$, and a covariate that is a convolution of others shifts the apparent coefficient. For $d=2$, the fourth-moment theorem shows that unless both centered standardized covariates have fourth moment $3$, $B_0$ is finite and often consists only of sign flips.
Load-bearing premise
The load-bearing premise of the positive theorem is the subset-sum uniqueness condition on the Gamma shape parameters: matching pole orders in the moment-generating-function proof concludes $\beta_i = \tilde{\beta}_i$ only because equal sums of $\alpha$'s over subsets are assumed to force equal index sets, and if two different subsets sum to the same value the argument degrades to identification up to permutation.
Editorial extensions
If this is right
- Under the Theorem 5 conditions, an unlinked-data regression estimate is uniquely anchored: any consistent estimator of $\beta_0$ is estimating the true coefficient, not an equivalent one.
- For independent covariates with different distributions, unlinked regression is identifiable only when the design prevents symmetries: if $X$ is spherically or elliptically symmetric, every $\beta$ on the corresponding sphere or ellipse is indistinguishable, so no method can single out $\beta_0$.
- If any covariate is a convolution of other covariates (for example, $X_3 \stackrel{d}{=} X_1 + X_2$ for Gammas), the coefficient vector is identified only up to shifts, so one must either exclude such redundancy or accept set-valued inference.
- In the $d=2$ case, computing the fourth moments of the centered standardized covariates gives a practical criterion: if not both kurtoses equal $3$, the solution set is finite and usually just sign flips, so sign is the only ambiguity.
- The multi-response variant with $m \ge 2$ response variables inherits a general weak identifiability result from overcomplete independent component analysis: with non-Gaussian independent sources and pairwise linearly independent columns, $B_0$ consists only of permutations and sign flips of the true matrix.
Reading between the lines
- Because randomly drawn continuous positive shape parameters have no equal subset sums with probability one, Theorem 5 covers 'almost every' Gamma-plus-nonzero-normal design, so its restrictive-looking condition is generic rather than exceptional.
- The $d=2$ fourth-moment criterion could be converted into a finite-sample pre-test: estimate $m_1$, $m_2$, and $c$ from the observed data, compute $w_1$ and $w_2$, and declare identifiability up to sign when $\min\{w_1,w_2\}<1$; the paper proves only the population version.
- A natural next step is to test the paper's closing conjecture — weak identifiability for independent unit-variance components with at most one Gaussian and a minimal representation — by searching for counterexamples among discrete or mixture distributions, where the moment-generating-function and moment arguments used here do not directly apply.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies identifiability of the regression coefficient vector in unlinked linear regression, where Y is distributed as β⊤X + ε but the data on X and Y are not paired. It reviews classical results for i.i.d. components (Marcinkiewicz, Linnik, Ghurye-Olkin), then turns to the non-i.i.d. case. The paper gives counterexamples showing that general analogues of the i.i.d. results fail (spherical symmetry, convolution invariance), proves a positive strong-identifiability result for independent Gaussian-plus-Gamma covariates under a subset-sum uniqueness condition (Theorem 5), and derives a fourth-moment condition for two-dimensional covariates that bounds the identifiability set (Theorem 6). It also discusses connections to independent component analysis and states an informal conjecture.
Significance. If the positive results hold, the paper provides a useful map of identifiability phenomena in unlinked linear regression and extends earlier work of Azadkia and Balabdaoui. The fourth-moment result for d=2 is elegant and checkable from the first four moments of the covariates, and the review of the i.i.d. case is well organized. The ICA connection is illuminating. However, the paper's negative message is weakened by a false convolution-invariance claim with a concrete counterexample, and the abstract overstates the scope of the negative findings by claiming impossibility rather than presenting counterexamples.
major comments (3)
- [Abstract] The abstract states that when the covariate components have different distributions, 'we show that it is not possible to prove similar theorems in the general case.' This is a meta-mathematical impossibility claim, but the paper only provides counterexamples (spherical symmetry, convolutions) showing that certain natural generalizations fail. Counterexamples do not establish that no theorem of a similar kind can be proved. The abstract, and the parallel sentence in §1, should be rephrased to say that the paper demonstrates obstacles to direct generalization by exhibiting counterexamples, rather than claiming a proof of impossibility.
- [§4.3] The convolution-invariance claim is false as stated. If Xj is equal in distribution to Σ_{i∈I} α_i X_i, the paper asserts that B0 contains β0+δ with δj = -β0j and δi = αiβ0j for i∈I. This does not hold in general: after the substitution, the remainder R = Σ_{k≠j} β0k X_k is generally not independent of the replacement combination when R contains any Xi with i∈I. For a concrete counterexample, take X1,X2∼Exp(1) independent and X3∼Gamma(2,1) independent, so X3 d= X1+X2. Let β0=(1,0,1). The construction yields β=(2,1,0). But β0^T X = X1+X3 has moment-generating function (1−t)^{-3}, whereas β^T X = 2X1+X2 has MGF ((1−2t)(1−t))^{-1}; the t^2 coefficients are 6 and 7, respectively, so these distributions are not equal. The invariance is valid only when β0i=0 for all i∈I, so that the remainder is independent of both Xj and the combination. The specific illustrative example with β0=(0,0,1) works, but the general construction and the surrounding discussion must be corrected.
- [§4.5] The paragraph 'Extension to more than two random variables' asserts, without proof, that applying Theorem 6 to subsets yields an upper bound of 3!·2^3 = 48 solutions for d=3 and, recursively, d!·2^d for general d. This is not established: Theorem 6 concerns the identifiability set for a fixed two-dimensional response, and the reasoning does not show that a candidate solution in higher dimension must be controlled by the two-dimensional marginal conditions in a way that yields the claimed global bound. The moment equations couple all coefficients through the full projection β^T X. Either a rigorous proof should be supplied, or the paragraph should be explicitly labeled as a heuristic sketch or open problem.
minor comments (6)
- [Theorem 6 statement] In the definition of B0 within Theorem 6, 'aX1 + bX1' should read 'aX1 + bX2'; the subsequent proof uses the correct form.
- [Theorem 4] The theorem states that there exists ρ>0 with B0 = Sρ, but if the true regression vector β0 equals zero, the sphere has radius zero. The statement should allow ρ≥0 or explicitly exclude the degenerate case β0=0, which is consistent with the proof (ρ = ||β0||).
- [Lemma 1 proof] In the appendix proof of Lemma 1, the displayed equivalence is missing a factor of Φε on the right-hand side inside the statement; it should read Φ_{β⊤0X} Φε = Φ_{β⊤1X} Φε, so that cancellation of Φε is justified.
- [Theorem 5 proof] The pole-matching argument in the proof of Theorem 5 does not explicitly handle the case where one side has no positive coefficients (c+ = 0) or no negative coefficients (c− = 0), even though such cases can occur. The contradiction can be obtained by taking the limit at the smallest pole of the other side, but this case should be stated, since the current text assumes that both R+(1) and R−(1) exist.
- [Theorem 3 note] The note at the end of Theorem 3 that the assertion 'continues to hold if X is replaced by M X' for an invertible matrix M is unclear: if M is not diagonal, the components of M X are generally not independent, so the proof does not apply. Please clarify the intended statement or remove the note.
- [Footnote] The footnote contains the placeholder 'Supported by XXX'; this should be filled in before publication.
Circularity Check
No significant circularity: Theorems 3–6 are derived from stated assumptions with complete proofs; self-citations are contextual and never load-bearing.
full rationale
All new results are proved from scratch in the appendix. Theorem 3 follows by rescaling scale-family variables to an i.i.d. sample and applying Marcinkiewicz's theorem; Theorem 4 and Corollary 1 use the defining invariance of spherical/elliptical symmetry; Theorem 5 works directly on moment-generating functions, matching Gamma poles R±(i) on both sides (Eqs. 10–12), and the subset-sum condition on the shape parameters converts matching exponent sums into matching index sets, yielding βi = β̃i for i < d and then βd = β̃d from φ(t) = 1 since μ ≠ 0; Theorem 6 reduces β^T X d= β0^T X to the second- and fourth-moment equalities (13)–(15) and solves the resulting quadratics. None of these steps invokes the theorem being proved, fits a parameter to data, or imports a result from this paper's authors as a premise. The external theorems used (Marcinkiewicz 1939; Linnik 1953; Kagan–Linnik–Rao 1973 via Eriksson–Koivunen; Cambanis–Huang–Simons 1981) are classical published benchmarks and thus genuine independent support. Self-citations to [2], [3], [37], [38] appear only in the introduction, the review of the i.i.d. case, and motivation; the two examples taken from [2] are re-derived by new methods (Theorem 6 Example 1 and Theorem 5), not imported as premises. The review half is explicitly labeled as a review, so restating [2]'s results there is restatement, not circular reasoning. Two non-circularity caveats flagged per the review rule: (i) the §4.3 convolution claim that Xj d= Σ αi Xi forces β0+δ ∈ B0 is false as stated — replacing Xj by the dependent combination changes the joint law whenever some β0i for i ∈ I is nonzero, so the reduction is valid only when the remaining part is independent of the combination; the paper's own example satisfies this, and the flaw is a correctness risk, not a circular step; (ii) the abstract's "not possible to prove similar theorems in the general case" is stronger than the exhibited counterexamples establish. These issues do not raise the circularity score because no step in the derivation chain is equivalent to its own input by definition, by fit, or by self-citation.
Assumptions & free parameters
assumptions (6)
- domain assumption The noise epsilon in the ULR model satisfies the condition of Lemma 1: its characteristic function has no open interval of zeros.
- domain assumption The covariate components X1,...,Xd are independent.
- ad hoc to paper In Theorem 5, the Gamma shape parameters alpha_i satisfy the subset-sum uniqueness condition: equality of sums over subsets I and J implies I=J.
- domain assumption In Theorem 3, the base density f has finite moments of all orders and is not a Gaussian density.
- domain assumption In Theorem 6, X1 and X2 are centered, unit variance, have finite fourth moments, and not both fourth moments equal 3.
- standard math Classical characterization theorems (Marcinkiewicz 1939, Linnik 1953, Kagan-Linnik-Rao 1973) are valid and applicable.
Cite this review
Pith. "Pith review of Identifiability in Unlinked Linear Regression: Some Results and Open Problems." pith.science (2026). https://pith.science/paper/BVHULBW5
@misc{pith2026250714986,
author = {Pith},
title = {Pith review of: Identifiability in Unlinked Linear Regression: Some Results and Open Problems},
year = {2026},
howpublished = {\url{https://pith.science/paper/BVHULBW5}},
note = {Machine review of arXiv:2507.14986}
}
abstract
A tacit assumption in classical linear regression problems is the full knowledge of the existing link between the covariates and responses. In Unlinked Linear Regression (ULR) this link is either partially or completely missing. While the reasons causing such missingness can be different, a common challenge in statistical inference is the potential non-identifiability of the regression parameter. In this note, we review the existing literature on identifiability when the $d \ge 2$ components of the vector of covariates are independent and identically distributed. When these components have different distributions, we show that it is not possible to prove similar theorems in the general case. Nevertheless, we prove some identifiability results, either under additional parametric assumptions for $d \ge 2$ or conditions on the fourth moments in the case $d=2$. Finally, we draw some interesting connections between the ULR and the well established field of Independent Component Analysis (ICA).
Figures
Reference graph
Works this paper leans on
-
[1]
J. Abowd, J. Abramowitz, M. Levenstein, K. McCue, D. Patki, T. Raghunathan, A. Rodgers, M. Shapiro, N. Wasi, and D. Zinsser. Finding Needles in Haystacks: Multiple-Imputation Record Linkage Using Machine Learning. Federal Reserve Bank of Boston Research Department Working Papers No. 22-11, 2021
work page 2021
-
[2]
M. Azadkia and F. Balabdaoui. Linear regression with unmatched data: A deconvo- lution perspective. Journal of Machine Learning Research , 25(197):1–55, 2024
work page 2024
-
[3]
F. Balabdaoui, C.R. Doss, and C. Durot. Unlinked monotone regression. Journal of Machine Learning Research, 22(172):1–60, 2021
work page 2021
-
[4]
O. Binette and R. Steorts. (Almost) all of entity resolution. Science Advances , 8(12):eabi8021, 2022
work page 2022
-
[5]
S. Cambanis, S. Huang, and G. Simons. On the theory of elliptically contoured dis- tributions. Journal of Multivariate Analysis , 11(3):368–385, 1981
work page 1981
-
[6]
A. Carpentier and T. Schl¨ uter. Learning relationships between data obtained indepen- dently. In Proceedings of the 19th International Conference on Artificial Intelligence and Statistics , pages 658–666, 2016
work page 2016
-
[7]
O. Chapelle, B. Sch¨ olkopf, and A. Zien. Semi-Supervised Learning. MIT Press, 2006
work page 2006
- [8]
Show all 45 references
-
[9]
P. Comon. Independent component analysis, a new concept? Signal Processing, 36(3):287–314, 1994
1994
-
[10]
Comon and C
P. Comon and C. Jutten. Handbook of Blind Source Separation: Independent compo- nent analysis and applications . Academic press, 2010
2010
-
[11]
Durot and D
C. Durot and D. Mukherjee. Minimax optimal rates of convergence in the shuffled regression, unlinked regression, and deconvolution under vanishing noise. arXiv:2404.09306, 2024
2024 arXiv
-
[12]
Enamorado, B
T. Enamorado, B. Fifield, and K. Imai. Using a probabilistic model to assist merging of large-scale administrative records. American Political Science Review, 112(2):353– 370, 2018
2018
-
[13]
Eriksson and V
J. Eriksson and V. Koivunen. Identifiability, separability, and uniqueness of linear ICA models. IEEE Signal Processing Letters , 11:601–604, 2004. 12
2004
-
[14]
K.W. Fang, S. Kotz, and K.W. Ng. Symmetric multivariate and related distributions . Chapman and Hall/CRC, 2018
2018
-
[15]
Gelman, D.K
A. Gelman, D.K. Park, S. Ansolabehere, P.N. Price, and L.C. Minnite. Models, assumptions, and model checking in ecological regressions. Journal of the Royal Sta- tistical Society Series A , 164(1):101–118, 2001
2001
-
[16]
Ghurye and I
S. Ghurye and I. Olkin. Identically distributed linear forms and the normal distribu- tion. Advances in Applied Probability, 5(1):138–152, 1973
1973
-
[17]
D. Hsu, K. Shi, and X. Sun. Linear regression without correspondence. In Advances in Neural Information Processing Systems (NIPS) , pages 1531–1540, 2017
2017
-
[18]
Hyv¨ arinen, J
A. Hyv¨ arinen, J. Hurri, and P.O. Hoyer. Independent component analysis . Springer, 2009
2009
-
[19]
Hyv¨ arinen and E
A. Hyv¨ arinen and E. Oja. Independent component analysis: algorithms and applica- tions. Neural Networks, 13(4-5):411–430, 2000
2000
-
[20]
Kagan, Y
A. Kagan, Y. Linnik, and C. Rao. Characterization Problems in Mathematical Statis- tics. Wiley, 1973
1973
-
[21]
Kamat and R
G. Kamat and R. Gutman. Analysis of linked files: A missing data perspective. Statistical Science, forthcoming, 2024
2024
-
[22]
Kim and J
J.K. Kim and J. Shao. Statistical Methods for Handling Incomplete Data . CRC Press, 2nd edition, 2021
2021
-
[23]
G. King. A Solution to the Ecological Inference Problem: Reconstructing Individual Behavior from Aggregate Data . Princeton University Press, 1997
1997
-
[24]
Y. Linnik. Linear forms and statistical criteria. Ukrain. Math. Zurnal , 5:207–243, 1953
1953
-
[25]
Little and D.B
R.J.A. Little and D.B. Rubin. Statistical Analysis with Missing Data . Wiley, 3rd edition, 2020
2020
-
[26]
Marcinkiewicz
J. Marcinkiewicz. Sur une propri´ et´ e de la loi de gauss. Mathematische Zeitschrift , 44(1):612–618, 1939
1939
-
[27]
McCullagh
P. McCullagh. Tensor methods in statistics . Chapman and Hall/CRC, 2018
2018
-
[28]
Meis and E
J. Meis and E. Mammen. Uncoupled isotonic regression with discrete errors. In Ad- vances in Contemporary Statistics and Econometrics , pages 123–135. Springer, 2021
2021
-
[29]
A. Meister. Deconvolution Problems in Nonparametric Statistics . Springer Berlin Heidelberg, 2009
2009
-
[30]
Mesters and P
G. Mesters and P. Zwiernik. Non-independent component analysis. The Annals of Statistics, 52(6):2506–2528, 2024
2024
-
[31]
Narayanan and V
A. Narayanan and V. Shmatikov. Robust de-anonymization of large sparse datasets. In IEEE Symposium on Security and Privacy , pages 111–125, 2008
2008
-
[32]
Newcombe and J.M
H.B. Newcombe and J.M. Kennedy. Record linkage: making maximum use of the dis- criminating power of identifying information.Communications of the ACM, 5(11):563– 566, 1962. 13
1962
-
[33]
G. Paass. Disclosure risk and disclosure avoidance for microdata. Journal of Business & Economic Statistics , 6(4):487–500, 1988
1988
-
[34]
Pananjady, M.J
A. Pananjady, M.J. Wainwright, and T.A. Courtade. Linear regression with shuffled data: statistical and computational limits of permutation recovery. IEEE Trans. Inform. Theory, 64(5):3286–3300, 2018
2018
-
[35]
G. P´ olya. Herleitung des Gaußschen Fehlergesetzes aus einer Funktionalgleichung. Mathematische Zeitschrift , 18(1):96–108, 1923
1923
-
[36]
Rigollet and J
P. Rigollet and J. Weed. Uncoupled isotonic regression via minimum wasserstein deconvolution. Information and Inference: A Journal of the IMA , 8(4):691–717, 2019
2019
-
[37]
Slawski and E
M. Slawski and E. Ben-David. Linear Regression with Sparsely Permuted Data. Elec- tronic Journal of Statistics , 13:1–36, 2019
2019
-
[38]
Slawski and B
M. Slawski and B. Sen. Permuted and Unlinked Monotone Regression in Rd: an approach based on mixture modeling and optimal transport. Journal of Machine Learning Research, 25(183):1–57, July 2024
2024
-
[39]
Slawski, B.T
M. Slawski, B.T. West, P. Bukke, G. Diao, Z. Wang, and E. Ben-David. A general framework for regression with mismatched data based on mixture modeling. Journal of the Royal Statistical Society Series A, forthcoming , 2024
2024
-
[40]
L. Sweeney. Computational disclosure control: A primer on data privacy protection . PhD thesis, Massachusetts Institute of Technology, 2001
2001
-
[41]
Tsakiris and L
M. Tsakiris and L. Peng. Homomorphic sensing. In International Conference on Machine Learning (ICML) , pages 6335–6344, 2019
2019
-
[42]
Unnikrishnan, S
J. Unnikrishnan, S. Haghighatshoar, and M. Vetterli. Unlabeled sensing with random linear measurements. IEEE Trans. Inform. Theory , 64(5):3237–3253, 2018
2018
-
[43]
Wang and A
K. Wang and A. Seigal. Identifiability of overcomplete independent component anal- ysis. arXiv:2401.14709, 2024
2024 arXiv
-
[44]
W.E. Winkler. Record linkage. Wiley Interdisciplinary Reviews: Computational Statis- tics, 2(5):503–511, 2010
2010
-
[45]
Moreover,
V. Zolotarev. One-dimensional stable distributions . American Mathematical Society, 1986. Appendix: Proofs Proof of Lemma 1 For some random variable Z, let Φ Z denote its characteristic function; i.e., Φ Z(t) = E[exp(itZ)], t ∈ R. Since X and ϵ are assumed to be independent, i...
1986
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.