{"id":"3abebc0c-1022-4565-a543-c19d37786f63","arxiv_id":"2507.13605","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A combined-likelihood approach for twin data achieves root-n consistent estimates of genetic impact without needing to know whether a trait's mean differs between the two members of a pair.","lead":"This paper develops a statistical method for estimating the genetic component of a trait from twin studies when the order of measurements within each twin pair is unknown. It shows that combining data from identical and fraternal twins yields accurate estimates of genetic impact with standard statistical guarantees.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The root-n result for rhoD rests on the untested equal-variance assumption sigmaM^2=sigmaD^2; a violation would bias the genetic-impact estimate, and the paper's informal check does not resolve this.","rationale":"The reader's conditional verdict is appropriate. The paper's central methodological novelty is that the combined likelihood converts the slow n^{-1/4} rate for ρD under the separate mixture into root-n by using MZ data to pin down σ^2. This mechanism only works if σ_M^2=σ_D^2. I checked the local structure: at μD1=μD2, the first-order effect of the mean difference in the DZ density is a quadratic form H that lies in the span of the score functions for ρD and σ^2 in the separate model, so without the MZ anchor, ρD and σ^2 can absorb the mean-difference perturbation and ρD cannot be root-n. With σ^2 anchored, the confounding is broken and root-n for ρD is plausible. This confirms that the equal-variance assumption is not a minor regularity condition but the mechanism behind Theorem 2. The paper's Section 5 check is informal, and the authors state that the LRT limiting distribution is beyond scope; no unequal-variance simulation is provided. A misspecified variance would bias bρD and hence bδ, and the bootstrap intervals in Table 4 would not have nominal coverage. The proposed simulation is a direct, low-cost check. If it shows robustness to moderate variance inequality, the conditional acceptance can be upgraded; if it shows bias or undercoverage, the paper's claims should be restricted to settings where the assumption is testable or the method should be modified.","tokens_in":11054,"tokens_out":34955,"duration_ms":362888,"concrete_test":"Run a simulation under violated equal variance: generate 1000 datasets with σ_M^2=1, σ_D^2=1.5, μD1=μD2=0, ρM=0.9, ρD=0.3, and nM=nD=400; fit the combined likelihood (3) and compute the median bias and 95% bootstrap coverage of bδ. If the median bias exceeds 0.05 or coverage falls below 0.90, the equal-variance assumption is load-bearing and the Table 4 results are conditional on it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Theorem 2) is that bρD is root-n consistent even when μD1=μD2. The heuristic in Section 3.2 is that lM_n provides a root-n estimate of σ^2, making σ^2 'effectively known' in lD_n. This is precisely why the assumption σ_M^2=σ_D^2=σ^2 in equation (3) is load-bearing: if the true DZ variance differs, the combined likelihood maximizes a misspecified DZ likelihood with a wrong variance, and bρD is not guaranteed consistent, let alone root-n. The paper's only check (Section 5) compares separate-mixture variance estimates and reports they are 'quite close', but it explicitly declines to derive the likelihood-ratio test limiting distribution as 'beyond the scope of this work'. No simulation under unequal variances is reported, so the sensitivity of bδ to this assumption is unknown. Since δ=ρM−ρD is the paper's target, a violation of the equal-variance assumption directly threatens the validity of the immune-trait intervals in Table 4. The companion separate-mixture rates in Theorem 1 (n^{-1/8} for μD1 and μD2) are also hard to reconcile with the root-n estimability of the marginal mean 0.5(μD1+μD2), underscoring that the asymptotic development needs independent scrutiny before the central claim is accepted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a bivariate mixture model for twin data in which the ordering within each pair is unknown. For monozygotic (MZ) twins the model reduces to a bivariate normal with equal means; for dizygotic (DZ) twins it is a 50-50 mixture of two bivariate normals with swapped means. The authors study maximum likelihood estimation from the separate MZ/DZ likelihoods and from a combined likelihood that imposes a common variance σ². Theorem 2 claims that under the combined likelihood, ρD and ρM are root-n consistent even when μD1 = μD2, while the separate mixture has slower rates (Theorem 1). This motivates inference on the genetic impact δ = ρM − ρD. A likelihood ratio homogeneity test for μD1 = μD2 is proposed (Theorem 3). Simulations and an application to six immune traits are presented.","tokens_in":11343,"tokens_out":18831,"duration_ms":189523,"significance":"If the central claim is correct, the combined-likelihood approach is a practical advance: it gives standard-rates inference for the genetic impact without genotype data, in a setting (unordered pairs with potential heterogeneity) where Pearson correlations are biased. The paper is honest about the difficulty of the homogeneity test and supplies a data application. The simulations support the claimed rates for ρD and δ under the assumed equal-variance model. However, the significance is currently conditional: the main theorems are not verifiable from the submitted text (proofs are in a missing supplement), Theorem 1's rates appear inconsistent with standard local expansions, and the equal-variance assumption on which the root-n result for ρD rests is only informally checked. These points must be resolved before the practical claims are accepted.","major_comments":[{"comment":"All three theorems are stated without proof in the main text; the reader is referred to a supplementary document that is not part of the submitted material. Since Theorem 2 is the central claim and Theorem 1 contains non-standard rates, the manuscript cannot be evaluated for correctness as presented. Please provide the complete proofs (ideally in the main text or an accessible supplement) and state the regularity conditions and the parameterization used.","section":"Section 3, Theorems 1–3"},{"comment":"The combined likelihood (3) is constructed under the equal-variance assumption σ_M² = σ_D² = σ², and the heuristic in Section 3.2 explicitly uses the MZ likelihood to anchor σ² for the DZ mixture. If the true variances differ, the DZ likelihood is misspecified and the claimed root-n consistency of ρD, and hence of δ, is not guaranteed; the intervals in Table 4 would then be unreliable. Section 5 checks this assumption only by noting that separate variance estimates are \"quite close\", with no test or numerical evidence. Please add a formal or simulation-based assessment of sensitivity to variance inequality (e.g., coverage of bootstrap intervals for δ under σ_D² ≠ σ_M²), or modify the method to be robust to this assumption.","section":"Section 2, Eq. (3) and Section 5"},{"comment":"The rates in Theorem 1 are surprising and need justification. With α = 0.5 fixed, the DZ mixture density at μD1 = μD2 = μD0 depends on (μD1 − μD2) only through quadratic and higher powers, so the standard local expansion gives bμD1 − bμD2 = Op(n^{−1/4}) for the difference and root-n for the mean (μD1 + μD2)/2, while ρD and σ_D² should be root-n since the model reduces to a single bivariate normal. The stated rates n^{−1/8} for the individual means and n^{−1/4} for ρD and σ_D² are not reconciled with this heuristic or with the separate-likelihood simulations in Table 1 (which do show sd ratios consistent with n^{−1/4} for ρD). Please clarify whether α is actually estimated, or provide the proof and explain the n^{−1/8} rate.","section":"Section 3.1, Theorem 1"}],"minor_comments":[{"comment":"The study is described as including 324 twins, with 96 MZ and 288 DZ twins; these numbers sum to 384. Please correct the total and verify the sample sizes used in Table 4.","section":"Introduction"},{"comment":"The displayed formula for Pearson's correlation coefficient is garbled: the second factor in the numerator appears to be missing Y_{i2}. Please correct the formula so that the convergence result (2) is connected to it.","section":"Section 2"},{"comment":"In Theorem 1, \"bσ2 D2\" appears to be a typo for \"bσ2_D\".","section":"Section 3.1"},{"comment":"The caption contains misspellings (\"qunatile\") and the phrase \"both both\"; please fix these presentation issues.","section":"Figure 1"},{"comment":"The adjusted limiting distribution relies on the empirically fitted constant a = 6.828 in a_n = 0.5 + 6.828/nM. Since this is a fitted quantity, please report its precision and show type I error simulations of the homogeneity test at the adjusted level, not only QQ plots.","section":"Section 3.3"},{"comment":"Please state the sample sizes used in the application and describe the bootstrap procedure (number of resamples, method for the CI).","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The supplementary proofs are essential; I could not assess the theorems from the submitted text. The empirical constant a in the homogeneity test is a fitted parameter; this is acceptable if properly validated, but the fitted nature should be disclosed. The equal-variance assumption is a limitation that should be highlighted in the abstract or discussion if no robustness evidence is provided."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a real advance for twin data with unordered pairs. The combined MZ/DZ likelihood is a new construction, and the claim that rho_D can be estimated at root-n speed even when the DZ means are equal is exactly the kind of result that makes the method usable in practice. The homogeneity test with the 0.5/0.5 chi-square mixture null is also new in this setting. The paper is honest about where the gain comes from: the MZ data anchor the variance, so the DZ mixture no longer has to estimate everything at the slow mixture rate.\n\nThe core theory, however, is in a supplementary document not included here, so I cannot verify Theorem 2, the load-bearing result. What I can verify from the simulations is that the combined estimator shrinks bias and variance roughly at root-n in the homogeneous case, which is consistent with the claim. The bias formula (2) is clearly correct, and the immune-trait application is a nice demonstration.\n\nNow the soft spots, in order of seriousness. First, the equal-variance assumption sigma_M^2 = sigma_D^2 is load-bearing. The heuristic in Section 3.2 says that with sigma^2 effectively known from the MZ data, the DZ mixture converges faster. If that assumption is violated, the combined likelihood is misspecified and there is no guarantee that rho_D is consistent. The paper checks this only informally in Section 5 by comparing separate variance estimates, and explicitly declines to derive the LRT for equal variances. No simulation under unequal variances is reported. That is the biggest gap.\n\nSecond, the sample size inconsistency: 96 MZ + 288 DZ does not equal 324. That is the kind of sloppy error that makes a reviewer nervous about the data handling.\n\nThird, the homogeneity test correction uses an empirically fitted constant (a = 6.828). That is fine, but it is fitted, not derived, so the test's calibration rests on simulation.\n\nIf the supplementary proofs hold up and the equal-variance assumption is given a sensitivity analysis or a proper test, this is a solid paper. I would send it to a serious referee, with the explicit request to scrutinize Theorem 2 and the variance assumption. The core idea is novel and worth the referee time.","headline":"Root-n result for rho_D is interesting and likely right, but the equal-variance assumption is load-bearing and the proof lives in an unavailable supplement.","tokens_in":11838,"tokens_out":3175,"would_cite":true,"duration_ms":33595,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62F03","62H20","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Pooling monozygotic and dizygotic twins into one likelihood with a shared variance makes dizygotic correlation estimation root-n consistent even for homogeneous traits, so the genetic impact $\\delta=\\rho_M-\\rho_D$ can be inferred with…","keywords":["correlation coefficient","maximum likelihood estimation","mixture distribution","unordered pairs","twin study","genetic impact","root-n consistency","homogeneity test"],"falsifier":"Simulate homogeneous DZ data with $\\mu_{D1}=\\mu_{D2}$ but $\\sigma_D^2\\neq\\sigma_M^2$, fit the combined likelihood, and track the empirical standard deviation of $\\hat\\rho_D$ as $n$ grows; if it does not shrink like $n^{-1/2}$, the equal-variance condition is essential. Equivalently, compute a likelihood-ratio test of $\\sigma_M^2=\\sigma_D^2$ on the immune-trait data and check whether bootstrap coverage of $\\delta$ degrades when the test rejects.","tokens_in":10877,"feed_emoji":"🧬","tokens_out":8357,"duration_ms":91598,"temperature":0.7,"pith_summary":"The paper addresses a classic blind spot in twin studies: the two measurements in a twin pair arrive without a meaningful label, and standard correlation measures either ignore the order or are biased by it. The authors model the unordered pair as a two-component mixture of bivariate normals and fit it jointly to monozygotic and dizygotic twins, assuming the two zygosity groups share a variance. The central result is that this combined likelihood estimates the dizygotic correlation $\\rho_D$ at root-n speed even in the homogeneous case where the dizygotic means are equal—the case that makes the usual separate mixture estimator slow down to $n^{-1/4}$. Consequently the genetic impact $\\delta=\\rho_M-\\rho_D$, Falconer's heritability-style contrast, is consistently estimable with bootstrap confidence intervals. If true, this gives a practical route to assessing genetic control of immune traits and any other unordered paired data.","feed_headline":"Joint likelihood restores standard accuracy to twin genetic estimates","feed_subtitle":"Combining MZ and DZ twins in one likelihood corrects the order bias that makes Pearson's correlation overstate genetic impact.","key_machinery":"The load-bearing object is the mixture density $f(y_1,y_2;\\mu_1,\\mu_2,\\rho,\\sigma^2)=0.5\\phi(y_1,y_2;\\mu_1,\\mu_2,\\rho,\\sigma^2)+0.5\\phi(y_1,y_2;\\mu_2,\\mu_1,\\rho,\\sigma^2)$, which says the observed pair is the labeled pair in either of two equally likely orders. The combined log-likelihood is $l_C=l_M+l_D$, where $l_M$ is the ordinary bivariate normal log-likelihood for MZ pairs with $\\mu_1=\\mu_2$ and $l_D$ is the mixture log-likelihood for DZ pairs, with $\\sigma^2$ shared across zygosity. Sharing the variance is what does the work: the regular MZ part estimates $\\sigma^2$ and $\\rho_M$ at the usual $n^{-1/2}$ rate, and because $\\sigma^2$ is then effectively known inside the DZ mixture, the poorly identified mean labels $\\mu_{D1}$ and $\\mu_{D2}$ may drift at $n^{-1/4}$ without dragging $\\rho_D$ and $\\sigma^2$ below $n^{-1/2}$. The likelihood-ratio statistic for homogeneity then inherits the boundary-type mixture limit $0.5\\chi_0^2+0.5\\chi_1^2$, with a small-sample adjustment $a_n=0.5+6.828/n_M$.","core_discovery":"The paper's central claim is Theorem 2: under the combined likelihood $l_C$ built from MZ and DZ data with a common variance $\\sigma^2$, the maximum likelihood estimators satisfy $\\hat\\rho_M-\\rho_{M0}=O_p(n^{-1/2})$ and $\\hat\\rho_D-\\rho_{D0}=O_p(n^{-1/2})$ even when the DZ means coincide, i.e. $\\mu_{D1}=\\mu_{D2}=\\mu_{D0}$. In that homogeneous case the separate DZ mixture alone only gives $\\hat\\rho_D-\\rho_{D0}=O_p(n_D^{-1/4})$ (Theorem 1). The combined estimator therefore keeps the genetic impact $\\delta=\\rho_M-\\rho_D$ at root-n consistency no matter whether the trait is homogeneous or heterogeneous, and simulations show $\\hat\\delta$ is nearly unbiased and symmetrically distributed, with bootstrap intervals near nominal coverage. The same likelihood yields a homogeneity test for $\\mu_{D1}=\\mu_{D2}$ whose null limiting distribution is $0.5\\chi_0^2+0.5\\chi_1^2$, with a simulation-calibrated small-sample adjustment.","pith_inferences":["The equal-variance assumption $\\sigma_M^2=\\sigma_D^2$ is checked only informally in Section 5; a natural extension would be a sensitivity analysis or a formal likelihood-ratio test, and the root-n claim for $\\rho_D$ would need re-examination if that equality fails.","The empirical correction $a_n=0.5+6.828/n_M$ is fit by simulation with $n_M$ as the only covariate; an analytic higher-order expansion could replace it and would clarify how the slow $n^{-1/4}$ mean estimators affect small-sample calibration.","In the immune-trait application, traits flagged as possibly heterogeneous (B cells, CD4, CD8) are exactly those where Pearson's $\\delta$ exceeds the mixture $\\delta$; this suggests a diagnostic rule that a large gap between Pearson and mixture estimates is itself evidence of order-labeling bias.","For other latent-label bivariate data, such as allele-specific expression or two-channel assays without strand assignment, the same combined-likelihood trick may recover root-n correlation estimation even when component means are equal, though the paper does not test these settings."],"forward_implications":["If the combined-mixture result holds, genetic impact $\\delta=\\rho_M-\\rho_D$ can be estimated with ordinary bootstrap inference for every trait, because $\\hat\\delta$ is root-n consistent whether or not the trait is homogeneous.","Pearson's correlation should not be used for DZ twins when a trait may be heterogeneous: the paper's formula shows it underestimates $\\rho_D$ by an amount growing with $(\\mu_{D1}-\\mu_{D2})^2$, so it overstates $\\delta$, while the mixture estimator removes that bias.","The homogeneity test with limiting distribution $0.5\\chi_0^2+0.5\\chi_1^2$ and the small-sample adjustment $a_n=0.5+6.828/n_M$ provides a practical check for whether DZ means differ.","The method transfers to any unordered paired data, not only twins, because the mixture only assumes random labeling and equal variance for the groups being compared."],"supporting_citations":[{"why":"Supplies the classical MLE asymptotics that give root-n normality in the heterogeneous case and anchor the proof of Theorem 2.","marker":"Serfling, 2000"},{"why":"Establishes the intrinsic slow-convergence and identifiability problems in mixture homogeneity testing that motivate the combined likelihood.","marker":"Dacunha-Castelle & Gassiat, 1999"},{"why":"Provides the likelihood-ratio testing framework for finite mixture homogeneity that underlies Theorem 3's null distribution.","marker":"Chen & Chen, 2001"},{"why":"Motivates the small-sample correction of the homogeneity test's limiting distribution by matching first moments.","marker":"Bartlett, 1937"},{"why":"Supplies Falconer's formula $\\delta=\\rho_M-\\rho_D$ that defines genetic impact in twin studies.","marker":"Frankham, 1996"},{"why":"Provides the immune-trait twin dataset and prior evidence on genetic architecture used in the application.","marker":"Roederer et al., 2015"},{"why":"Formalizes the unordered-pair problem that the mixture model addresses.","marker":"Hinkley, 1973"},{"why":"Supplies additional inference methods for unidentified paired data that the proposed mixture approach extends.","marker":"Davies & Phillips, 1988"}],"fun_headline_variants":["Combined twin likelihood restores root-n genetic impact","Joint MZ-DZ likelihood fixes genetic impact estimation","Twin genetic impact root-n consistent with combined likelihood","Unified twin likelihood beats mixture convergence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that MZ and DZ twins have the same variance, because the MZ data anchor the variance in the DZ mixture and the claimed root-n rate for $\\rho_D$ collapses if that equality fails; the paper checks this assumption only informally in Section 5, with no formal test.","fun_headline_variants_meta":{"raw":{"variants":["Combined twin likelihood restores root-n genetic impact","Joint MZ-DZ likelihood fixes genetic impact estimation","Twin genetic impact root-n consistent with combined likelihood","Unified twin likelihood beats mixture convergence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000244,"raw_usage":{"total_tokens":1488,"prompt_tokens":854,"completion_tokens":634,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":574}},"tokens_in":470,"tokens_out":634,"duration_ms":7203,"temperature":1.0,"reasoning_tokens":574,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:20:22.097574+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate homogeneous DZ data with $\\mu_{D1}=\\mu_{D2}$ but $\\sigma_D^2\\neq\\sigma_M^2$, fit the combined likelihood, and track the empirical standard deviation of $\\hat\\rho_D$ as $n$ grows; if it does not shrink like $n^{-1/2}$, the equal-variance condition is essential. Equivalently, compute a likelihood-ratio test of $\\sigma_M^2=\\sigma_D^2$ on the immune-trait data and check whether bootstrap coverage of $\\delta$ degrades when the test rejects.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the classical MLE asymptotics that give root-n normality in the heterogeneous case and anchor the proof of Theorem 2."},{"cited_title":"& Gassiat, E","cited_arxiv_id":null,"evidence_quote":"Establishes the intrinsic slow-convergence and identifiability problems in mixture homogeneity testing that motivate the combined likelihood."},{"cited_title":"& Chen, J","cited_arxiv_id":null,"evidence_quote":"Provides the likelihood-ratio testing framework for finite mixture homogeneity that underlies Theorem 3's null distribution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the small-sample correction of the homogeneity test's limiting distribution by matching first moments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Falconer's formula $\\delta=\\rho_M-\\rho_D$ that defines genetic impact in twin studies."},{"cited_title":", Quaye, L","cited_arxiv_id":null,"evidence_quote":"Provides the immune-trait twin dataset and prior evidence on genetic architecture used in the application."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Formalizes the unordered-pair problem that the mixture model addresses."},{"cited_title":"& Phillips, A","cited_arxiv_id":null,"evidence_quote":"Supplies additional inference methods for unidentified paired data that the proposed mixture approach extends."}],"review_version":1}