{"id":"f814743d-77bd-4813-9fbf-c1ac3cfe582f","arxiv_id":"2509.06005","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A Mallows-type criterion for selecting and averaging spatial weights matrices in multivariate SAR models, with asymptotic optimality and selection consistency.","lead":"This paper develops rules for picking or blending spatial neighbor matrices in multivariate spatial autoregressive models. The authors prove the chosen model is as good as the best candidate in large samples, and test the methods on Sina Weibo posting data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Penalty term relies on an unverified approximate derivative of D̂_k; conditions C7/C10/C13 are assumed, not established, so the optimality/consistency theorems are conditional on an unproved rate.","rationale":"The reader's weakest-assumption analysis identifies exactly the same load-bearing concern: the feasible Mallows-type criterion in eq. (8) contains a penalty term built from an approximate derivative of D̂_k, and conditions (C7), (C10), and (C13) are high-level rate conditions on this term that are never verified from primitives. My reading of the proofs confirms that these conditions are directly load-bearing: (B2a) in the proof of Theorem 3.1 is nothing but condition (C7), and the analogous steps in Theorems 3.2 and 4.1 rely on (C10) and (C13). The paper's Remark 3 is unusually explicit about the gap: the exact derivative is not derived, the approximation ignores randomness of Σ̂, and the authors state without proof that the rate conditions are satisfied by the approximation. Because the asymptotic optimality and selection consistency claims depend on this unverified condition, the correct verdict is CONDITIONAL rather than ACCEPT. I do not see a stronger reason to reject: the proof skeletons are standard, the simulation evidence is broadly supportive, and the high-level conditions are stated clearly. Other potential concerns, such as the divergence of K or the eigenvalue conditions in (C5), are less central and are standard in this literature. Therefore no change to the reader's verdict is needed.","tokens_in":37495,"tokens_out":10768,"duration_ms":111816,"concrete_test":"Using the Section 5.1 design (n = 300, q = 2, K = 4), compute the implemented penalty exactly as coded, and also compute a numerically exact Jacobian ∂vec(D̂_k)/∂y^T by finite differences or automatic differentiation of the full estimation routine including the covariance estimator Σ̂. For n = 100, 400, 1600, estimate sup_k |tr(∂vec(D̂_k)/∂y^T (y^T⊗Ω̂) ∂vec(P̃_k)/∂vec(D̂_k)^T)|/ξ_n and the analogous ratio with ξ*_n when W* is in the candidate set. If the ratio does not converge to zero in probability, or if the exact Jacobian differs from the Appendix A approximation by more than o_p(ξ_n), then conditions (C7)/(C10)/(C13) fail and the theorems' assumptions are violated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing gap is the penalty term in the feasible criterion (Section 3, eq. 8): 2 tr(∂vec(D̂_k)/∂y^T (y^T⊗Ω̂) ∂vec(P̃_k)/∂vec(D̂_k)^T). Theorems 3.1, 3.2, and 4.1 place rate conditions (C7), (C10), (C13) on this term, and the proof of Theorem 3.1 uses (C7) directly as (B2a). Remark 3 and Appendix A state that the exact derivative is 'very difficult to derive' and instead use an approximation that ignores the randomness of Σ̂. No proof shows that this approximation satisfies the required o_p(ξ_n) or o_p(ξ*_n) rates. The conditions are written for the exact derivative, while the implemented criterion uses the approximation; if the omitted ∂Σ̂/∂y^T contributions are not o_p(ξ_n), the penalty is not asymptotically negligible, the uniform comparison (B2a) breaks, and the optimality/consistency conclusions are unsupported. This is not an internal contradiction, but it makes the paper's central claims conditional on an unproved high-level assumption. Simulations are suggestive but do not establish the required rates.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies model selection and model averaging over a finite set of candidate spatial weights matrices for the multivariate spatial autoregressive model Y = WYD + XB + E. It proposes a Mallows-type criterion (eq. 8) whose penalty contains a derivative term ∂vec(D̂_k)/∂y^T, establishes asymptotic optimality of the selected model when the true weights matrix is not among the candidates (Theorem 3.1), selection consistency when it is (Theorem 3.2), and asymptotic optimality of a model averaging estimator (Theorem 4.1). The theorems are conditioned on high-level rate conditions, notably (C7), (C10), and (C13), involving the derivative penalty. The paper also contains extensive simulations, including non-normal errors and geographically constructed weights, a comparison with high-order MSAR, and an application to Sina Weibo data.","tokens_in":37802,"tokens_out":7114,"duration_ms":86464,"significance":"If the theoretical claims are fully supported, the paper is a useful extension of the univariate SAR model-selection literature (Zhang and Yu 2018) to multivariate responses. The allowance for a growing number of candidate matrices, the treatment of the case where the true weights matrix is not in the candidate set, and the practical application to social-network data are all valuable. The paper also ships code and data, and the simulation study is unusually thorough, covering normal and t-distributed errors, several candidate sets, and a comparison with high-order MSAR. The main weakness is that the central theorems rest on rate conditions for an approximated derivative whose required rates are asserted but not verified.","major_comments":[{"comment":"The load-bearing gap is the penalty term in the feasible criterion, eq. (8): 2 tr(∂vec(D̂_k)/∂y^T (y^T⊗Ω̂) ∂vec(P̃_k)/∂vec(D̂_k)^T). The proofs of Theorems 3.1, 3.2, and 4.1 use Conditions (C7), (C10), and (C13) to control this term (e.g., (C7) gives (B2a) in Appendix B). However, Remark 3 states that the exact derivative ∂vec(D̂_k)/∂y^T is 'very difficult to derive', and Appendix A derives instead an approximation that ignores the randomness of the covariance estimator Σ̂_k. No formal result shows that this approximation satisfies the required o_p(ξ_n), o_p(ξ*_n), or o_p(ξ̃_n) rates uniformly over k, nor that the difference between the exact and approximated penalty is asymptotically negligible. If the omitted ∂Σ̂_k/∂y^T contributions are not o_p of the relevant rates, the uniform comparison in (B2a) fails and the optimality/consistency conclusions are unsupported. The statement in Rema","section":"Section 2, Condition (C5)"},{"comment":"Condition (C5) as printed is not dimensionally coherent. Since fX = I_q⊗X ∈ R^{nq×pq}, the expression fX^T W_k (I_{nq} − D^T⊗W*)^{-1} fX/n is not well-defined as written because fX^T is pq×nq while W_k is n×n. Presumably the intended expression involves (I_q⊗W_k) or an equivalent block-diagonal embedding. Because (C5) is used in the theoretical framework and in verifying other conditions, this needs to be corrected before the assumptions can be checked.","section":"Section 3, Appendix C"}],"minor_comments":[{"comment":"There are several typos and nonstandard encodings: 'Purcha' for 'Prucha' (Introduction), 'misspecificaitions' (Concluding Remarks), 'Techincal' (Introduction), and 'user¡¯s' in the abstract.","section":"Throughout"},{"comment":"In Case 2 of Table 1, the row label 'SAR BasedY2' for the first SAR block appears to be a typo for 'BasedY1'; the subsequent block is labeled BasedY2.","section":"Section 5.1, Table 1"},{"comment":"The notation n_pq for the number of parameters (q² + pq + q(q+1)/2) is easily confused with the sample-size/product dimensions n×pq. Please use a different symbol or clarify.","section":"Section 3, after eq. (9)"},{"comment":"The derivation in Appendix A is very dense and some expressions are difficult to parse (e.g., the definitions of F_k^{ij}, the derivative ∂M_k/∂d_{st}, and the final assembling of ∂Q/∂d). A short explanation of the matrix layout and the notational conventions would improve reproducibility.","section":"Appendix A"},{"comment":"The text says code and data are openly available in an online supplementary material, but the arXiv version contains no link or repository identifier. Please provide a working link.","section":"End of paper"}],"recommendation":"major_revision","confidential_remarks":"The main issue is the unverified derivative approximation underlying Conditions (C7), (C10), and (C13). This is a genuine gap, but it seems fixable: the authors could either prove the required rates for the approximation used in Appendix A or modify the conditions to be stated directly for the implemented criterion and verify them. If that is supplied, the paper would be a solid contribution. The manuscript also relies substantially on Zhu et al. (2020) for the estimation theory; that is a standard black-box reliance and does not itself raise circularity concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one-line take: this paper extends Mallows-type weights-matrix selection and averaging to multivariate SAR, and that's a real contribution; but the headline theorems are conditional on a rate condition on an approximate derivative that the authors do not verify, so the central claims are not fully established.\n\nWhat's new and good: the extension from univariate SAR (Zhang and Yu 2018) to multivariate responses is genuine. The degrees-of-freedom term requires a new derivative computation because the estimated D appears through both the fitted values and the response covariance; the Stein's lemma derivation in Section 3 is careful, and the proof skeleton is standard Mallows-type asymptotics. Simulations are reasonably extensive: normal and t errors, geographic-distance weights, comparisons with high-order MSAR and univariate SAR. The Weibo application gives a concrete use and the results are cleanly reported. The paper is also honest about the hard step: Remark 3 and Appendix A state that the exact derivative is very difficult and that they approximate it by ignoring randomness of Σ̂.\n\nWhere it is soft: Conditions C7, C10, and C13 are high-level and simply assumed to hold for the approximation. Since those conditions are exactly what make the penalty term asymptotically negligible in the proofs (B2a and analogues), the asymptotic optimality and selection consistency are conditional on an unverified rate. If the omitted ∂Σ̂/∂y^T contributions are not o_p(ξ_n), the unbiasedness of the criterion can fail. The BIC analogy in Remark 3 is not quite right: in BIC the penalty rate is a simple deterministic sequence, whereas here the penalty depends on the data and on the approximation. Simulations are encouraging but cannot establish rates. This is not a fatal internal contradiction and not circularity; it is an incomplete proof. I would push the authors to either derive the needed bounds from primitives or state explicitly that the feasible criterion's asymptotic claims require the approximation condition and treat it as part of the assumptions.\n\nWho this is for: spatial econometricians and network analysts who care about weights-matrix specification. It deserves full peer review; I would send it, with expectation of substantial revision on the derivative-rate gap. I would cite it for the MSAR extension despite the gap.","headline":"A legitimate extension of Mallows-type weights-matrix selection to multivariate SAR, but the headline theorems rest on an unverified rate condition for an approximate derivative, so the central claims are conditional; worth a serious referee.","tokens_in":38266,"tokens_out":2190,"would_cite":true,"duration_ms":27925,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Mallows-type criterion for multivariate spatial autoregressive models selects the true spatial weights matrix when present, and otherwise matches the best misspecified candidate's prediction loss.","keywords":["spatial weights matrix","multivariate spatial autoregressive model","model selection","model averaging","Mallows criterion","asymptotic optimality","selection consistency","social network data"],"falsifier":"Simulate an MSAR model with known D and W*, and with at least one misspecified candidate W_k. Estimate the true degrees of freedom df_k by Monte Carlo covariance between fitted values and responses, and compare it with the penalty term computed from the Appendix A approximation. If the difference, scaled by n or by the risk R_k, does not vanish as n grows, then the criterion is not an unbiased risk estimator and the optimality theorems cannot hold in that setting.","tokens_in":37375,"feed_emoji":"🕸️","tokens_out":7491,"duration_ms":82140,"temperature":0.7,"pith_summary":"The paper extends spatial weights matrix selection and model averaging from univariate to multivariate spatial autoregressive (MSAR) models. Its central claim is that a Mallows-type criterion—an estimate of each candidate's squared prediction loss plus a degrees-of-freedom penalty—picks a model whose loss is asymptotically as small as the loss of the best candidate, and picks the true weights matrix when it is among the candidates. A companion model averaging estimator, chosen by minimizing the same kind of criterion over convex weights, is shown to reach the infeasible best weighted combination's loss. If true, the methods give a principled answer to a routine practical question: which neighborhood structure should be used when many are plausible. Simulations with normal and non-normal errors, and a Sina Weibo application on finance and economics posting behavior, support the theory in finite samples.","feed_headline":"One risk criterion picks the best spatial-neighborhood matrix","feed_subtitle":"For multivariate spatial autoregressive models, selection and averaging match the infeasible best candidate in loss.","key_machinery":"The central object is the criterion C_hat(k) = ||P_hat(k) y - y||^2 + 2(tr(P_hat(k) Omega_hat) + tr(partial vec(D_hat(k))/partial y^T (y^T ⊗ Omega_hat) partial vec(P_hat(k))/partial vec(D_hat(k))^T)), and its model-averaging analogue C_hat(w) = w^T H^T H w + 2 w^T h. C_hat(k) estimates R_k + tr(Omega); the first term is the in-sample fit and the second is twice the effective degrees of freedom, so minimizing it balances fit against model complexity. The same penalty structure is applied to the convex combination of candidate models, turning weight choice into a constrained quadratic program.","core_discovery":"The central discovery is that the prediction risk of a fitted MSAR model can be estimated unbiasedly up to an additive constant by a Mallows-type statistic, and that minimizing that statistic over candidate spatial weights matrices is asymptotically optimal. Specifically, if all candidates are misspecified, the selected model satisfies L_hat(k)/inf_k L_k →_p 1; if the true weights matrix is in the candidate set, P(W_hat(k) = W*) →_p 1. For averaging, the chosen weight vector satisfies L(hat(w))/inf_w L(w) →_p 1. The proof requires the penalty term to track the effective degrees of freedom, which the paper obtains through Stein's lemma and an approximated derivative of the estimated spatial d","pith_inferences":["Inference: the approximate derivative in Appendix A is the spot to attack—if a data-generating process violates the convergence-rate conditions (C7), (C10), or (C13), selection consistency and optimality could fail even though the large-sample theorems look general. A Monte Carlo comparison of the penalty with the true degrees of freedom would expose this.","Inference: the averaged weights matrix sum_k w_hat(k) W_k gives a fitted convex combination of connectivity mechanisms, inviting interpretation of the weights as an estimate of mixture proportions, although the paper does not provide standard errors for these weights.","Inference: the quadratic form of the averaging criterion makes model screening natural when the number of candidate matrices is large; the paper notes screening as future work, but the structure of C_hat(w) suggests a tractable route."],"forward_implications":["When the true spatial weights matrix is in the candidate set, the selected model recovers it with probability tending to one, so the selection result can be read as evidence about the true connectivity mechanism.","When the true matrix is not in the candidate set, the selected predictor's squared loss is asymptotically the same as the loss of the candidate that would have been best if we had known the truth.","The model averaging estimator is asymptotically optimal over all convex combinations of candidates and can beat every single candidate model in squared prediction loss.","The covariance estimator used in the penalty does not need to be consistent, so a single dense candidate matrix can be used for the correction term without breaking the theory.","On the Sina Weibo data, the bivariate MSAR selection and averaging are more stable and predict finance and economics posting behavior better than univariate SAR, with averaging weights pointing toward uniform followee influence or influence proportional to the followee's follower count."],"supporting_citations":[{"why":"Supplies the MSAR model and the least-squares type estimator whose fitted values define the prediction risk being minimized, plus regularity conditions used as (C2).","marker":"Zhu et al. (2020)"},{"why":"Establishes the univariate SAR weights-matrix selection and averaging template that this paper extends to multivariate responses.","marker":"Zhang and Yu (2018)"},{"why":"Provides the Mallows C_p risk-estimation idea that the proposed criterion is built on.","marker":"Mallows (1973)"},{"why":"Stein's lemma is used to derive the unbiasedness of the risk estimator and the degrees-of-freedom penalty.","marker":"Stein (1981)"},{"why":"Defines degrees of freedom as the covariance between fitted values and responses, which the penalty term estimates.","marker":"Efron (2004)"},{"why":"Least-squares model averaging by Mallows criterion supplies the averaging framework and the optimality argument.","marker":"Hansen (2007)"},{"why":"Provides the discrete-index-set asymptotic optimality theorem and moment techniques used in the proofs of Theorems 3.1 and 4.1.","marker":"Li (1987)"},{"why":"Gives moment bounds for quadratic forms in independent variables used to control the random fluctuation terms in the proofs.","marker":"Whittle (1960)"}],"fun_headline_variants":["Risk criterion selects optimal spatial weights matrix","Multivariate spatial models auto-select weights matrix","Mallows-type rule picks best spatial neighbors","Optimal spatial weights selection without true matrix"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The theory assumes that the approximate derivative of the estimated spatial dependence matrix, obtained by ignoring randomness in the estimated covariance, converges fast enough; the paper does not derive the exact derivative or prove that the approximation meets this rate.","fun_headline_variants_meta":{"raw":{"variants":["Risk criterion selects optimal spatial weights matrix","Multivariate spatial models auto-select weights matrix","Mallows-type rule picks best spatial neighbors","Optimal spatial weights selection without true matrix"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1329,"prompt_tokens":705,"completion_tokens":624,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":569}},"tokens_in":449,"tokens_out":624,"duration_ms":7497,"temperature":1.0,"reasoning_tokens":569,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T04:38:14.240639+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate an MSAR model with known D and W*, and with at least one misspecified candidate W_k. Estimate the true degrees of freedom df_k by Monte Carlo covariance between fitted values and responses, and compare it with the penalty term computed from the Appendix A approximation. If the difference, scaled by n or by the risk R_k, does not vanish as n grows, then the criterion is not an unbiased risk estimator and the optimality theorems cannot hold in that setting.","supporting_citations":[{"cited_title":", author Huang, D","cited_arxiv_id":null,"evidence_quote":"Supplies the MSAR model and the least-squares type estimator whose fitted values define the prediction risk being minimized, plus regularity conditions used as (C2)."},{"cited_title":", year 1973","cited_arxiv_id":null,"evidence_quote":"Provides the Mallows C_p risk-estimation idea that the proposed criterion is built on."},{"cited_title":", year 1981","cited_arxiv_id":null,"evidence_quote":"Stein's lemma is used to derive the unbiasedness of the risk estimator and the degrees-of-freedom penalty."},{"cited_title":", year 2004","cited_arxiv_id":null,"evidence_quote":"Defines degrees of freedom as the covariance between fitted values and responses, which the penalty term estimates."},{"cited_title":", year 2007","cited_arxiv_id":null,"evidence_quote":"Least-squares model averaging by Mallows criterion supplies the averaging framework and the optimality argument."},{"cited_title":", year 1960","cited_arxiv_id":null,"evidence_quote":"Gives moment bounds for quadratic forms in independent variables used to control the random fluctuation terms in the proofs."}],"review_version":1}