{"id":"551cb36a-d821-489a-85a1-eacd65057392","arxiv_id":"2501.13625","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper proves a variational formula for the mutual information in high-dimensional linear regression with AR(1) dependent rows, and shows empirically that VAMP often reaches the predicted optimal error.","lead":"This paper derives an exact formula for how much information a long high-dimensional time series carries about an unknown hidden signal. It gives a benchmark for whether practical estimation algorithms like VAMP are optimal on dependent data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lower-bound proof of Theorem II.1 relies on an incorrect Jacobian derivative; the main MI theorem is not rigorously established as written.","rationale":"The paper's principal deliverable is Theorem II.1, a rigorous single-letter MI formula; the MMSE statements and VAMP experiments are downstream of it. The adaptive-interpolation framework is credible, and much of the proof (concentration lemmas, time-derivative identity, endpoint computations) is standard. However, the lower-bound half is the delicate direction, and the Jacobian step in Lemma IV.5 is internally incorrect as written: the derivative claimed positive is not the derivative of the field F1 defined in (IV.25). This is not a disagreement with consensus; it is an error inside the proof. If it is a typo, the corrected Jacobian still needs to be verified; if the positivity fails, inequality (IV.24) may be false. Either way, the central claim is not currently fully supported. The reader's identified weak point, the per-block MMSE conjecture (II.23), is also real: the abstract overstates the MMSE result, and the numerical experiments cannot confirm Bayes-optimality near phase transitions. But that concern attaches to an additional claim, while the Jacobian gap threatens the main theorem itself. Hence I partially agree with the reader and would keep the verdict conditional, with the conditions being a corrected lower-bound proof and explicit labeling of (II.23) as conjectural.","tokens_in":40276,"tokens_out":12074,"duration_ms":110732,"concrete_test":"Specialize Lemma IV.5 to k=1, lambda=0 (standard Gaussian design), where (IV.25) becomes R1' = c/(rho - E<Q> + sigma^2), R2' = rho - E<Q>. Compute the exact Jacobian J(t) = det(partial R(t,epsilon)/partial epsilon) by differentiating E<Q> via the posterior (IV.5). Check (a) whether J >= 1 actually holds, and (b) whether the paper's displayed positive derivative formula appears anywhere in this calculation. If the formula is absent or J is not provably >= 1, the lower-bound proof must be fixed before Theorem II.1 can be accepted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is not the conjectural block-MMSE formula but a broken step in the proof of the central MI theorem. In Lemma IV.5 (lower bound), the Jacobian determinant of the map epsilon -> R(t,epsilon) is asserted to be at least 1 because partial_{R1,i}F1 = c/(pi k) int delta_i^2(theta)/(k^{-1} sum_j delta_j(theta) R1,j + sigma^2) > 0. But by (IV.25)-(IV.26), F1,i = A_i(rho - E<Q>_{t,epsilon}) = c/pi int delta_i(theta) dtheta / (sum_j l_j(rho - E<Q_j>) delta_j(theta) + sigma^2), which depends on R1 only through the posterior overlaps E<Q_j>; the displayed derivative is not partial F1/partial R1, and even the corresponding derivative of A_i with respect to its own argument is negative and has a squared denominator. The Liouville lower bound used to apply Lemma IV.3 is therefore unsupported. Since Lemma IV.5 supplies the liminf half of Theorem II.1, the rigorous replica formula is not established unless this calculation is corrected or replaced.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies Bayesian inference in a high-dimensional stochastic regression model whose design matrix is generated by a diagonal AR(1) dynamics, with the number of features growing proportionally to the number of observations. The main result (Theorem II.1) is a variational formula for the normalized mutual information between the observations and a dense signal, proved by the adaptive interpolation method for the case where the diagonal transition matrix has k distinct eigenvalues; Theorem II.2 extends this to general eigenvalue distributions by a Wasserstein-approximation argument. Theorem II.3 gives a formula for the measurement MMSE and relates it to the per-block MMSEs. The paper also presents numerical experiments suggesting that VAMP achieves near-Bayes-optimal performance in this non-rotationally-invariant setting, and it transparently labels the per-block MMSE formula (II.23) as a conjecture.","tokens_in":40530,"tokens_out":7776,"duration_ms":74474,"significance":"If the main theorem is rigorously established, the paper provides a first single-letter formula for the information content of a high-dimensional AR(1) time series about a dense signal, going beyond the i.i.d. and right-rotationally-invariant designs previously treated in the literature. The reduction of the known λ=0 and λ=λI cases to prior results is a useful sanity check, and the extension to general diagonal matrices by approximating the eigenvalue distribution is an elegant step. The numerical VAMP study is also valuable despite the lack of theoretical guarantees. However, the proof of the central lower bound contains a broken Jacobian-derivative step, and several load-bearing concentration estimates are delegated to prior works without verification. The paper is therefore not yet a complete rigorous treatment, although the overall approach and the claimed formulas are plausible.","major_comments":[{"comment":"The Liouville lower bound used to justify the change of variables from ε to R(t, ε) is not established. In the ODE (IV.25), F1,i is defined as A_i(ρ−E⟨Q_1⟩,...,ρ−E⟨Q_k⟩), so it depends on R1 only through the posterior overlaps E⟨Q_i⟩; the displayed derivative ∂R1,i F1 = c/(πk) ∫ δ_i²(θ)/(k^{-1}Σ_j δ_j(θ)R1,j + σ²) is not the partial derivative of F1 with respect to R1,i, and it does not follow from (IV.26). In fact the natural derivative of A_i with respect to its own argument would be negative and would contain a squared denominator. Since this positive Jacobian is the stated reason for applying Lemma IV.3, the liminf half of Theorem II.1 is not rigorously proved as written; the calculation must be corrected or replaced by a valid estimate of the Jacobian determinant.","section":"Section IV-A, Lemma IV.5 (Eqs. (IV.25)-(IV.26))"},{"comment":"The overlap concentration lemma is explicitly proved only in outline: the text says 'we outline the steps and omit the details' and asserts that inequalities (IV.18)-(IV.20) follow from the proofs in [4], [5]. This lemma is load-bearing for the fundamental identity Lemma IV.3 and hence for both bounds in Theorem II.1. The transfer from the models in [4], [5] to the present block-KMS design is not automatic, and the lemma also assumes the Jacobian regularity that is the subject of the previous comment. The full proof, or a precise statement of which results in [4], [5] apply verbatim and why, must be supplied.","section":"Section IV-A, Lemma IV.2"},{"comment":"The proof of Lemma IV.7 delegates the crucial concentration estimates to prior work: (E.9) is said to follow from the 'same proof as Lemma 9.1 of [7]', (E.10) is declared 'equivalent to Lemma 9.2 of [7]', and only (E.14) receives a new argument because the independence used in [7] fails. Since the present design is neither i.i.d. nor right-rotationally invariant, the transfer of these lemmas requires verification of their hypotheses or a self-contained proof. Without (E.9)-(E.11), the proof of Theorem II.3 is incomplete.","section":"Appendix E, estimates (E.9)-(E.11)"},{"comment":"The abstract and introduction claim derivation of 'minimum mean-square errors', but the rigorously proven statement is the measurement MMSE (II.21) together with the relation (II.22); the per-block and signal MMSEs in (II.23) are explicitly conjectural and rely on replica symmetry and uniqueness of the global minimizer of iRS on Γ. The numerical experiments in Section III compare VAMP against the conjectured curve (II.23), so the match is evidence for the conjecture rather than a proof of it. The text should be revised to state this limitation in the abstract and in the discussion of Figures 2 and 3, and the conditions under which the derivative in (II.20) can be interchanged with the variational formula should be stated.","section":"Section II-B and Section III"}],"minor_comments":[{"comment":"The caption appears to describe the panels inconsistently: it refers to 'On the right MMSE versus cN' and 'On the left we see MMSE ... versus 1/σ²', while the body text refers to Figure 2a and Figure 2b in the opposite order. Please align the caption with the actual panel layout.","section":"Figure 2 caption"},{"comment":"The sentence 'This allows as to apply Lemma IV.3' contains a typo and should read 'This allows us to apply Lemma IV.3'.","section":"Section IV-A, Step 4"},{"comment":"In the term E[⟨(r2,i(t) − (ρ − Qi))Z^T Λ_i,N u_t⟩], the placement of Qi inside the Gibbs bracket while r2,i(t) is deterministic should be clarified; it is currently ambiguous which quantities are quenched and which are averaged over the posterior.","section":"Section IV-A, proof of Lemma IV.1"},{"comment":"The bounded-difference proof is only sketched: the text states that showing ψ′(s) ≤ Cp^{-1} would imply the result, but the bound on ψ′(s) is not displayed. Please complete the argument or give a precise reference.","section":"Appendix D, Lemma D.5"},{"comment":"The statement that KMS matrices 'asymptotically share the same eigenspace' is used in (A.37) and in Appendix E, but it is stated without proof or a precise reference. A formal statement of the asymptotic joint eigenvalue distribution would make the argument easier to verify.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central MI theorem is plausible and the adaptive-interpolation framework is appropriate, but the proof as written contains a substantive gap in the Jacobian calculation of Lemma IV.5, and the self-acknowledged omissions in Lemma IV.2 and Appendix E are load-bearing. These issues are fixable within the scope of the paper, so major revision is appropriate rather than rejection. The conjectural status of (II.23) is disclosed in the text, but the abstract and the empirical discussion should not present the signal MMSE as a derived result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nHere's my read of Tieplova, Lahiry, Barbier. The problem they solve is real: non-sparse high-dimensional regression with an AR(1) design whose covariance is block-diagonal but not rotationally invariant. As far as I know, the multi-eigenvalue block case is new, and their replica formula reduces correctly to the i.i.d. and λI cases. That part is worth taking seriously.\n\nThe main theorem (II.1) is the contribution. The proof follows the adaptive-interpolation template, and many steps are standard. But I checked the lower-bound proof in Lemma IV.5, and the stress-test note is right: the claimed Jacobian derivative is not the derivative of F1 with respect to R1. F1 as defined is A_i(ρ−E⟨Q⟩); the displayed expression with δ_i^2/(k^{-1}Σ δ_j R1,j + σ^2) has no clear relation to that function, and the denominator is not the one in A_i. Since the Liouville argument needs the Jacobian determinant ≥1, the liminf half of Theorem II.1 is not established as written. This is a load-bearing gap, not a typo in a peripheral lemma. I would not call the theorem false—the adaptive-interpolation method is robust and the formula is plausible—but the proof as it stands is incomplete.\n\nAlso, the abstract says they derive MMSE, but the per-block MMSE formula (II.23) is explicitly conjectural. They do prove the measurement MMSE (II.21) and its connection to block MMSEs, but the block MMSE itself is not proven. That overstatement needs fixing.\n\nMinor issues: overlap concentration (Lemma IV.2) is outlined but details are deferred; Appendix E waves at [7] and leaves (E.9)-(E.11) as assertions; no code for the experiments. None of these are fatal on their own, but together with the Jacobian gap they mean the paper needs real revision before I'd trust Theorem II.1.\n\nWho is this for? People working on information-theoretic limits of structured random linear models and AMP. The problem is relevant, and if the proof is repaired, this would be a solid contribution.\n\nRecommendation: send it to a serious referee—the question is important enough—but the referee should be asked to scrutinize Lemma IV.5 specifically and the distinction between the conjectured block MMSE and the proved measurement MMSE. I would not desk-reject it.","headline":"Plausible and novel MI formula for block-AR(1) regression, but the lower-bound proof has a broken Jacobian step and the abstract overstates the MMSE result.","tokens_in":41032,"tokens_out":4776,"would_cite":false,"duration_ms":44225,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AR(1) time-series regression gets exact information limits","keywords":["high-dimensional time series","stochastic regression","mutual information","replica method","adaptive interpolation","vector approximate message passing","Kac-Murdock-Szegö matrix","minimum mean-square error"],"falsifier":"Run exact Bayesian inference on a small instance with $k=2$ eigenvalue blocks placed on opposite sides of the phase transition seen in Figure 2, compute the per-block MMSE by Monte Carlo, and compare it with $\\tilde{r}_{2,i}$ at the global minimum of $i_{\\mathrm{RS}}$. If the block MMSE deviates from $\\tilde{r}_{2,i}$ while the measurement MMSE still follows (II.21), the conjecture (II.23) is false even though the mutual-information formula remains correct.","tokens_in":40078,"feed_emoji":"📈","tokens_out":7634,"duration_ms":65493,"temperature":0.7,"pith_summary":"The paper asks what can be learned about a dense high-dimensional signal from observations of a time series whose covariates follow an AR(1) process, when the number of features and the number of samples grow at the same rate. It proves that the normalized mutual information between the observations and the signal converges to the value of a finite-dimensional variational problem: the infimum over one set of parameters and the supremum over another of a potential built from the AR(1) correlation spectrum. It also proves a closed-form formula for the measurement MMSE and states a conjecture for the per-block MMSE. The rigorously proven part covers the mutual information and the measurement MMSE; with the additional block-MMSE conjecture, the description becomes complete.","feed_headline":"AR(1) time-series regression gets exact information limits","feed_subtitle":"A single-letter formula gives mutual information and MMSE when features and samples grow together, with no sparsity assumed.","key_machinery":"The load-bearing object is the replica-symmetric potential $i_{\\mathrm{RS}}(r_1,r_2)$, whose two vector arguments act as control parameters for the block structure: $r_1$ couples to scalar denoising channels $\\beta\\mapsto\\sqrt{r_{1,i}}\\beta + Z$, while $r_2$ enters through the spectral density $\\delta_i(\\theta)$ of the Kac-Murdock-Szegö covariance matrix of each AR(1) column. The argument runs through adaptive interpolation: one interpolates between the original time-series channel and $k$ decoupled scalar channels, controls the derivatives via overlap concentration, and uses the known limiting eigenvalue distribution of KMS matrices to evaluate the log-det terms. The fixed-point equations (II.9) define the set $\\Gamma$ of critical points, and the inf-sup formula selects the global minimum.","core_discovery":"For the stochastic regression model $Y_\\mu = p^{-1/2} x_\\mu^\\top \\beta_0 + Z_\\mu$ with $x_{\\mu+1} = A_p x_\\mu + \\xi_\\mu$, where $A_p$ is diagonal with $k$ fixed eigenvalues, the paper establishes $\\lim_{p\\to\\infty} i_p = \\inf_{r_1\\in[0,\\infty)^k}\\sup_{r_2\\in[0,\\rho]^k} i_{\\mathrm{RS}}(r_1,r_2)$, with the replica-symmetric potential given by (II.6) and $\\delta_i(\\theta) = (1 - 2\\lambda_i\\cos\\theta + \\lambda_i^2)^{-1}$ the spectral density of the AR(1) column covariance. The proof, via adaptive interpolation, shows that the mutual information per parameter is governed by this low-dimensional potential even though the design matrix is not right-rotationally invariant. A second theorem proves the limiting measurement MMSE equals $\\frac{\\sigma^2}{\\pi}\\int_0^\\pi \\frac{\\sum_i l_i\\delta_i(\\theta)\\tilde{r}_{2,i}}{\\sum_i l_i\\delta_i(\\theta)\\tilde{r}_{2,i}+\\sigma^2}\\,d\\theta$ and relates it to the block MMSEs. The per-block MMSE formula (II.23), asserting that each block's error equals the saddle-point value $\\tilde{r}_{2,i}$, remains a conjecture, and the numerical experiments show VAMP matching these predictions away from phase transitions but becoming unstable near them.","pith_inferences":["Editorial extension: because only the spectral density $\\delta_i(\\theta)$ enters the formula, the same variational structure should also hold for other stationary Gaussian processes whose column covariance is asymptotically Toeplitz, such as ARMA processes; this is not tested in the paper.","Editorial extension: the observed VAMP instability exactly at phase transitions suggests the inf-sup formula can have multiple competing global minima; checking whether the block MMSE is discontinuous there would settle the conjecture and could inform when spectral initialization is needed.","Editorial extension: the same adaptive-interpolation proof likely extends to generalized linear observations on top of the AR(1) design, because the interpolation step decouples the temporal correlation from the likelihood."],"forward_implications":["For any fixed number $k$ of AR(1) eigenvalues, the exact asymptotic mutual information is computed by a $k$-dimensional variational problem, reducing the inference problem to a finite optimization.","The measurement MMSE is available through formula (II.21), requiring only the saddle point of $i_{\\mathrm{RS}}$ rather than a full posterior computation.","With an arbitrary limiting eigenvalue distribution for $A_p$, Theorem II.2 gives the same kind of formula with functions $r_1(\\lambda), r_2(\\lambda)$ in place of vectors, extending the result to continuously many AR(1) components.","If the per-block conjecture (II.23) holds, the posterior error on each block of coefficients is asymptotically $\\tilde{r}_{2,i}$, giving a complete block-by-block description of estimation limits.","The empirical VAMP results indicate that a practical algorithm can reach the predicted MMSE outside phase-transition regions even without right rotational invariance."],"supporting_citations":[{"why":"supplies the adaptive interpolation and overlap-concentration framework used to prove the replica formula","marker":"[4]"},{"why":"introduces the adaptive interpolation method that turns the replica prediction into a rigorous identity","marker":"[5]"},{"why":"the right-rotationally invariant counterpart this paper extends; its upper and lower bound strategy is adapted to the block model","marker":"[8]"},{"why":"the i.i.d. Gaussian design case recovered when $A_p=0$, providing the baseline single-letter formula","marker":"[3]"},{"why":"introduces the same stochastic regression model with AR(1) covariates in the sparse penalized setting","marker":"[11]"},{"why":"the I-MMSE relation used to convert the mutual-information derivative into the measurement MMSE","marker":"[28]"},{"why":"defines the VAMP algorithm whose empirical performance is compared with the theoretical MMSE","marker":"[48]"},{"why":"gives the replica analysis of eigenvalue-spectrum designs used for the $A_p=\\lambda I_p$ special case","marker":"[61]"}],"fun_headline_variants":["Exact info limits for high-dim time series without sparsity","Time series regression gets exact info limits, no sparsity","AR(1) time series: exact mutual info without sparsity","High-dim time series: exact info limits, VAMP robust","No sparsity needed: exact info limits for time series"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The unproven step is that the posterior error on each block settles at the single saddle-point value $\\tilde{r}_{2,i}$; this needs the replica-symmetric global minimum to be unique, and it can fail where the system has competing optimal states, which is exactly where the experiments show the algorithm becoming unstable.","fun_headline_variants_meta":{"raw":{"variants":["Exact info limits for high-dim time series without sparsity","Time series regression gets exact info limits, no sparsity","AR(1) time series: exact mutual info without sparsity","High-dim time series: exact info limits, VAMP robust","No sparsity needed: exact info limits for time series"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000272,"raw_usage":{"total_tokens":1667,"prompt_tokens":1017,"completion_tokens":650,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":565}},"tokens_in":633,"tokens_out":650,"duration_ms":5984,"temperature":1.0,"reasoning_tokens":565,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:47:11.202114+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run exact Bayesian inference on a small instance with $k=2$ eigenvalue blocks placed on opposite sides of the phase transition seen in Figure 2, compute the per-block MMSE by Monte Carlo, and compare it with $\\tilde{r}_{2,i}$ at the global minimum of $i_{\\mathrm{RS}}$. If the block MMSE deviates from $\\tilde{r}_{2,i}$ while the measurement MMSE still follows (II.21), the conjecture (II.23) is false even though the mutual-information formula remains correct.","supporting_citations":[{"cited_title":"Barbier, F","cited_arxiv_id":null,"evidence_quote":"supplies the adaptive interpolation and overlap-concentration framework used to prove the replica formula"},{"cited_title":"Barbier and N","cited_arxiv_id":null,"evidence_quote":"introduces the adaptive interpolation method that turns the replica prediction into a rigorous identity"},{"cited_title":"Barbier, N","cited_arxiv_id":null,"evidence_quote":"the right-rotationally invariant counterpart this paper extends; its upper and lower bound strategy is adapted to the block model"},{"cited_title":"Barbier, M","cited_arxiv_id":null,"evidence_quote":"the i.i.d. Gaussian design case recovered when $A_p=0$, providing the baseline single-letter formula"},{"cited_title":"Basu and G","cited_arxiv_id":null,"evidence_quote":"introduces the same stochastic regression model with AR(1) covariates in the sparse penalized setting"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the I-MMSE relation used to convert the mutual-information derivative into the measurement MMSE"},{"cited_title":"Rangan, P","cited_arxiv_id":null,"evidence_quote":"defines the VAMP algorithm whose empirical performance is compared with the theoretical MMSE"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"gives the replica analysis of eigenvalue-spectrum designs used for the $A_p=\\lambda I_p$ special case"}],"review_version":1}