{"id":"1c733a97-4d2b-4d33-b7a0-1f6b2e584929","arxiv_id":"2501.04607","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A Bayesian mixed-frequency VAR produces monthly GDP estimates for all 50 U.S. states plus DC, consistent with official quarterly data, and gives state nowcasts up to three months before the BEA.","lead":"This paper builds a large statistical model that estimates monthly economic output for every U.S. state from 1964 onward, using faster monthly indicators combined with official quarterly and yearly GDP data. The resulting monthly state GDP numbers, released months before official data, could help economists and policymakers track regional booms, busts, and spillovers in near real time.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unquantified two-stage MCMC approximation in §2.5 replaces monthly U.S. indicators by quarterly aggregates in the block that identifies pre-2005 state GDP; without a small-scale exact-MCMC comparison, the historical monthly series is not validated.","rationale":"The paper's central claim has two parts: a historical monthly state GDP series back to 1964 and a nowcasting product with a multi-month lead. The nowcasting evaluation is post-2005 and is therefore not directly distorted by the annual-monthly approximation, since quarterly state GDP is observed after 2005. But the historical series is the larger and more novel deliverable, and every pre-2005 monthly observation is generated through the approximate two-stage algorithm. The replaced conditioning set is the only channel through which monthly U.S. indicators enter the annual/quarterly block; if their within-quarter movements help identify the quarterly path of state GDP, the cut posterior will differ from the stated model's posterior in a way that is not measured. This is not an outside-consensus objection; it is an internal validity question about whether the reported algorithm targets the model described in Section 2.3. A small-scale exact comparison is feasible because the exact sampler's bottleneck is dimension, not the presence of the annual-monthly restriction itself. The other concerns in the reader's verdict—missing replication code, post-hoc COVID exclusion, and the lead-time wording—are real but secondary; they affect reproducibility and presentation, not the estimator's correctness. I therefore agree with the reader's weakest assumption and see no reason to change the CONDITIONAL verdict.","tokens_in":29920,"tokens_out":6961,"duration_ms":66855,"concrete_test":"Estimate a small-scale version of the model—for example, 3–5 states, one monthly U.S. indicator (say employment growth), one quarterly U.S. variable, with annual state GDP through 2004 and quarterly thereafter—using both the paper's two-stage approximate sampler and an exact MCMC sampler that directly implements the annual-monthly restriction (5). Compare the pre-2005 posterior distributions of monthly state GDP expressed as year-on-year growth (the quantity plotted in Figure 1): posterior medians, 68% credible intervals, and quarterly aggregates. If any state-month differs by more than 0.1 percentage point in median year-on-year growth, or credible interval widths change by more than 10%, the Section 2.5 approximation is not negligible and the historical monthly series is not validated as it stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the approximate MCMC decomposition in Section 2.5. Equation (8) is exact, but the second factor is replaced by p(y^S_{a,q}|y^S_a, y^{US}_q, y^{US}_{m,q}), conditioning on quarterly aggregates of monthly U.S. indicators instead of the monthly series themselves. The paper justifies this by saying 'the loss of information is likely to be small' and notes the approximation affects only pre-2005 draws of quarterly state GDP. However, those draws are then used as conditioning inputs in the first factor, p(y^S_{a,m}|y^S_{a,q}, y^{US}_q, y^{US}_m), which generates every pre-2005 monthly state GDP estimate. The resulting algorithm is therefore a cut/modular posterior, not the posterior of the stated model; it coincides with the exact posterior only if monthly U.S. indicators carry no extra information for quarterly state GDP beyond their quarterly aggregates. No simulation, sensitivity analysis, or error bound is provided for this condition. The full 51-state exact sampler is admittedly prohibitive, but whether the approximation is harmless is checkable in a smaller instance, and that check has not been reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a large Bayesian mixed-frequency vector autoregression (MF-VAR) that produces monthly real GDP estimates for the 50 U.S. states plus Washington, DC, from 1964 to 2024, using monthly U.S. indicators, quarterly U.S. GDP, quarterly state GDP from 2005, and annual state GDP before 2005. Temporal and cross-sectional aggregation constraints are imposed so that the latent monthly state series aggregates exactly to official BEA quarterly state and U.S. GDP. Estimation is based on a horseshoe-prior shrinkage and an approximate two-block MCMC algorithm that separates an annual/quarterly model from a quarterly/monthly model. The paper reports historical monthly state GDP estimates, business-cycle dating and connectedness analyses, and a real-time nowcasting exercise for 2007Q1-2024Q1 in which the joint model is compared with a state-specific MF-VAR benchmark.","tokens_in":30131,"tokens_out":11721,"duration_ms":110458,"significance":"If the results hold, the paper offers a new and potentially valuable data product: monthly state GDP estimates that are exactly consistent with official BEA aggregates by construction, plus a nowcasting tool with a lead over official releases. The strengths are the exact temporal and cross-sectional constraints, the real-time out-of-sample evaluation against a reasonable benchmark, the use of a horseshoe prior in a high-dimensional MF-VAR, and the transparency about the computational obstacles. The central caveat is that the pre-2005 historical monthly series depends on an unquantified approximation in Section 2.5, so the historical component of the product is not yet validated to the same standard as the post-2007 nowcast results.","major_comments":[{"comment":"The two-block MCMC algorithm conditions the annual/quarterly block on quarterly aggregates of the U.S. monthly variables rather than on the monthly series themselves. The paper's only justification is the sentence 'the loss of information is likely to be small,' and no simulation, sensitivity analysis, or error bound is given. Since draws of quarterly state GDP from this approximate block are then fed into the first factor in Eq. (8) that generates the monthly state GDP series, the approximation propagates into every pre-2005 monthly estimate and into the business-cycle and connectedness results based on those estimates. I request a small-scale exact-MCMC comparison (for example, on a few states with the same data frequencies) or a simulation study demonstrating that posterior medians and credible intervals for state GDP are insensitive to replacing monthly U.S. indicators with their quarterly aggregates.","section":"Section 2.5, Eq. (8)"},{"comment":"The out-of-sample evaluation covers only 2007Q1-2024Q1 and therefore only the regime in which quarterly state GDP is observed throughout; it does not validate the 1964-2004 historical monthly series that is produced under the annual-only regime and the approximate algorithm. The abstract's nowcast claim concerns the modern regime, but the paper's historical data product is a central output. The authors should either state this limitation prominently and supply the exact-MCMC validation from the previous comment, or provide an additional validation exercise for the annual-only regime (for example, a pseudo-out-of-sample experiment using annual data only).","section":"Section 3.3, Tables 1-4"},{"comment":"The claim that fixing the cross-sectional weights at sample averages and ignoring temporal variation in state GDP shares 'does not affect our results' is asserted without supporting evidence. Since the cross-sectional restriction in Eq. (6) is the main channel through which U.S. GDP information is allocated across states, this sensitivity claim is not self-evident over a sample with large shifts in regional composition; please report the numerical comparison with time-varying weights or qualify the claim.","section":"Footnote 9 and Section 3.2.1"}],"minor_comments":[{"comment":"The statement in Section 3.2.1 that the model-based estimates 'align' with the BEA estimates at observed dates should be rephrased: this alignment is imposed by the temporal aggregation constraints and therefore cannot be read as evidence of in-sample fit.","section":"Section 3.2.1, Figure 1"},{"comment":"The conclusion says the nowcasts are available 'four month ahead of the BEA's first estimates,' while the abstract says three months and Section 3.3 mentions two- and five-month leads for different horizons; please reconcile the timing claims and fix the singular/plural error.","section":"Section 4, Conclusion"},{"comment":"Equations (3)-(5) are presented as exact restrictions on 'exact' growth rates, but the standard Mariano-Murasawa linear form is derived for log-differenced variables; please clarify the sense in which these equations are exact for x_t/x_{t-1}-1 and, if they are approximations, quantify the approximation error.","section":"Section 2.3, Eqs. (3)-(5)"},{"comment":"The state-space notation in Equations (7)-(13) of Appendix A.2 would benefit from a table defining the dimensions of the blocks (N_HF, N_LF, p) and the relationship between the two-step algorithm and the two blocks, to make the implementation reproducible.","section":"Online Appendix A.2, Eqs. (7)-(13)"},{"comment":"The evaluation drops 2020Q2-Q4 because of COVID-19 outliers, but the historical monthly estimates cover the pandemic period; a brief statement of how the smoothed historical estimates are affected by those observations would help readers interpret Figure 1 and the business-cycle results.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be a useful empirical contribution once the Section 2.5 approximation is checked; I do not see a reason to reject, but the historical product's validity is the key issue. I would ask the authors to provide the exact-MCMC comparison and replication code and data before acceptance, especially given the promise to maintain the monthly state GDP series as an ongoing public resource."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. This is the first monthly state-level GDP series for all 50 states plus DC back to 1964, and it is produced by a serious extension of the authors' own MF-VAR framework. The out-of-sample nowcast evaluation is the right test, and the paper mostly passes it. The soft spot is real but localized: the approximate MCMC algorithm in Section 2.5 drives every pre-2005 monthly estimate, and the paper gives no evidence that the approximation is harmless.\n\nWhat's actually new: the product, and the joint cross-sectional constraint that apportions U.S. GDP across states in real time. The temporal and cross-sectional aggregation equations are exact identities. The nowcast design uses real-time vintages, a state-specific benchmark, and RMSFE/CRPS; the gains at m3, once monthly U.S. GDP information is in, are consistent and sensible. The horseshoe prior is a reasonable way to handle a 50-plus equation VAR. I do not see a circularity problem: in-sample consistency with BEA data is imposed by construction, but the paper's claim to test is the out-of-sample nowcast accuracy, which is legitimate.\n\nSoft spots, in rough order. First, the approximation. Equation (8) is exact, but the second factor is replaced by a posterior conditioning on quarterly aggregates of monthly U.S. variables rather than the monthly series. The paper says information loss is 'likely to be small' and notes it only affects pre-2005 quarterly state draws. But those draws feed the first factor and therefore every pre-2005 monthly observation. Because the full 51-state exact sampler is prohibitive, that does not mean the approximation is wrong; it means the paper should validate it on a smaller instance or with a sensitivity analysis. That check is absent, and the historical series is a central output. Second, the lead-time claim is inconsistent across the text: the abstract says three months, Section 3.3 says nowcasts at m1 have a five-month gain, and the conclusion says four months ahead of BEA's first estimates. Should be reconciled. Third, no replication code or data is shipped, even though the authors promise to maintain the series online. For a data-product paper, that is a real gap. Fourth, dropping 2020Q2-Q4 from evaluation and freezing parameters during the pandemic is defensible, but the post-hoc exclusion should be more clearly flagged as a robustness choice.\n\nWho it is for: regional macroeconomists, state policy analysts, and anyone doing business-cycle dating or connectedness across states. It deserves a serious referee. I would send it out, with a request for a small-scale exact-MCMC diagnostic and a cleaner statement of the lead-time results.","headline":"A genuinely useful measurement paper whose pre-2005 historical series rests on an unvalidated MCMC shortcut; referees should ask for a small-scale exact-MCMC check and a clearer statement of the lead-time claims.","tokens_in":30705,"tokens_out":2468,"would_cite":true,"duration_ms":23716,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Monthly GDP for every U.S. state can be estimated directly from official annual and quarterly data, and nowcast months before the official release, using one jointly estimated Bayesian model.","keywords":["Mixed frequency","Monthly GDP","State economies","Regional nowcasting","Vector autoregression","Business cycles","Bayesian analysis","Cross-state heterogeneity"],"falsifier":"Run the exact, non-approximate estimation routine on a three-state version of the same model over a pre-2005 subsample, for example 1990 through 2004, and compare its posterior draws of monthly state GDP growth with those from the approximate two-block algorithm; if the credible intervals for the differences exclude zero, the approximation is not negligible and the historical monthly series is compromised.","tokens_in":29683,"feed_emoji":"📈","tokens_out":8816,"duration_ms":87080,"temperature":0.7,"pith_summary":"This paper tries to establish that a single Bayesian mixed-frequency vector autoregression can turn the official lower-frequency state GDP releases—annual before 2005, quarterly after—plus a set of monthly indicators into a coherent monthly GDP series for all 50 states and Washington, DC, stretching back to the 1960s. The defining feature is that the model is estimated jointly across states and imposes two aggregation constraints: the monthly numbers must temporally average to the official quarterly and annual state figures, and the state figures must cross-sectionally sum to U.S. GDP. If the claim holds, regional researchers and policymakers get a direct measure of state output at monthly frequency, instead of waiting months for official releases or settling for proxy coincident indexes. The paper further claims that, once the latest U.S. GDP figure is in hand, the model produces state GDP nowcasts accurate enough to be useful three months before the official state data arrive, and that the jointly estimated, constrained model beats a state-by-state mixed-frequency VAR that lacks these features.","feed_headline":"Monthly state GDP now arrives months early","feed_subtitle":"One joint model keeps monthly state estimates consistent with official data and beats state-by-state nowcasts.","key_machinery":"The central object is a mixed-frequency vector autoregression written as a state-space model whose latent states are monthly values of variables observed only quarterly or annually, plus a cross-sectional adding-up restriction. Temporal consistency is enforced by exact aggregation identities: quarterly growth is a weighted average of five adjacent monthly growth rates, and annual growth a weighted average of 23 adjacent monthly rates. Cross-sectional consistency is enforced by a separate measurement equation that makes state GDP sum to U.S. GDP, with a small estimated error. The horseshoe prior shrinks the many VAR coefficients equation by equation, and the approximate MCMC algorithm splits the three-way annual–quarterly–monthly mismatch into two two-way blocks, avoiding the computationally heavy annual–monthly restriction while still conditioning the monthly draws on monthly data.","core_discovery":"The central claim is that state-level GDP can be estimated and nowcast at monthly frequency with one large mixed-frequency VAR in which the unobserved monthly state series is the object of interest. The model writes each state's monthly GDP as part of a state-space system where quarterly and annual observations enter through a weighted temporal aggregation rule, and where state GDP adds up to U.S. GDP through a cross-sectional restriction with a stochastic error that absorbs the overseas accounting wedge and vintage differences. Estimation is joint across all 51 units, so shocks can spill across states and the more timely U.S. GDP releases can be apportioned among states rather than ignored. The paper reports that this joint constrained model produces historical monthly estimates that align with official data at the observed low frequencies, and that in real-time evaluation from 2007 to 2024 its nowcasts are more accurate than those from a benchmark that neither allows cross-state dependencies nor imposes the cross-sectional constraint, with the largest improvement appearing at the third month of the quarter.","pith_inferences":["An extension left implicit is that the same two-block approximation should carry over to any panel with a changing frequency mismatch, so the computational device is not tied to U.S. states; a regional output panel in another country with annual-then-quarterly data could reuse it directly.","A natural decomposition experiment would run the out-of-sample exercise four ways—with and without cross-state VAR dynamics, and with and without the cross-sectional constraint—to quantify how much of the third-month gain comes from each ingredient; the paper only compares the full model against the fully restricted benchmark.","Because the cross-sectional weights are held at fixed annual shares, a testable refinement is to let them vary each year or quarter; states with trending output shares would be the place to look for differences, even though the paper reports the fixed-weight choice is not driving its results.","The paper notes weekly estimates are possible but costly; a weekly extension would matter most for pandemic-period tracking, where intra-month movements were large, and the model's latent-state structure is already a natural fit for that higher frequency."],"forward_implications":["State-level business cycles can be dated at monthly frequency from 1964 onward: the paper's median estimates imply Florida and Georgia saw three recessions since 1964, while Iowa, North Dakota, and Alaska saw thirteen or fourteen.","Cross-state spillovers can be measured at a monthly horizon: the paper's variance-decomposition analysis shows many states shift from mostly own-state shocks to strong macro and cross-state spillovers within three months of a shock.","Real-time nowcasts made at the end of the third month of a quarter are accurate enough to replace waiting for the official state release, and the biggest accuracy gain arrives when U.S. GDP for that quarter becomes known.","A joint model with the cross-sectional constraint beats a state-by-state mixed-frequency VAR on average RMSE and CRPS from the third month onward, so conditioning on U.S. GDP and other states' data is doing real work.","The monthly historical series and updated nowcasts are maintained and posted online, so the estimates are intended as a continuing product for regional analysis rather than a one-off exercise."],"supporting_citations":[{"why":"Supplies the baseline mixed-frequency VAR and state-space simulation smoother that the paper extends to a three-way frequency mismatch and cross-sectional constraints.","marker":"Schorfheide and Song (2015)"},{"why":"Provides the inter-temporal aggregation restrictions that link unobserved monthly state GDP to observed quarterly and annual GDP, and the cross-sectional restriction used here.","marker":"Koop et al. (2020b)"},{"why":"Introduces stochastic hierarchical aggregation constraints for regional nowcasting, the direct device for forcing state estimates to agree with U.S. GDP.","marker":"Koop et al. (2024)"},{"why":"Defines the state-level indicators, including wages, employment, hours, and unemployment, that the model borrows to track within-quarter state activity.","marker":"Crone and Clayton-Matthews (2005)"},{"why":"Documents the gap that higher-frequency state indicators fill by showing richer data improve state economic tracking, motivating direct monthly GDP estimates.","marker":"Baumeister et al. (2024)"},{"why":"The horseshoe prior that performs automatic coefficient shrinkage in the large joint VAR, making estimation of the 50-plus equation system feasible.","marker":"Carvalho et al. (2010)"},{"why":"Provides the exact implementation of the horseshoe prior and the Gibbs-sampling steps used in the parameter block.","marker":"Korobilis (2022)"}],"fun_headline_variants":["Monthly state GDP nowcasts, three months before BEA","State GDP monthly estimates arrive a quarter early","First monthly GDP for every state, months ahead","MF-VAR links states for early monthly GDP","Nowcast monthly state GDP with cross-state spillovers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that, for the pre-2005 sample, replacing the monthly U.S. indicators with their quarterly averages in one step of the estimation discards almost no information; if that loss is not small, the historical monthly state GDP series is biased in a way the paper does not measure.","fun_headline_variants_meta":{"raw":{"variants":["Monthly state GDP nowcasts, three months before BEA","State GDP monthly estimates arrive a quarter early","First monthly GDP for every state, months ahead","MF-VAR links states for early monthly GDP","Nowcast monthly state GDP with cross-state spillovers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000617,"raw_usage":{"total_tokens":2823,"prompt_tokens":863,"completion_tokens":1960,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":1887}},"tokens_in":479,"tokens_out":1960,"duration_ms":16101,"temperature":1.0,"reasoning_tokens":1887,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:28:27.472066+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the exact, non-approximate estimation routine on a three-state version of the same model over a pre-2005 subsample, for example 1990 through 2004, and compare its posterior draws of monthly state GDP growth with those from the approximate two-block algorithm; if the credible intervals for the differences exclude zero, the approximation is not negligible and the historical monthly series is compromised.","supporting_citations":[{"cited_title":"Using stochastic hierarchical aggregation constraints to nowcast regional economic aggregates","cited_arxiv_id":null,"evidence_quote":"Introduces stochastic hierarchical aggregation constraints for regional nowcasting, the direct device for forcing state estimates to agree with U.S. GDP."},{"cited_title":"and Alan Clayton-Matthews (2005)","cited_arxiv_id":null,"evidence_quote":"Defines the state-level indicators, including wages, employment, hours, and unemployment, that the model borrows to track within-quarter state activity."}],"review_version":1}