{"id":"7431bf9b-9df7-4ada-aaac-1b48f5cdd46e","arxiv_id":"2505.12422","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"LP impulse response estimates are decomposed into time-stamped contributions, revealing that many influential estimates are concentrated in a few historical episodes.","lead":"This paper decomposes local projection estimates of impulse responses into a sum of historical event contributions, with weights interpretable as purified shocks or proximity scores. The tool reveals that many influential macro estimates are driven by a handful of episodes, such as World War II, 1970s monetary loosening, and the 1964 Mount Agung eruption.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The empirical 'dominant historical episode' claims rest on point-estimate concentration statistics and contribution rankings that are never given sampling uncertainty; the paper's own robustness checks show how easily top contributors move, so the headline findings need a stability check.","rationale":"The reader's weakest_assumption correctly identifies the identification conditions for the narrative reading and the RF always.split.variables regularization. My concern is adjacent but distinct: even granting the identifying assumptions, the empirical headline depends on point estimates of concentration and top-episode rankings that are never accompanied by sampling uncertainty or a stability analysis. The paper does include substantive robustness checks for fiscal and climate applications, which is credit to the authors, but those checks are deterministic sample modifications rather than a characterization of how much the top-episode attribution itself is subject to sampling variation. Because the decomposition w is a nonlinear functional of the full sample, a small perturbation can reorder contributions; the dramatic collapse in Figure 6 and the sample-dependence in Figure 9 illustrate this. This does not undermine the core methodological contribution, which is an algebraic identity with useful interpretive lenses, nor does it require rejection. It does suggest that the empirical 'dominant drivers' statements should be framed as sample-specific diagnostics unless they survive a formal stability check. I therefore keep the reader's ACCEPT verdict unchanged, while adding a concrete test that would settle whether the concern is material.","tokens_in":35611,"tokens_out":20015,"duration_ms":215270,"concrete_test":"Run a delete-one jackknife and a moving-block bootstrap on two flagship specifications: the R&R monetary policy inflation IRF at h=48 and the RZ recession-regime fiscal IRF at h=12. For each replication, recompute w, c_th, WC, CC, and the identity/rank of the top three contributing episodes. Report the frequency with which each episode (e.g., 1976-1978 loosening, WWII) appears in the top three, and bootstrap confidence intervals for WC and CC. If the top episodes are stable across replications and the WC/CC intervals are tight and far from low-concentration benchmarks, the concern is resolved; if the top contributor changes frequently or the intervals are wide enough to include diffuse distributions, the empirical claims should be softened to sample-specific observations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central decomposition is algebraically correct and the proximity/dual interpretations check out. The load-bearing concern is in the empirical translation: the paper identifies 'dominant drivers' such as Nixon's pressure, WWII, and Mount Agung from point estimates of w and c_th, and reports no uncertainty about which episodes are top contributors or about the concentration statistics WC and CC. This matters because w is the second row of (X'X)^{-1}X', a nonlinear function of the whole sample; a single influential observation can change the top-contribution ranking. The paper's own adversarial checks make this concrete: trimming the top and bottom 1% of fiscal weights collapses both state-dependent IRFs to zero (Figure 6), and starting the climate sample in 1985 eliminates the long-run effect attributed to Mount Agung (Figure 9). Thus the headline claim that easily identifiable events are dominant drivers is not yet shielded from sampling and specification variability. The causal reading of those contributions also inherits the LP identifying assumptions stated in footnote 2; if exogeneity, control sufficiency, or stability fails, the attribution to Nixon/WWII/Agung is not identified, even though the decomposition remains a valid descriptive diagnostic. The concern is not that the algebra is wrong, but that the empirical conclusions are stated more strongly than the accompanying inference supports.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces a decomposition of local projection (LP) estimates of impulse responses into a sum over time of contributions c_th = w_t y_{t+h}, where w is the second row of (X'X)^{-1}X' (Eqs. 3-4). For OLS, the paper gives two interpretations of w: as standardized, purified shocks via the Frisch-Waugh-Lovell theorem (Eqs. 6-8), and as proximity scores between a hypothetical intervention and past interventions in an orthonormalized regressor space (Eqs. 9-15). The proximity interpretation is extended to random forest LPs by expressing predictions as weighted averages of outcomes. Two concentration statistics, WC and CC, quantify the share of top-Q% weights and contributions. The framework is applied to monetary, fiscal, climate, and financial shocks; the authors report that estimates are often concentrated in a small number of identifiable episodes such as Nixon-era monetary loosening, WWII, and the 1963 Mount Agung eruption.","tokens_in":1605,"tokens_out":1768,"duration_ms":71483,"significance":"If the empirical claims can be supported with inference, this is a useful and original diagnostic. The OLS decomposition is an exact, parameter-free identity, and the FWL interpretation is rigorous. The proximity interpretation provides an intuitive bridge to machine-learning LPs, and the RF weight recovery is clearly described. The paper ships R code and slides, and the fiscal and climate applications include transparent robustness checks that, if anything, undercut the headline episode attribution. The main value is a diagnostic that quantifies how many historical observations actually support an IRF, which is a real gap in the LP toolkit.","major_comments":[{"comment":"The headline concentration statistics (WC, CC) and the top-contribution episode rankings are reported as point estimates with no measure of sampling uncertainty. Because w is a nonlinear function of the full sample through (X'X)^{-1}, a single influential observation can change the top-contribution ranking, and the paper's own adversarial checks demonstrate exactly this fragility: trimming the top and bottom 1% of fiscal weights collapses both state-dependent IRFs to zero (Figure 6), and starting the climate sample in 1985 eliminates the long-run effect attributed to Mount Agung (Figure 9). Please add bootstrap or wild bootstrap confidence intervals, or subsample stability analyses, for WC, CC, and the identity of top contributors, and calibrate the claims in the Abstract and Section 4 to what survives that inference.","section":"Section 3; Tables 1-4"},{"comment":"The narrative attribution of contributions to events such as 'Nixon's interference with the Fed' or 'stagflation' is a causal reading of the decomposition; it inherits the identifying assumptions of the LP listed in footnote 2 (exogeneity, correct controls, coefficient stability). Footnote 2 correctly states that the decomposition itself is mechanical, but the Introduction and Abstract do not carry that caveat. Please either soften the episode-attribution language to descriptive weighting, or provide evidence that the attributions are stable under alternative control sets and placebo shock series with the same autocorrelation structure. A placebo exercise would be a concrete way to show that top contributions are not generated by a null shock process.","section":"Section 3.1; footnote 2"},{"comment":"The random forest design forces shock variables into every split via always.split.variables, and the paper states that without this adjustment the shock series would be effectively ignored by the RF. The nonlinear findings, especially the null contractionary monetary response and the sparse proximity weights, may therefore be artifacts of this selective regularization. Sensitivity to min.node.size is reported, but sensitivity to mtry and to always.split.variables is not. Please report results for a grid of these tuning parameters, or a data-driven choice such as cross-validation, and show the corresponding weights and IRFs so the reader can assess how much of the nonlinear concentration is model-induced.","section":"Appendix A.2.1; Section 3.1 nonlinear results"},{"comment":"The proximity interpretation of w in the correlated case uses the scenario vectors X^delta_tau and X^0_tau, but the empirical proximity plots do not state how Z_tau is chosen. Since beta_h is invariant to Z_tau in the linear model while w's decomposition into cosine and norm components in (12)-(15) is not, the 'proximity' narrative depends on an arbitrary choice of the scenario. Please specify the rule used for Z_tau in Figures 2 and 14 and check the sensitivity of the reported proximity patterns to alternative scenario definitions, such as the sample mean, the last observation, or an explicit policy counterfactual.","section":"Section 2.2.2; Figures 2 and 14"}],"minor_comments":[{"comment":"The notation dIRF with a hat is ambiguous; suggest writing the estimator as \\widehat{IRF}_{s\\to y}(h,1) = \\hat\\beta_h and defining w as a row vector consistently.","section":"Eq. (3)"},{"comment":"Var(\\tilde s_t) in (7) should be the sample variance of the residual series to be numerically exact; please state this explicitly.","section":"Eqs. (6)-(8)"},{"comment":"The approximation w_t \\hat\\nu_{t+h} \\approx IF_t is informal; define the approximation error and state that it holds exactly only under the i.i.d. cross-sectional conditions described in the text.","section":"Eq. (24)"},{"comment":"The repeated quotation marks in Table 1 should be replaced with actual numbers or an explicit statement that WC is horizon-invariant for linear models.","section":"Table 1"},{"comment":"The caption refers to 'Figure ??'; the cross-reference should be fixed to Figure 10.","section":"Figure 23 caption"},{"comment":"The in-text citation 'Zeev et al. (2023)' should be expanded to 'Ben Zeev, Ramey, and Zubairy (2023)' to match the reference list.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of econ.EM and the decomposition is a genuine contribution. The main risk is that the empirical 'dominant episode' findings are published without inference on rankings and concentration; I recommend requiring the additional evidence described in major comments 1-3. The reliance on an SSRN working paper (Goulet Coulombe 2025) for the proximity interpretation is acceptable, but the editor may want to verify that the key result is publicly accessible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The paper's core decomposition is elementary—beta-hat_h equals the sum of w_t y_{t+h} with w_t the second row of (X'X)^{-1}X'—and the authors turn that identity into two useful readings: purified shocks via FWL, and proximity scores via the dual solution. That coefficient-level extension of Goulet Coulombe (2025) plus the WC/CC concentration statistics are the genuine novelties, and they are worth having as standard diagnostics. The empirical applications are honest, transparent, and carefully hedge causal readings.\n\nThe soft spot is where the stress-test note lands. The dominant-episode claims—Nixon's pressure, WWII, Mount Agung—are based on point-estimate weights and contribution rankings with no sampling uncertainty attached. The concentration statistics are also point estimates. That matters because w is a nonlinear function of the whole sample, and the paper's own robustness checks show how easily the picture changes: trimming the top and bottom 1% of fiscal weights kills the state-dependent IRFs, and moving the climate sample to 1985 eliminates the long-run damage attributed to Agung. The decomposition holds mechanically, but the empirical headline that these events drive the result is not yet shielded from sampling or specification variability.\n\nTwo smaller concerns. The Random Forest results force shock variables into every split via always.split.variables; this is disclosed and defended, but it is a consequential regularization, and the finding that contractionary R&R shocks have no effect rides on it. Also, the causal reading of episode contributions inherits the usual LP identifying assumptions, which the authors acknowledge in footnote 2. Neither is fatal—the diagnostic value is real—but the conclusions should be softened until there is some sense of how stable the top-contribution rankings are.\n\nBottom line: a solid paper that deserves a serious referee. The right revision would add uncertainty quantification for the concentration statistics and rankings, or reframe the empirical claims as suggestive. I would bring it to reading group and likely cite it once the diagnostics become standard.","headline":"A useful and mostly sound diagnostic framework for local projections, but the headline claims about dominant historical episodes need uncertainty quantification before they can be taken at face value.","tokens_in":36372,"tokens_out":2734,"would_cite":true,"duration_ms":26855,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that every local projection estimate decomposes into weighted contributions from individual historical episodes, with weights interpretable as purified shocks or proximity scores.","keywords":["local projections","impulse response functions","historical decomposition","proximity weights","Frisch-Waugh-Lovell","Random Forest","external validity","macroeconomic shocks"],"falsifier":"For each of the four applications, remove the single highest-contribution episode identified by the decomposition: World War II for the recession fiscal multiplier, the 1964 Mount Agung shock for climate, the 1976-1978 loosening sequence and Nixon episodes for monetary policy, and the 2007-2008 financial distress cluster for financial shocks, then re-estimate; the paper's central concentration claim predicts each estimate should largely collapse, and verifying this uniformly would settle how much of each result depends on one event.","tokens_in":35388,"feed_emoji":"🔍","tokens_out":9043,"duration_ms":83533,"temperature":0.7,"pith_summary":"The paper aims to open the black box of local projections, the workhorse regression method for estimating impulse responses, by decomposing any estimate into a sum of contributions from historical periods. Each period contributes the product of a weight and the observed future outcome, and the weights admit two readings: as purified, standardized shocks, and as proximity scores between the policy intervention being studied and past interventions. Because many machine-learning predictors are also linear combinations of outcomes, the decomposition applies to nonlinear local projections as well. In monetary, fiscal, climate, and financial applications, the paper finds that estimates are often heavily concentrated in a few recognizable episodes, with direct consequences for how much external validity the estimates have.","feed_headline":"Most local projections are driven by a handful of events","feed_subtitle":"A new decomposition names the shocks—Nixon, WWII, Mount Agung—behind influential macroeconomic results.","key_machinery":"The central object is the weight vector $w = [(X'X)^{-1}X']_{2,:}$, the row of the OLS projection matrix that selects the shock coefficient. The Frisch-Waugh-Lovell theorem turns this into $w_t = s_t^*/T$, giving the weights a purified-shock meaning. The dual solution of least squares turns it into a proximity score: in the whitened feature space $F_t = X_t U \\Lambda^{-1/2}$, each $w_t$ is the inner product $\\langle F^\\delta_\\tau - F^0_\\tau, F_t\\rangle$, so the coefficient is a similarity-weighted average of past outcomes. In Random Forests, weights are recovered by averaging, across trees, the indicator that observation $t$ falls in the same leaf as the hypothetical intervention; the nonlinear impulse response is the difference between two such weight vectors, one for shock $\\delta$ and one for no shock.","core_discovery":"This paper's central claim is that any local projection coefficient is a weighted sum of historical outcomes: $\\hat{\\beta}_h = \\sum_{t=1}^{T} w_t y_{t+h}$, where $w$ is the second row of $(X'X)^{-1}X'$ and $y_{t+h}$ is the outcome realized $h$ periods after observation $t$. The paper then gives the weights two interpretations. By the Frisch-Waugh-Lovell theorem, $w_t = s_t^*/T$, where $s_t^*$ is the shock $s_t$ residualized on the controls and standardized, so each contribution is a purified shock times the realized outcome. Through the dual form of least squares, $w_t$ is instead the difference between the inner products of a hypothetical intervention and each historical intervention in an orthonormal feature space, which makes OLS an estimator that upweights past episodes resembling the intervention being studied. Because Random Forest predictions and many other machine-learning predictors are linear combinations of the outcome, the same decomposition carries over to nonlinear local projections, where weights differ by horizon, shock sign, and context.","pith_inferences":["A natural extension would make the decomposition a pre-registered robustness check: report $WC$, $CC$, and the dominant episode for every LP estimate, and treat estimates whose top episode accounts for most of the coefficient with caution.","The same machinery could be applied to panel local projections to separate time-period from cross-sectional contributions, and to two-stage least squares by decomposing the first and second stages separately.","If concentration is as high as the applications suggest, then nonlinear and state-dependent effects estimated from the same short samples inherit the same fragility; the paper's trimming exercises indicate that removing one episode often collapses the estimate, a test that could be standardized for all four applications.","Because the decomposition is mechanical, it does not by itself validate the causal reading; it equips a substantive identification argument with a precise picture of where the evidence lives."],"forward_implications":["Each LP estimate can be plotted as an evidence curve, with cumulative contributions over time showing whether support is broad or concentrated in one episode.","The concentration statistics $WC$ and $CC$, which report the share of absolute weights and contributions held by the top 10% of observations, give a one-number diagnostic for external validity.","The framework explains the monetary price puzzle as misattribution of 1970s stagflation to tightening, and shows that narrative monetary shocks identify effects almost entirely from politically motivated loosening in the 1970s.","State-dependent fiscal multipliers estimated from military spending shocks are almost entirely driven by World War II, and long-run climate damage estimates are heavily influenced by the 1964 Mount Agung eruption paired with the post-war boom, making those long-run results fragile.","Nonlinear Random Forest local projections can be read with the same tools, revealing that contractionary narrative monetary shocks yield null effects while expansionary effects are concentrated in identifiable episodes such as Nixon's pressure on the Federal Reserve."],"supporting_citations":[{"why":"Supplies the standard result that OLS coefficients are weighted averages of the outcome, the starting point for the decomposition.","marker":"Davidson and MacKinnon (2004)"},{"why":"Provides the similarity-based, dual view of OLS predictions that the paper extends from predictions to coefficients.","marker":"Goulet Coulombe (2025)"},{"why":"Establishes that Random Forest predictions are convex combinations of outcomes, enabling weight retrieval.","marker":"Lin and Jeon (2006)"},{"why":"Supplies the narrative monetary policy shock series whose historical concentration the paper analyzes.","marker":"Romer and Romer (2004)"},{"why":"Is the temperature-shock local projection whose long-run GDP estimates the paper re-examines and finds fragile.","marker":"Bilal and Känzig (2024)"},{"why":"Provides the state-dependent fiscal multiplier estimates shown to be driven by World War II and the Korean War.","marker":"Ramey and Zubairy (2018)"},{"why":"Supplies the excess bond premium shock series used in the financial application.","marker":"Gilchrist and Zakrajšek (2012)"},{"why":"Motivates the expanded information set that corrects the price puzzle in the paper's monetary application.","marker":"Bernanke et al. (2005)"},{"why":"Frames the nonlinearity caveat and alternative weighting perspective the paper positions itself against.","marker":"Kolesár and Plagborg-Møller (2024)"},{"why":"Gives the dual interpretability machinery for machine-learning forecasts that the paper adapts to impulse response coefficients.","marker":"Goulet Coulombe et al. (2024)"}],"fun_headline_variants":["Local projections are just weighted sums of history","How a few events dominate local projection estimates","New decomposition exposes key events in local projections","From Nixon to Mount Agung: what drives local projections","Local projections: a closer look at the contributing shocks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The narrative reading of the decomposition, that specific episodes like Nixon's pressure or World War II are what drive an estimate, presupposes the standard local projection assumptions of exogenous shocks, correctly specified controls, and stable coefficients; if those fail, the weights still decompose the OLS number mechanically, but the historical attribution is not identified.","fun_headline_variants_meta":{"raw":{"variants":["Local projections are just weighted sums of history","How a few events dominate local projection estimates","New decomposition exposes key events in local projections","From Nixon to Mount Agung: what drives local projections","Local projections: a closer look at the contributing shocks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000241,"raw_usage":{"total_tokens":1532,"prompt_tokens":964,"completion_tokens":568,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":497}},"tokens_in":580,"tokens_out":568,"duration_ms":5940,"temperature":1.0,"reasoning_tokens":497,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:34:13.410170+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For each of the four applications, remove the single highest-contribution episode identified by the decomposition: World War II for the recession fiscal multiplier, the 1964 Mount Agung shock for climate, the 1976-1978 loosening sequence and Nixon episodes for monetary policy, and the 2007-2008 financial distress cluster for financial shocks, then re-estimate; the paper's central concentration claim predicts each estimate should largely collapse, and verifying this uniformly would settle how much of each result depends on one event.","supporting_citations":[{"cited_title":"and MacKinnon, J","cited_arxiv_id":null,"evidence_quote":"Supplies the standard result that OLS coefficients are weighted averages of the outcome, the starting point for the decomposition."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the similarity-based, dual view of OLS predictions that the paper extends from predictions to coefficients."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the narrative monetary policy shock series whose historical concentration the paper analyzes."},{"cited_title":"and K \\\"a nzig, D","cited_arxiv_id":null,"evidence_quote":"Is the temperature-shock local projection whose long-run GDP estimates the paper re-examines and finds fragile."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the state-dependent fiscal multiplier estimates shown to be driven by World War II and the Korean War."},{"cited_title":"and Zakraj s ek, E","cited_arxiv_id":null,"evidence_quote":"Supplies the excess bond premium shock series used in the financial application."}],"review_version":1}