{"id":"dee2ab88-4662-4271-9f22-45ecdf14e0a6","arxiv_id":"2507.22204","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review chapter that surveys and recommends methods for detecting and estimating multiple change points in macroeconomic time series, panel, and factor models.","lead":"This preprint reviews statistical methods for finding sudden changes, called break points, in macroeconomic relationships. It is a practitioner's guide that compares sequential tests, bootstrap methods, information criteria, and penalized regression for detecting multiple breaks.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unit-root negative gap in Section 2 is load-bearing: the review asserts no sequential multiple-break procedure exists under unit roots, but does not bound the claim to predictive regressions; a published sequential I(1)-noise/trend-break procedure would falsify it.","rationale":"This preprint is a handbook review, so its central claim is the accuracy and completeness of its survey, and the strongest practical claim is that the recommended sequential-testing workflow is current best practice. The reader correctly identifies the negative gap claims as the weakest load-bearing component. Among these, the Unit roots paragraph is the most consequential and the least precisely scoped: the sentence denying existing sequential procedures is written generally, and the surrounding discussion only covers single-break bootstrap results in predictive regressions. A published sequential multiple-break procedure for integrated settings would directly falsify the claim and make the guide incomplete for a core macroeconometric case. The proposed literature audit, centered on Kejriwal and Perron (2010), is a single check that settles whether this concern actually lands. Absent that verification, the review appears internally consistent and signals its own limitations in the local-projections discussion about autocorrelated errors; no other internal inconsistency or clear error was found. The verdict therefore remains UNCHANGED: the risk is real but depends on an external literature fact that has not been established in this pass.","tokens_in":25040,"tokens_out":21552,"duration_ms":269646,"concrete_test":"Run a targeted literature audit in EconLit, Web of Science, or Google Scholar for items published before 29 Jul 2025 matching terms such as 'sequential procedure', 'multiple breaks' or 'multiple change-points', and 'unit root' or 'integrated' or 'I(1)'. Specifically verify whether Kejriwal and Perron (2010) or an equivalent paper provides a sequential procedure for estimating the number of breaks in a trend model with integrated noise. If such a paper exists and is not cited or excluded by a scope definition in Section 2, the review's negative gap claim fails and the guide should be revised. If the search returns no in-scope sequential procedure, the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2, Unit roots, states: 'The asymptotic properties of a sequential testing procedure for breaks in the presence of unit roots is to our knowledge not yet available in the literature.' This is one of the negative gap claims on which the guide's completeness rests. It is worded generally, immediately after a discussion confined to zero-vs-one break tests in predictive regressions with a unit-root regressor. If the intended domain is only those predictive-regression settings, the sentence should say so; as written, it appears to deny the existence of any sequential multiple-break procedure with integrated data. The econometrics literature contains sequential procedures for multiple breaks with I(1) noise/trend-break settings, for example Kejriwal and Perron (2010) for trend breaks with an integrated or stationary noise component. If such a procedure falls within the chapter's stated scope—inference on slope-parameter break-points under macro-relevant assumptions—the guide is incomplete on a topic it explicitly flags as important. This is the most load-bearing unverified absence claim because unit-root behavior is a central macroeconometric concern and the guide's usefulness depends on its coverage claims being correct.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a practitioner-oriented survey of methods for detecting and estimating multiple change-points in slope parameters of time-series models, panel data models, factor models, and nonlinear models, with emphasis on macroeconometric applications. It organizes the literature around sequential testing procedures, discusses information criteria, lasso, and mixed-integer programming alternatives, and explicitly flags conjectures and open gaps. The chapter contains no new derivations, simulations, or code.","tokens_in":25268,"tokens_out":11591,"duration_ms":134188,"significance":"If accurate, the survey would be a valuable and much-needed guide for empirical macroeconomists: it is organized around slope-parameter breaks rather than mean shifts, covers endogenous regressors, unit roots, panels, factor models, and recent penalized methods, and it gives concrete practitioner advice such as the Bonferroni correction, bootstrap under the null, and first-stage break checks for 2SLS. It is also transparent in labeling conjectures and open problems. However, because it is a review with no derivations or reproducible code, its accuracy rests entirely on correct citation and scoping of the primary literature; the unit-root gap claim discussed below fails that check as written.","major_comments":[{"comment":"The sentence 'The asymptotic properties of a sequential testing procedure for breaks in the presence of unit roots is to our knowledge not yet available in the literature' is too broad as written. Published work develops sequential procedures for multiple breaks in regressions with integrated variables and/or integrated noise; Kejriwal and Perron (2010, 'Testing for multiple structural changes in cointegrated regression models') is one example, and it falls within the chapter's stated scope of slope-parameter break inference in macro-relevant time-series models. If the intended claim is restricted to the predictive-regression setting with a unit-root regressor and martingale-difference errors discussed immediately above, the sentence should be rewritten to state that restriction and to acknowledge the cointegrated and integrated-noise literature. Because the chapter's usefulness depends on the accuracy of its 'gap' claims, this should be corrected before publication.","section":"Section 2, Unit roots"}],"minor_comments":[{"comment":"The statement that the Bai-Perron dynamic programming algorithm computes the tests in O(T^3) operations 'regardless of the number of candidate breaks' is not supported by Bai and Perron (2003a), which describe an O(mT^2) algorithm for m breaks; please correct or justify the complexity claim.","section":"Section 2, OLS"},{"comment":"The sentence beginning 'Once can also consistently estimate...' contains a typo: 'Once' should be 'One'.","section":"Section 3, first paragraph"},{"comment":"The phrase 'apart form the intercept' should read 'apart from the intercept', and the phrase 'denotes the the ith row' should read 'denotes the ith row'.","section":"Section 5, Factor-augmented forecasting equations"},{"comment":"The reference for Corradi and Swanson (2014) contains the typo 'Testinag' and should read 'Testing'.","section":"References"},{"comment":"The proposed two-step strategy for determining which parameters change is presented without a formal consistency result; if this is a new suggestion, it should be labeled as a conjecture or supported by a citation.","section":"Section 2, Partial structural change"}],"recommendation":"major_revision","confidential_remarks":"This is an invited handbook chapter rather than a primary research article, so the absence of new results is expected. The heavy overlap between the recommended workflows and the authors' own prior work (Boldea, Cornea-Madeira, and Hall 2019; Hall, Han, and Boldea 2012; Boldea, Drepper, and Gan 2020) is natural in a specialist survey, but the editor may wish to ask for an explicit disclosure and for a careful check that alternative methods are not underweighted for this reason. The unit-root gap claim should be corrected before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a solid review chapter, not a research paper, and it should be judged as such. The authors know the literature well and the organization is genuinely useful: focusing on slope-parameter breaks rather than mean shifts, on multiple breaks, and on assumptions relevant to macro applications. The practical recommendations are sensible—use 2SLS plus the bootstrap sequential procedure for endogenous regressors, prefer modified BICs in some settings, be cautious with sup-Wald tests under heteroskedasticity. I found several places where the chapter flags its own conjectures clearly, which is honest and helpful for readers.\n\nThe main soft spot is a real one, and the stress-test note lands. In Section 2 (Unit roots), the sentence \"The asymptotic properties of a sequential testing procedure for breaks in the presence of unit roots is to our knowledge not yet available in the literature\" is overbroad as written. Kejriwal and Perron (2010) provide a sequential procedure for multiple trend breaks with integrated or stationary noise. If the authors mean sequential slope-break testing in predictive regressions with a unit-root regressor, they should say exactly that. As written, the claim is a negative gap statement on which the guide's completeness partly rests, and it is falsifiable. This needs a scope correction, not a major rewrite.\n\nA smaller issue: the paper says Bai and Perron's dynamic programming runs in O(T^3) operations regardless of the number of candidate breaks. The standard statement is O(mT^2) for m breaks. This is a minor technical slip, but in a practitioner guide it matters because bootstrapping feasibility is part of the sales pitch.\n\nThe self-citation pattern is not a problem here. The cited work by Boldea, Cornea-Madeira, and Hall and by Hall and coauthors is directly relevant, and the chapter also cites competitors like Perron-Yamamoto, Qian-Su, and Baltagi-Kao-Wang. The review is not a vanity exercise. It also does not pretend to offer new theorems or simulations; it is explicitly a handbook survey, so lack of novelty is not a flaw.\n\nWho is this for? Applied macroeconomists who want a map of available change-point tests, and graduate students entering the area. It deserves a serious referee, mainly to check the scope of the negative claims and the computational assertion. I would send it out, expect the authors to fix the unit-root sentence and the complexity statement, and then publish it. It will be a useful reference even if it does not break new ground.","headline":"A careful, genuinely useful handbook review of multiple change-point methods for macroeconomists, with no new results but a few overbroad 'gap' claims that need tightening.","tokens_in":25756,"tokens_out":2857,"would_cite":true,"duration_ms":36083,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims that a sequential testing workflow can reliably detect and estimate multiple change-points in the slope parameters of the main macroeconometric model classes.","keywords":["change-points","break-points","time series","panel data","factor models","sequential testing","information criteria","lasso"],"falsifier":"A targeted literature search that surfaces a published sequential testing procedure for multiple breaks in the presence of unit roots, or a sequential multiple-break test for heterogeneous panels, would directly falsify the review's central gap claims; a simulation showing the recommended two-stage least squares bootstrap breaks down under autocorrelated errors would test the paper's own boundary.","tokens_in":24834,"feed_emoji":"📈","tokens_out":8599,"duration_ms":83151,"temperature":0.7,"pith_summary":"This review argues that a practitioner can reliably detect and estimate multiple change-points in macroeconomic models by following a sequential testing workflow centered on sup-F and sup-Wald statistics. It covers univariate regressions with exogenous or endogenous regressors, multivariate and high-dimensional time series, panel data models, and factor models, and focuses on changes in slope parameters — the coefficients on explanatory variables — rather than on changes in the mean of the dependent variable. The review recommends using two-stage least squares with a wild bootstrap sequential procedure to determine the number and location of breaks when regressors are endogenous, and notes that once break locations are fixed, generalized method of moments can estimate the parameters as if the locations were known. It also identifies settings where methods do not yet exist, such as multiple breaks under unit roots, sequential multiple-break inference in heterogeneous panels, and bootstrap-valid tests with autocorrelated errors. Getting this right matters because undetected parameter change distorts policy recommendations and forecasts.","feed_headline":"Sequential testing locates multiple breaks in macro models","feed_subtitle":"Macro relationships shift; the recommended bootstrap workflow tells you when and by how much.","key_machinery":"The load-bearing object is the sequential testing strategy for slope-parameter change-points. Its three components are tests of zero breaks against a fixed number $k$, against an unknown number up to a maximum $M$ (the double-maximum tests), and of $\\ell$ versus $\\ell + 1$ breaks; all are built from sup-F or sup-Wald statistics maximized over candidate partitions. A dynamic programming algorithm computes the objective in $O(T^3)$ operations regardless of the number of candidate breaks, which makes bootstrapping the full procedure feasible. The bootstrap, applied by resampling residuals under the null and either keeping regressors fixed or recursively reconstructing lagged dependent variables, is the device that extends validity to the non-covariance-stationary environments common in macro data. In time series, the procedure consistently estimates the break fractions $\\tau_j^0$ but not the break dates themselves; in panels and factor models, the break-point $T_j^0$ can be consistently estimated.","core_discovery":"The central claim is that the literature now supplies a coherent, assumption-tailored toolkit for testing and estimating multiple discrete change-points in the slope parameters of the linear models macroeconomists use most. In univariate time series, the tools are the three test families for no breaks versus a known number, no breaks versus an unknown number up to a bound, and $\\ell$ versus $\\ell+1$ breaks, together with a dynamic programming algorithm that makes the search fast and a wild bootstrap that extends validity to data whose second moments change over time. With endogenous regressors, the same tests are built on two-stage least squares estimates after checking the first stage, and the review recommends this bootstrap two-stage least squares procedure because generalized method of moments based break detection has not been theoretically justified. In panels and factor models, extensions of the same logic work on transformed data or on the second moments of estimated factors, and these deliver consistent estimators of the break dates themselves, whereas time-series models only deliver consistent break fractions. The review also claims that modified information criteria, adaptive lasso, and mixed-integer programming are useful complements for breaks that are close together or near the sample edge.","pith_inferences":["The review's conjecture that a moving block bootstrap would make the sequential tests valid under autocorrelated errors, if proved, would fill a conspicuous gap and make the workflow directly applicable to multi-step local projection regressions.","If the claimed gaps are real, the most valuable near-term targets are a sequential multiple-break procedure for unit-root regressors and one for heterogeneous panels; either would extend the guide's scope.","The paper's comparison suggests an empirical extension: applied studies could benchmark the sequential workflow against modified information criteria and lasso methods on the same macroeconomic datasets to see which detects adjacent or edge breaks more often.","Because factor-model tests operate on second moments of estimated factors, one could adapt them to detect joint breaks in loadings and in factor-augmented forecasting equations, the setting the single-break test in the review addresses."],"forward_implications":["Use the two-stage least squares sequential bootstrap procedure to detect breaks when regressors are endogenous; only after the break locations are fixed should generalized method of moments be used to estimate the slope parameters.","Always use bootstrap critical values, even when asymptotic critical values exist, because finite-sample behavior is better in simulations.","Modified BIC-type information criteria, adaptive lasso, and mixed-integer programming can detect breaks that sequential tests miss, in particular breaks close together or near the sample edges.","In homogeneous panels with interactive fixed effects, sequentially testing on common-correlated-effects transformed data consistently estimates the number and location of breaks, and a dedicated software package is available.","In factor models, tests based on the second moments of estimated factors can detect common breaks in loadings, and the likelihood-ratio version diverges faster under the alternative than the Wald version."],"supporting_citations":[{"why":"Supplies the sup-F and sup-Wald test families, the sequential testing logic, and the dynamic programming algorithm that the whole guide builds on.","marker":"Bai and Perron (1998)"},{"why":"Establishes the sup-Wald approach to unknown break-points and its asymptotic distribution for models with heteroskedasticity and autocorrelation.","marker":"Andrews (1993)"},{"why":"Extends the sequential testing strategy to models with endogenous regressors estimated by two-stage least squares.","marker":"Hall, Han, and Boldea (2012)"},{"why":"Proves validity of the wild bootstrap for the sequential sup-F and sup-Wald tests, the method the review recommends for practice.","marker":"Boldea, Cornea-Madeira, and Hall (2019)"},{"why":"Generalizes the asymptotic theory to mixingale regressor-error products, broadening coverage to many macro processes.","marker":"Perron and Qu (2006)"},{"why":"Provides the fixed-regressor wild bootstrap for change-point tests when covariance stationarity fails, an ingredient for the bootstrap approach.","marker":"Hansen (2000)"},{"why":"Extends the sequential testing strategy to multiple breaks in interactive-effects panel models with common-correlated-effects transformation.","marker":"Ditzen, Karavias, and Westerlund (2024)"},{"why":"Develops multiple-break tests for factor models based on second moments of estimated factors, with limiting distributions from the linear regression case.","marker":"Baltagi, Kao, and Wang (2021)"},{"why":"Provides quasi-likelihood estimation and testing for multiple breaks in multivariate regressions with common breaks across equations.","marker":"Qu and Perron (2007)"},{"why":"Provides single common-break tests in factor models based on sub-sample means of estimated factor second moments.","marker":"Han and Inoue (2015)"}],"fun_headline_variants":["Sequential tests find multiple breaks in macro slope parameters","Bootstrap method sharpens multiple change-point detection in macro","New guide to multiple break testing in macroeconometrics","Consistent break-date estimators for panels and factor models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The guide's usefulness stands or falls on whether the gaps it identifies are really gaps: if any published method already handles sequential multiple-break testing under unit roots, in heterogeneous panels, or with generalized method of moments, the recommended workflows and scope statements would be incomplete.","fun_headline_variants_meta":{"raw":{"variants":["Sequential tests find multiple breaks in macro slope parameters","Bootstrap method sharpens multiple change-point detection in macro","New guide to multiple break testing in macroeconometrics","Consistent break-date estimators for panels and factor models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000342,"raw_usage":{"total_tokens":1856,"prompt_tokens":895,"completion_tokens":961,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":897}},"tokens_in":511,"tokens_out":961,"duration_ms":10400,"temperature":1.0,"reasoning_tokens":897,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:58:14.628831+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A targeted literature search that surfaces a published sequential testing procedure for multiple breaks in the presence of unit roots, or a sequential multiple-break test for heterogeneous panels, would directly falsify the review's central gap claims; a simulation showing the recommended two-stage least squares bootstrap breaks down under autocorrelated errors would test the paper's own boundary.","supporting_citations":[],"review_version":1}