{"id":"feb37dfb-953c-471e-bab1-431fd71422e9","arxiv_id":"2505.04423","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"Averaging GNAR forecasts across the top five random graphs selected by recent one-step-ahead errors beats AR benchmarks at all horizons and beats the Bank of England at 4-6 months, though the Bank comparison lacks significance testing.","lead":"This paper tests whether forecasting UK inflation with network models built on random graphs can beat simple benchmarks and even the Bank of England. The best version, which averages several randomly selected network models, wins at medium horizons but the headline advantage over the Bank rests on a short sample without statistical tests.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"BoE outperformance claim depends on unverified target and horizon alignment; Table 5 may compare monthly RaGNAR forecasts with quarterly BoE MPR forecasts.","rationale":"The reader's conditional verdict is on point. The paper is otherwise a well-executed out-of-sample application: walk-forward selection on past RMSE, averaging over the top five networks, reproducible code and data, and consistent improvement over simple benchmarks. But the abstract's strongest claim, beating the Bank of England, depends entirely on Table 5, whose comparability is not established. This is not a disagreement with the BoE being hard to beat; it is a correctness risk about the comparison. The concern is directly testable with the supplied data and code, so I recommend no change to the reader's CONDITIONAL verdict: accept only if the BoE forecast object, horizon mapping, and information set are shown to match, and preferably with a formal equal-predictive-accuracy test.","tokens_in":35481,"tokens_out":13699,"duration_ms":137561,"concrete_test":"Use the repository's data and code (github.com/henrypalasciano/Forecasting-Inflation-with-RaGNAR) to extract the exact BoE series used and verify whether the BoE MPR values are quarterly average CPI forecasts or monthly point forecasts. If they are quarterly, recompute Table 5 with RaGNAR forecasts aggregated to the same quarterly target on the same release dates and data vintage. Then run a Diebold-Mariano test with HAC/Newey-West standard errors on the h=4, 5 and 6 forecast error series for the best AvGNAR model versus the BoE. If BoE RMSEs change materially, or if the Diebold-Mariano p-values exceed 0.05, the headline outperformance claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.3 (Table 5) is the only support for the abstract's headline claim that RaGNAR is materially more accurate than the Bank of England's 4-6 month forecasts. The paper never states what the BoE forecast object is: the MPR publishes quarterly CPI inflation projections, yet Table 5 reports BoE RMSEs for horizons 1-6 months and compares them with RaGNAR's monthly year-on-year forecasts. If the BoE numbers are quarterly average CPI forecasts, or are interpolated from quarterly fan-chart modes, then the reported BoE RMSEs of 0.85, 0.93 and 0.92 at h=4, 5 and 6 are not measuring the same target as RaGNAR's monthly forecasts, and the apparent RaGNAR advantage (e.g., global-alpha AvGNAR(P2,S1) h=4 RMSE 0.30 in Table 5) is not a valid comparison. The paper also does not align data vintages: BoE forecasts are made in real time, whereas RaGNAR uses the full ONS dataset as of the latest release. Finally, the evaluation window since end-2019 is short, and the reported plus-or-minus uncertainties are over random graph draws, not over time; no equal-predictive-accuracy test supports the phrase 'materially more accurate.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces RaGNAR, a forecasting procedure that fits Generalised Network Autoregressive (GNAR) models to many random Erdős–Rényi–Gilbert graphs defined on 114 UK CPI component series. Each month the graphs are ranked by rolling one-step-ahead RMSE at the CPI node, the best networks are retained, and forecasts up to 12 months ahead are produced either from a single best network or by averaging over networks and model orders. The empirical sections report RMSE and MAPE performance over 2010–2024 against AR, random-walk, averaged-AR, and Chronos benchmarks, and compare 4–6 month forecasts with Bank of England forecasts from the quarterly Monetary Policy Reports. The paper also proves two propositions about neighbour-set size distributions and interprets the most frequent edges in the best networks.","tokens_in":35756,"tokens_out":7286,"duration_ms":74635,"significance":"If the results are valid, the paper makes a useful practical contribution: it provides a very fast, parsimonious, reproducible forecasting method that appears to beat simple univariate benchmarks at medium horizons, and it offers a constructive way to exploit disaggregated CPI information through random networks. Strengths include a genuinely out-of-sample network-selection protocol, transparent reporting of Monte Carlo variability over random graph draws, proofs of the random-graph propositions, and a public code repository. The main unresolved issue is whether the headline comparison with the Bank of England measures the same forecast target and horizon; currently that claim is not sufficiently supported.","major_comments":[{"comment":"The Bank of England comparison is not an apples-to-apples comparison as presented. The paper never states whether the BoE numbers are month-specific year-on-year CPI forecasts, quarterly average CPI forecasts, or projections of some other object; it also does not define forecast origins or data vintages. Since BoE Monetary Policy Reports publish quarterly projections, the monthly horizons 4, 5, and 6 in Table 5 may be compared with a different target than RaGNAR's monthly year-on-year forecasts. Please specify the exact BoE forecast object and align horizons, target definitions, and information sets, or restrict the abstract claim to a properly aligned subsample.","section":"Section 5.3, Tables 5 and 7"},{"comment":"The claim that RaGNAR is 'materially more accurate' than the Bank of England has no inferential support. The evaluation window since end-2019 contains only a small number of quarterly observations, and the reported ±1 standard deviations are Monte Carlo variation across the 100 random graph draws, not sampling uncertainty of forecast-error differences. Please add a small-sample equal-predictive-accuracy test (e.g., Diebold–Mariano or a bootstrap over the dated forecast errors) and state the number of observations used at each horizon.","section":"Section 5.3"},{"comment":"The MAPE definition in Eq. (15) uses a nonstandard denominator |X_t| + 1. Because the reported MAPE comparisons with the Bank of England in Table 7 depend on this modified metric, differences may partly reflect the offset rather than forecast accuracy. Please justify the modification and report results under the conventional MAPE definition, or explicitly discuss the sensitivity of the conclusions to the denominator choice.","section":"Section 5.4, Eq. (15)"}],"minor_comments":[{"comment":"There is a typo in the abstract: 'Bank of Englan's' should be 'Bank of England's'.","section":"Abstract"},{"comment":"Tables 5 and 7 should state the exact dates and number of BoE observations available at each horizon, and clarify how the BoE series was constructed from the Monetary Policy Reports, including whether modes, means, or fan-chart ranges were used.","section":"Section 5.3"},{"comment":"The PACF windows used to motivate the P1/P2 order sets include 2005–2024, which overlaps the 2010–2024 evaluation period; please clarify whether the averaging sets were chosen before the evaluation window or discuss this as a limitation of the benchmark construction.","section":"Appendix B"},{"comment":"The labels 'liquid fuels' and 'fuels & lubricants' in Figures 6–8 may be confusing without ONS series codes; adding the series identifiers would improve reproducibility.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The core methodology and benchmark comparison are likely salvageable and could be publishable on their own. The BoE comparison in its current form should not be the basis of the abstract's headline claim. I would recommend that the authors either provide a rigorous alignment and significance test for the BoE comparison, or remove the BoE claim from the abstract and present the BoE comparison as exploratory. A major revision that clarifies or narrows the claim seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things. First, this paper is a competent, honest application of random-graph GNAR models to UK CPI forecasting, with a genuinely out-of-sample design and reproducible code. Second, its headline claim—beating the Bank of England at 4–6 month horizons—is not yet supported, because the paper never specifies what the BoE forecast numbers in Table 5 actually are. If those are quarterly average CPI projections, as the MPR publishes, then comparing them to RaGNAR's monthly year-on-year forecasts is invalid. The stress-test note is right to flag this. The paper says comparisons are “limited to the corresponding months” but does not state the target variable or horizon definition of the BoE series, nor does it align data vintages. The evaluation window since end-2019 is short, no equal-predictive-accuracy test is run, and the reported ± intervals are over random graph draws, not over time. That is a load-bearing weakness for the abstract's central claim.\n\nWhat is actually new: the GNAR framework and random-graph idea come from Knight and Leeming, but the authors adapt it sensibly to single-node CPI forecasting—selecting networks each month using past one-step-ahead RMSE, averaging over top networks, and comparing multiple model classes. The benchmark results are more convincing: across the full 2010–2024 evaluation, averaged GNAR forecasts consistently beat simple AR and AvAR benchmarks, especially at longer horizons with relative RMSEs around 0.85. The paper is well written and honest about what didn't work, like the CPI basket-weight graph. That earns credit. The propositions on neighbour-set distributions are correct but minor.\n\nSoft spots besides the BoE comparison: the novelty is incremental—local models and single-node focus are reasonable tweaks, not breakthroughs. The model averaging sets P1/P2 look arbitrary, though the PACF rationale in Appendix B is plausible. The local-αβ models occasionally show very large standard deviations at h=12, which is underexplained.\n\nBottom line: this paper deserves a serious referee—it is not a desk reject. But a responsible referee should demand that the authors identify the BoE forecast object precisely, align vintages and horizons, and run Diebold-Mariano or similar tests. If the BoE comparison cannot be cleaned up, the claim should be downgraded to “competitive with” rather than “materially more accurate.” The benchmark results can stand on their own.","headline":"Solid empirical application of random-graph GNAR models to UK CPI, but the headline BoE outperformance rests on a comparison that is not clearly apples-to-apples and needs more work before it can be believed.","tokens_in":36294,"tokens_out":1964,"would_cite":true,"duration_ms":22790,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M10","62P20","91B84"],"pacs":[],"model":"deepseek-v4-flash","headline":"A network autoregression that averages forecasts over many random graphs predicts UK CPI inflation more accurately than the Bank of England's four-to-six-month projections.","keywords":["inflation forecasting","UK CPI","Generalised Network Autoregressive processes","random graphs","Erdős–Rényi–Gilbert model","model averaging","forecast comparison","disaggregated price indices"],"falsifier":"Using the public ONS component series and the authors' released code, recompute RaGNAR forecasts on aligned data vintages and compare them month-by-month with the Bank of England's published four-to-six-month forecasts; if the RMSE advantage disappears when vintages and target definitions are matched, the central claim collapses. Alternatively, apply the identical procedure to US or euro-area CPI components and compare against the respective central bank's published forecasts.","tokens_in":35267,"feed_emoji":"📈","tokens_out":8614,"duration_ms":74751,"temperature":0.7,"pith_summary":"This paper claims that UK CPI inflation can be forecast more accurately than conventional benchmarks and the Bank of England's published medium-term projections by a simple ensemble called RaGNAR. The method generates thousands of random graphs on the 114 disaggregated CPI component series, fits a Generalised Network Autoregressive (GNAR) process to each, keeps the graphs that predicted best over the past thirty months, and averages forecasts across the top five graphs and several model orders. The authors report that this ensemble beats autoregressive, random-walk, and Chronos benchmarks at every horizon from one to twelve months, with the largest gains at six months and beyond, and that it beats the Bank of England's four-to-six-month forecasts on the overlapping sample. A sympathetic reader would care because the results suggest that medium-term inflation forecasting does not require a large judgmental model suite: rapid, replicable forecasts from public data can compete with a central bank's published numbers.","feed_headline":"Random-graph method beats Bank of England inflation forecasts","feed_subtitle":"Averaging autoregressions over thousands of random graphs on CPI component series cuts forecast error at medium horizons.","key_machinery":"The load-bearing object is the GNAR(p,s) process, an autoregression in which each node's value depends on its own past values plus, for each lag, the average of values from nodes at graph-distance stages one and two, with neighbour sets defined by the graph and averaged with uniform weights. The second ingredient is the random-network ensemble: because selecting among all $2^{6441}$ possible graphs on 114 nodes is infeasible, graphs are sampled from the Erdős–Rényi–Gilbert model with edge probability $\\pi=0.03$, which yields small stage-1 neighbour sets (typically three to four nodes). Each month the graph selection uses the CPI node's one-step-ahead RMSE over the previous thirty months, and the final forecast averages the top five graphs and several model orders, which reduces variance and improves robustness; the paper also derives exact distributions for neighbour-set sizes as a function of $\\pi$.","core_discovery":"The central claim is that RaGNAR — selecting and averaging GNAR processes fitted to Erdős–Rényi–Gilbert random graphs — delivers more accurate UK CPI inflation forecasts than standard benchmarks at all horizons and than the Bank of England's four-to-six-month forecasts. Each month, 10,000 random graphs with edge probability $\\pi=0.03$ are generated on 114 CPI component series, GNAR(p,s) models are fitted at the CPI node using 150 training observations, and graphs are ranked by the root mean squared error of their last thirty one-step-ahead forecasts. The top five graphs are re-fitted and iterated forward to twelve months, with forecasts averaged across graphs and across model orders such as {1,13,25} with neighbour stages {1}, {2}, or {1,2}. Relative to the AvAR(P2) benchmark, relative RMSEs reach 0.87–0.90 at one month and fall to about 0.83–0.85 at six months for the local-$\\alpha\\beta$ class; against the Bank of England, several AvGNAR models improve on the published four-to-six-month RMSEs, with a median improvement near 19% for the global-$\\alpha$ class. The authors also claim that the neighbour sets of the best graphs identify economically interpretable leading components, such as oils and fats, fuels and lubricants, and liquid fuels, which anticipate CPI movements.","pith_inferences":["My inference: the headline comparison with the Bank of England likely depends on aligning forecast vintages and target definitions; if the Bank's published numbers are quarterly averages or use data available later than the RaGNAR information set, the measured gap could narrow, and the paper does not report such an alignment.","My inference: the dominance of single components such as liquid fuels in the selected neighbour sets suggests that the method's edge may come mainly from tracking a few volatile item prices; a sparse factor model or a small VAR on the top five components might recover much of the gain at lower cost.","My inference: applying the same random-network ensemble to CPI component data from other countries, where the Bank of England comparison is replaced by the local central bank's published forecasts, would directly test whether the result is a UK-specific artefact or a general property of network autoregressions on disaggregated price data."],"forward_implications":["Central banks could produce competitive medium-term inflation forecasts from public disaggregated price data alone, in hours on a single processor, without expert judgement.","Averaging forecasts across multiple random graphs — rather than searching for a single 'true' network — is a cheap way to stabilise network-based forecasts and outperforms any single graph.","The method flags a small set of CPI components (oils and fats, fuels and lubricants, liquid fuels) as leading indicators, which could be monitored in real time for early inflation signals.","Because the gains over benchmarks are largest at six to twelve months, RaGNAR is best suited to the policy horizon where the Bank of England's own forecasts are weakest."],"supporting_citations":[{"why":"Defines GNAR(p,s) processes and the estimation approach that RaGNAR fits to each random graph.","marker":"Knight et al. (2020)"},{"why":"Provides the Erdős–Rényi–Gilbert random-graph model used to generate the candidate networks.","marker":"Gilbert (1959)"},{"why":"Introduced the idea of fitting GNAR models to random networks, which RaGNAR adapts for CPI forecasting.","marker":"Leeming (2019)"},{"why":"Describes the Bank of England's COMPASS/MAPS forecasting platform, the official benchmark RaGNAR is compared against.","marker":"Burgess et al. (2013)"},{"why":"Provides the recent review of Bank of England forecast performance and processes that motivates and frames the comparison.","marker":"Bernanke (2024)"},{"why":"Establishes the random-walk-of-order-four benchmark that RaGNAR must beat.","marker":"Atkeson and Ohanian (2001)"},{"why":"Reviews inflation forecasting methods and supports the choice of AR and averaging benchmarks.","marker":"Faust and Wright (2013)"},{"why":"Recent UK CPI forecasting study using disaggregated components; supplies benchmark practice and the data setting.","marker":"Joseph et al. (2024)"},{"why":"Defines the Chronos transformer-based zero-shot models used as additional benchmarks.","marker":"Ansari et al. (2024)"}],"fun_headline_variants":["Random-graph model outforecasts Bank of England on CPI","Averaging random graphs beats BoE inflation predictions","RaGNAR: random graphs sharpen UK CPI forecasts","Graph-averaging method tops BoE on medium-term CPI","Fast random-graph inflation forecasts beat BoE"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim of beating the Bank of England assumes the two sets of forecasts predict the identical thing over the identical window with the same data available to both.","fun_headline_variants_meta":{"raw":{"variants":["Random-graph model outforecasts Bank of England on CPI","Averaging random graphs beats BoE inflation predictions","RaGNAR: random graphs sharpen UK CPI forecasts","Graph-averaging method tops BoE on medium-term CPI","Fast random-graph inflation forecasts beat BoE"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000281,"raw_usage":{"total_tokens":1745,"prompt_tokens":1104,"completion_tokens":641,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":562}},"tokens_in":720,"tokens_out":641,"duration_ms":6285,"temperature":1.0,"reasoning_tokens":562,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:29:14.583493+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Using the public ONS component series and the authors' released code, recompute RaGNAR forecasts on aligned data vintages and compare them month-by-month with the Bank of England's published four-to-six-month forecasts; if the RMSE advantage disappears when vintages and target definitions are matched, the central claim collapses. Alternatively, apply the identical procedure to US or euro-area CPI components and compare against the respective central bank's published forecasts.","supporting_citations":[{"cited_title":"P., and Nunes, M","cited_arxiv_id":null,"evidence_quote":"Defines GNAR(p,s) processes and the estimation approach that RaGNAR fits to each random graph."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Erdős–Rényi–Gilbert random-graph model used to generate the candidate networks."},{"cited_title":"(2019) New Methods in Time Series Analysis: Univariate Testing and Network Autoregression Modelling, Ph.D","cited_arxiv_id":null,"evidence_quote":"Introduced the idea of fitting GNAR models to random networks, which RaGNAR adapts for CPI forecasting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the Bank of England's COMPASS/MAPS forecasting platform, the official benchmark RaGNAR is compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the recent review of Bank of England forecast performance and processes that motivates and frames the comparison."},{"cited_title":"and Ohanian, L","cited_arxiv_id":null,"evidence_quote":"Establishes the random-walk-of-order-four benchmark that RaGNAR must beat."},{"cited_title":"and Wright, J","cited_arxiv_id":null,"evidence_quote":"Reviews inflation forecasting methods and supports the choice of AR and averaging benchmarks."},{"cited_title":"(2024) Forecasting UK inflation bottom up, International Journal of Forecasting, 40, 1521--1538","cited_arxiv_id":null,"evidence_quote":"Recent UK CPI forecasting study using disaggregated components; supplies benchmark practice and the data setting."},{"cited_title":"F., Stella, L., Turkmen, C., Zhang, X., Mercado, P., Shen, H., Shchur, O., Rangapuram, S","cited_arxiv_id":null,"evidence_quote":"Defines the Chronos transformer-based zero-shot models used as additional benchmarks."}],"review_version":1}