{"id":"892ad532-8a52-4167-9546-04799d412afd","arxiv_id":"2607.05450","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Temporal disaggregation inflates in-sample fit while recursive multi-step error compounding degrades out-of-sample cumulative accuracy, with Linear Regression stable and LSTM U-shaped across grains.","lead":"Finer time grains improve training fit and sample size but can wreck multi-step forecasts because recursive models compound errors over longer horizons. The paper shows this trade-off on public procurement data and argues that cumulative planning metrics, not RMSE alone, should decide model and grain.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Single-series recursive-topology claim is load-bearing but under-supported; fixed-H hyperparams and one CPV series leave the structural generalization open.","rationale":"The reader correctly isolates the generalization assumption as the softest load-bearing point: the tables support the recursive-topology story inside this dataset, yet §11 itself flags single-series, single-seed, and no significance tests. My concern is the same one, sharpened to the precise claim that LR stability vs. recursive collapse is topology-driven rather than series-driven. No internal inconsistency appears; the metric-dissensus diagnostic and the naïve recursive controls are well-motivated. Therefore the verdict stays CONDITIONAL—useful if multi-series replication (and code) follow—rather than moving to REJECT or ACCEPT. The concrete test is the minimal experiment that would settle whether the directional findings survive outside CPV-72.","tokens_in":13453,"tokens_out":566,"duration_ms":7429,"concrete_test":"Re-run the identical six-grain, 8-fold expanding-window protocol on at least two additional CPV divisions (or an independent public series of comparable length), keeping the same fixed hyper-parameters and seeds. If Linear Regression TPFE remains inside a ~1 pp band while Holt-Winters/ARIMAX still degrade and LSTM still exhibits the intermediate-grain peak then recovery at Daily, the topology claim is corroborated; if the U-shape or LR flatness disappears on any series, the structural generalization fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that the paradox is driven by recursive feedback topology (not complexity) rests on Linear Regression remaining flat (~16–17% TPFE) while recursive models degrade and LSTM shows a U-shape that recovers only at Daily. That contrast is clear inside Table 1 for this series, but the paper treats the directional patterns as structural properties of model–granularity interaction (§9–11). The weakest link is that every model–grain cell is estimated on one 13-year CPV-72 series, with deep models single-seed and hyperparameters fixed after the first fold (§7). Under those conditions the U-shape and the LR-vs-recursive split could still be series-specific (procurement seasonality, intermittency, or exogenous contract-value dynamics) rather than topology-driven. The formal AR(1) bias-propagation sketch (Eqs. 6–7) is only illustrative; it does not prove that non-recursive linear projection must stay scale-invariant once the data-generating process changes. Without multi-series or multi-seed evidence, the strongest claim over-reaches from a transparent single-series demonstration to a general structural law.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper defines a “Granularity Paradox”: for a fixed planning horizon, finer temporal grains raise N and in-sample fit but lengthen the multi-step horizon H, so recursive/state-dependent models compound forecast error out of sample, while coarse grains avoid recursion but starve estimators. It formalizes recursive bias growth with a simple AR(1) sketch (Eqs. 6–7), introduces cumulative metrics (CFE/TAFE/TPFE) and a consensus–dissensus diagnostic that contrasts pointwise vs cumulative directional behaviour, and benchmarks 10 models (naïve, statistical, XGBoost, LSTM, N-BEATS) over six grains on a 13-year Portuguese public-procurement series (CPV 72) with 8-fold expanding-window backtesting. Empirically, Linear Regression stays near 16–17% TPFE across grains; recursive seasonal models collapse at Daily (Holt-Winters Test R² −151, TPFE 425.85%); LSTM shows a U-shaped TPFE path that recovers only at Daily; and pointwise metrics often improve while TPFE degrades for recursive models.","tokens_in":13867,"tokens_out":1674,"duration_ms":22240,"significance":"If the recursive-topology account generalizes, the paper makes a practically important point: temporal grain is an upstream design choice that can dominate architecture, and evaluation without a goal-dependent cumulative metric systematically misranks models for budgeting/volume planning. Strengths include a transparent multi-model, multi-grain table of 8-fold means; explicit naïve recursive controls; a clear non-recursive contrast (Linear Regression); and a usable consensus–dissensus diagnostic that does not require extra model runs. The AR(1) bias-compounding sketch is standard and supports the qualitative mechanism. The main scientific value is as a careful demonstration and evaluation methodology rather than as a fully established structural law of forecasting.","major_comments":[{"comment":"Abstract and §9–10 treat the LR-vs-recursive split, LSTM U-shape, and “recursive feedback topology, not model complexity” as structural properties of model–granularity interaction. All cells in Table 1 come from one CPV-72 series (§7, §11). The AR(1) sketch (Eqs. 6–7) is illustrative and does not prove scale-invariance of non-recursive linear projection under other DGPs. Either add multi-series evidence (other CPV divisions or public series) or systematically scope abstract/discussion/conclusion to a single-series demonstration with directional hypotheses, not a general structural claim.","section":"Abstract; §9 Discussion; §10 Conclusion; §11"},{"comment":"The LSTM U-shaped TPFE curve (Monthly 19.66% → Bi-Weekly 35.94% → Daily 4.35%, Test R² 0.66) is a headline result and the main evidence that high-capacity recursive models can “overcome” the H penalty. §7 states single-seed LSTM (seed 42), hyperparameters grid-searched on the first fold and held fixed across grains, and no multi-seed CIs or Diebold–Mariano tests. Fixed hyperparams can confound grain with capacity–data match; a single Daily seed can drive the recovery. Report multi-seed means/dispersion for LSTM/N-BEATS (and preferably XGBoost) at least at Monthly/Bi-Weekly/Daily, or move the U-shape from a confirmed threshold to a provisional pattern and soften abstract claims accordingly.","section":"§7 Empirical Evaluation; Table 1; §9 point 3"},{"comment":"The topology claim rests on Linear Regression being non-recursive (“projects predictions directly as a function of time,” §9.2) while ARIMAX/SARIMAX/Holt-Winters/Persistence feed predictions back. The manuscript does not specify the Linear Regression feature set (time index only vs lags/exogenous), how multi-step forecasts are produced at each grain, or whether the same exogenous contract-value series used in ARIMAX/SARIMAX enters LR. Without that protocol, the flat 16.09–16.96% TPFE band cannot be cleanly attributed to absence of recursive feedback rather than to a different information set or direct multi-step setup. Add an explicit forecasting protocol for every model class (recursive vs direct, features, exogenous use).","section":"§7; §9.2; Table 1 Linear Regression rows"},{"comment":"TPFE is defined as absolute cumulative error over the full test window scaled by total observed volume (Eq. 18) and is treated as the primary, goal-dependent criterion. For H=1 (Annual) TAFE equals the single-step absolute error, so cross-grain TPFE comparisons mix one-shot level error with multi-step cumulative bias. The paper already notes R² is undefined at H=1; it should also discuss whether Annual TPFE is commensurate with Daily TPFE for ranking, and whether CFE sign (over- vs under-prediction) matters for procurement budgeting. A short sensitivity (e.g., reporting signed CFE or horizon-normalized cumulative error) would strengthen the metric argument in §6 and §10.","section":"§6 Eqs. 16–18; Table 1 Annual rows; §10"}],"minor_comments":[{"comment":"Table 2’s global log sparklines are described in prose but are hard to verify in the text-only manuscript; ensure the published version has readable glyphs and a self-contained caption defining the [0,1] maps and clipping (R² at −1, TPFE at 100%).","section":"Table 2"},{"comment":"Abstract TPFE band for Linear Regression is “16.3–17.0%” while Table 1 and §9 give 16.09–16.96% (and “16.1–17.0%” in one place). Align all stated ranges with Table 1.","section":"Abstract; §9.2"},{"comment":"Holt-Winters Daily Test R² = −151 is extreme; a one-sentence check that seasonal period, initialization, and recursive update match the Daily grain (and that the figure is not a numerical overflow) would help readers trust the collapse narrative.","section":"Table 1; §9.1"},{"comment":"§4 related work is appropriate; a brief pointer to multi-horizon non-recursive DL (e.g., direct/MIMO or TFT-style) in the discussion would clarify that the paradox is about recursive deployment, not “deep learning” per se—consistent with your N-BEATS vs LSTM contrast.","section":"§4; §9.4"},{"comment":"Reproducibility: Portal BASE is public and filters are described, but code is “on request.” Depositing the preprocessing, fold indices, and configs would match the paper’s methodological emphasis.","section":"§11"}],"recommendation":"major_revision","confidential_remarks":"Fit is reasonable for a forecasting/ML methods venue if claims are scoped to a rigorous single-series study plus evaluation methodology. I would not reject on novelty grounds: TPFE and consensus–dissensus are incremental but useful, and the multi-grain recursive vs non-recursive contrast is cleanly executed. The load-bearing issue is over-generalization from one series and single-seed DL. If the authors add multi-series or multi-seed evidence, this could become a solid contribution; if they only rewrite claims, minor_revision might then suffice on a second round."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core is a transparent six-grain × ten-model backtest on one 13-year Portuguese IT procurement series that shows recursive models can look better in-sample and on RMSE/MAE while cumulative planning error (TPFE) blows up as H grows. Linear Regression stays flat at ~16–17% TPFE across Annual through Daily; Holt-Winters collapses at Daily (Test R² −151, TPFE 425%); LSTM traces a U-shape that only recovers at Daily. That pattern, plus the consensus-dissensus diagnostic that flags when pointwise metrics improve while TPFE degrades, is the actual contribution.\n\nWhat is new is not the idea that aggregation or multi-step strategy matters—MAPA, ADIDA, hierarchies, Ben Taieb, Marcellino are properly cited—but the controlled isolation of grain as the experimental factor and the explicit topology contrast (non-recursive linear projection vs recursive feedback). The AR(1) bias sketch is standard and only illustrative; the table is the evidence. TPFE/CFE are cleanly defined and not circular. The metric warning is practical and well-supported inside this design: RMSE can fall while the planning gap explodes.\n\nSoft spots are real but proportional. Everything rests on one CPV-72 series, single-seed deep models, and hyperparameters fixed after the first fold. The paper itself flags this in §11 and still sometimes writes as if the directional patterns are structural properties of model–grain interaction. That over-reaches; the LR-vs-recursive split and the LSTM U-shape could still be series-specific. No multi-seed intervals, no Diebold–Mariano, code on request only. Those are fixable limits, not internal contradictions.\n\nThis is for applied forecasting people who care about budget/volume horizons and for methodologists who want a clean demonstration that grain choice is upstream of architecture. It is not a foundational result. I would send it to peer review: the design is honest, the tables are readable, and the cumulative-metric point is worth refereeing even if the structural claim needs multi-series tempering. Engage if you work on multi-horizon evaluation or procurement-style series; otherwise skim the tables and the diagnostic.","headline":"Clear single-series map of recursive compounding vs sample size, with a useful cumulative-metric warning; topology claim is suggestive but not yet structural.","tokens_in":14404,"tokens_out":561,"would_cite":false,"duration_ms":6476,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Finer time grains inflate in-sample fit but compound recursive forecast error over longer horizons; the damage tracks feedback topology, not model complexity.","keywords":["Forecasting","Temporal Aggregation","Error Propagation","Deep Learning","Granularity Paradox","Public Procurement","Multi-step Ahead Forecasting","Cumulative Metrics"],"falsifier":"Repeat the six-grain, ten-model expanding-window backtest on several additional independent series (different CPV divisions or domains); if recursive models no longer show systematic TPFE degradation at fine grains relative to non-recursive baselines, or if the LSTM U-shape disappears, the structural claim fails.","tokens_in":14375,"feed_emoji":"📉","tokens_out":650,"duration_ms":6587,"temperature":0.7,"pith_summary":"When you fix a planning window such as one year and choose how finely to slice the series, you face a trade-off the paper names the Granularity Paradox. Daily or weekly data give estimators more observations and prettier training diagnostics, but they force multi-step recursive models to chain hundreds of predictions, so small biases explode. Annual data remove that chaining yet starve the estimator of sample size. Across six grains and ten models on a long public-procurement series, recursive and seasonal methods collapse at high frequency while a plain linear regression of the series on time stays flat near 16–17% total percentage error. An LSTM only recovers after the sample becomes large enough to overcome the same penalty, tracing a U-shaped error curve. Pointwise scores such as RMSE hide the cumulative planning gap; a consensus–dissensus check against total percentage forecast error exposes which models are systematically biased. The practical message is that grain choice is a design decision that controls error dynamics, and evaluation without a goal-dependent cumulative metric systematically misranks models.","feed_headline":"Finer time grains inflate fit but compound recursive forecast error","feed_subtitle":"Linear models stay flat; recursive seasonal methods and mid-grain LSTMs collapse under longer horizons.","key_machinery":"The Granularity Paradox trade-off between sample size N and recursive horizon H, together with the consensus–dissensus diagnostic that compares directional changes in pointwise metrics (RMSE, MAE, R²) against cumulative Total Percentage Forecast Error (TPFE) across grains.","core_discovery":"Finer temporal disaggregation improves in-sample fit and sample size yet degrades out-of-sample multi-step accuracy for recursive models because it lengthens the forecast horizon and compounds feedback error; the effect is driven by recursive topology rather than complexity, as linear regression remains stable across all six grains while recursive seasonal models and intermediate-grain LSTMs degrade.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Finer grains boost fit yet compound recursive multi-step error","Granularity paradox: more data worsens recursive forecast accuracy","Linear models hold flat as recursive ones collapse at fine grains","Disaggregation inflates N and fit while lengthening error horizons","Recursive topology, not complexity, drives the granularity tradeoff"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The directional patterns (recursive collapse at fine grains, LSTM U-shape, linear-regression flatness) are assumed to be structural features of model–grain interaction that will hold beyond the single 13-year IT-procurement series used in the benchmark.","fun_headline_variants_meta":{"raw":{"variants":["Finer grains boost fit yet compound recursive multi-step error","Granularity paradox: more data worsens recursive forecast accuracy","Linear models hold flat as recursive ones collapse at fine grains","Disaggregation inflates N and fit while lengthening error horizons","Recursive topology, not complexity, drives the granularity tradeoff"]},"model":"grok-4.5","effort":"low","cost_usd":0.007506,"raw_usage":{"total_tokens":1829,"prompt_tokens":867,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":75060000,"prompt_tokens_details":{"text_tokens":867,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":877,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":867,"tokens_out":85,"duration_ms":8581,"temperature":1.0,"reasoning_tokens":877,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T20:39:31.099275+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Repeat the six-grain, ten-model expanding-window backtest on several additional independent series (different CPV divisions or domains); if recursive models no longer show systematic TPFE degradation at fine grains relative to non-recursive baselines, or if the LSTM U-shape disappears, the structural claim fails.","supporting_citations":[],"review_version":1}