{"id":"62b9ed97-63a6-4936-92b1-ccf117343123","arxiv_id":"2608.08633","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"For periodic scalar upwind advection, the parent-average history has exact rank P+(P-1)min{L,r-1}; a flux-divergence queue attains this bound, and small Courant numbers make deep delayed information numerically unrecoverable.","lead":"This paper works out exactly how much extra information is needed, beyond coarse cell averages, to predict future coarse averages in a simple model: scalar advection with an upwind finite-volume scheme on a periodic grid. It gives precise dimension counts, shows when that information is numerically recoverable, and when numerical diffusion makes it irrelevant.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the rank law, encoder bounds, queue attainment, conditioning estimates, and dissipative decay claims are internally consistent.","rationale":"I independently re-derived the key steps: the row-span equivalence between {R, RA^t} and {R, RS^t}, the D_t layer exposure bound, the block-determinant independence argument, the product-local injection, the saturated recurrence, the T_q inverse and norm formulas, and the Fourier contraction in Theorem 3. All are correct. The numerical experiments are consistent with the theory: the rank ladder matches the formula, the effective rank drops for small λ as predicted by the λ^q scaling, the step-function flattening follows the dissipative envelope, and the delayed-collision experiment shows the claimed sequence of hidden, then observable, then dissipated. The reader's weakest assumption correctly identifies the explicit scope restrictions (continuous encoders on open supports; declared bounded-perturbation model) rather than an internal inconsistency. Since these restrictions are stated as hypotheses in Sections 4 and 6 and acknowledged in Section 8.1, they do not change the verdict. ACCEPT with moderate confidence remains appropriate.","tokens_in":22297,"tokens_out":24774,"duration_ms":275040,"concrete_test":"As an independent check of the central rank formula, compute rank O_L exactly in rational arithmetic (e.g., with sympy) for P=3, r=4, λ=1/3, L=0,...,6 and confirm the predicted sequence 3, 5, 7, 9, 9, 9, 9; if any entry differs, the proof of Theorem 1 contains a hidden linear-algebra flaw.","verdict_should_be":"UNCHANGED","load_bearing_attack":"I read the central claims in good faith and found no load-bearing error. Theorem 1 is sound: the change from RA^t to RS^t is triangular with nonzero diagonal, the D_t layer-difference rows give the rank upper bound, and the exposed-layer submatrix has determinant ±r^{-P}, giving the matching lower bound. The encoder lower bounds follow from invariance of domain exactly under the stated continuous-encoder/open-support hypotheses, and Section 8.1 explicitly disclaims discontinuous or non-open extensions. The saturated flux-divergence queue attains the centralized bound, and the recurrence used for autonomous advancement checks out. Theorem 2's inverse, norms, and singular-value bounds are correct, and Theorem 3 is a clean Fourier contraction estimate with a valid perturbation iteration. The only delicate points are the explicitly declared scope restrictions: continuous encoders on open supports for the topological bounds, and the bounded-perturbation model for floating-point arithmetic. Because these are stated hypotheses rather than hidden gaps, they do not threaten the theorem's validity as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper treats coarse finite-volume parent averages as a predictive-state question for the periodic scalar advection equation discretized by first-order upwind with forward Euler. Its main object is the finite-horizon observation matrix O_L = [R; RA_λ; ...; RA_λ^L], and Theorem 1 establishes the exact rank law rank O_L = P + (P−1) min{L, r−1} for P parents with r child cells per parent. From this it derives lower bounds of (P−1) min{L, r−1} additional real coordinates for continuous centralized encoders and min{L, r−1} coordinates per parent for product-local encoders, and constructs an anchored flux-divergence queue that attains the centralized bound and, at saturation, yields a minimal autonomous predictive state of dimension Pr−r+1. Theorem 2 gives the exact conditioning κ∞(T_q) = ((2−λ)/λ)^{q−1} for the triangular collar-to-queue map, and Theorem 3 proves contraction of the nonconstant fine-grid component under a declared bounded-perturbation arithmetic model, with separate control of mean drift. Numerical experiments illustrate the rank ladder, tolerance-dependent effective rank, queue conditioning, step-function flattening, and delayed coarse separation followed by dissipative decay.","tokens_in":22473,"tokens_out":22468,"duration_ms":214335,"significance":"When accepted as stated, the paper delivers a rare exact benchmark separating three properties that are often conflated in reduced and learned coarse models: the algebraic dimension of predictive memory, the conditioning with which delayed information can be recovered, and the time over which numerical dissipation erases it. The proofs of the rank theorem, the encoder lower bounds, the attaining queue construction, and the contraction estimate are internally consistent; I checked the row-rank construction, the triangular collar-to-queue inversion, and the Fourier eigenvalue estimate, and found no gaps. The encoder lower bounds are correctly scoped to continuous maps on open supports, and the bounded-perturbation model is declared rather than disguised as a derived IEEE roundoff bound. The public reproducibility package with executable certificates is an additional strength. The contribution is a well-posed solvable benchmark rather than a broad turbulence-closure theorem, and the paper is appropriately modest about that distinction.","major_comments":[],"minor_comments":[{"comment":"The statement uses δ^n both for the perturbation vector and for its spatial mean: the displayed sentence 'Here δ^n = 1/N 1^T δ^n is the spatial mean of the arithmetic perturbation at step n' cannot be read literally. Please use an overbar for the mean and state the mean-drift assumption as |mean(δ^n)| ≤ η_m.","section":"§6, Theorem 3"},{"comment":"The abstract's sentence 'Centralized prediction therefore requires one fewer additional coordinate than the number of parent cells per exposed layer' states the lower bound without the qualifier 'for continuous encoders on open supports.' Section 8.1 gives the restriction correctly; adding the same qualifier to the abstract and contribution list would avoid an over-general reading by downstream users.","section":"Abstract and §1.1"},{"comment":"The effective-rank curves for several λ overlap with the exact-rank curve and with each other, as the caption explains, but the degree of overlap makes the effective-rank loss hard to read in monochrome print. A short table of rankτ(O_L) values at L = 5, or a few marked points, would make the λ-dependence of the loss immediately visible.","section":"§7.2, Figure 1"},{"comment":"The remark correctly notes that the condition ζ > E_floor can fail when E_floor grows with N, but the mechanism would be clearer if it stated directly that E_floor = η/(1−ρ_N) scales like η N^2/[2π^2 λ(1−λ)] for fixed λ and η.","section":"§6, Remark 1"},{"comment":"The caption states that dashed curves are the theoretical bounded-perturbation envelope, but it does not give the value of η used. Since η = 200 ε_mach √N is a declared diagnostic rather than a derived roundoff bound, include it in the caption or legend for reproducibility.","section":"§7.5, Figure 4 caption"}],"recommendation":"minor_revision","confidential_remarks":"The technical content appears correct and the paper is a good fit for math.NA. The changes I would request before acceptance are the notation fix in Theorem 3 and the scope qualifier in the abstract; both are local and do not require another full review round."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plainly: this is the rare submission where the scope is the result. The paper proves a rank formula for finite-horizon parent-average observations of periodic scalar upwind: rank O_L = P + (P-1) min{L,r-1}, gives matching centralized and product-local encoder lower bounds under continuity/open-support assumptions, constructs an anchored flux-divergence queue that attains the bound, and shows the saturated state is autonomous and minimal. The conditioning theorem (κ∞(T_q)=((2-λ)/λ)^(q-1), σ_min ~ λ^q) and the dissipative contraction bound under bounded perturbations are also correct. I checked the central proofs line by line: the triangular change of basis, the exposed-layer determinant, the local injection via invariance of domain, the saturated recurrence, the inverse of T_q, and the Fourier contraction all check out. The code and data package is a plus; the experiments reproduce the rank ladder and the effective-rank loss at small λ as expected. The citation pattern is standard and fair, with no red flags.\n\nWhat is genuinely new is the exact synthesis: separating algebraic observability from stable recoverability from dissipative erasure in one concrete FV hierarchy. That is useful to people building learned closures or reduced models, because it gives a solvable benchmark with a sharp dimension count.\n\nSoft spots, in proportion: they are mostly declared scope restrictions, not hidden gaps. The encoder lower bounds use topological dimension arguments and require continuous encoders on open supports. Allow discontinuous encoders or non-open supports and the bounds can fail; the paper explicitly says this in Section 8.1, but it means the \"exact memory\" claim is conditional in a way that matters for ML practice, where discontinuous maps are common. The dissipative-erasure theorem assumes the computed update is the exact upwind update plus a per-step perturbation with separately bounded mean and nonconstant parts. That is a reasonable model of arithmetic, not a derived IEEE roundoff bound, and the η used in the flattening plots is hand-chosen (200 ε_mach√N). Again the authors flag it. The effective-rank tolerance C_svd=100 is a declared diagnostic. None of this undermines the theorems; it just tells you how far the conclusions travel.\n\nBottom line: a narrow but solid paper, clear about its own limits, with real proofs and reproducibility. I would send it to a referee rather than desk-reject. For a reading group on closure models or observability of numerical schemes, it is worth an hour. I probably wouldn't cite it in my own work unless I start working on coarse-state dimension, but I would be glad it is in the literature.","headline":"A correct, carefully scoped exact-rank and conditioning analysis for coarse upwind prediction; the topological lower bounds and perturbation model are narrower than the title might suggest, but the authors say so.","tokens_in":22995,"tokens_out":2506,"would_cite":false,"duration_ms":28399,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65M08","65F35","93B07","65M12"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves an exact observation-rank law for coarse upwind advection: parent averages alone are not a predictive state, and the missing information is counted exactly, then shown to be partly unrecoverable or dissipatively erased.","keywords":["finite-volume methods","coarse graining","predictive memory","observability","numerical dissipation","stable recovery","upwind scheme","observation rank"],"falsifier":"Compute the singular values (or a rank-revealing factorization) of $O_L$ exactly, for example in rational arithmetic, for $P=8$, $r=6$, $\\lambda=1/2$ across $L=0,\\dots,9$; the theorem is false if the rank ever differs from $P+(P-1)\\min\\{L,5\\}$.","tokens_in":22096,"feed_emoji":"🌊","tokens_out":7467,"duration_ms":70767,"temperature":0.7,"pith_summary":"Coarse finite-volume averages are not, in general, an exact predictive state: two fine-grid states with the same parent averages can generate different future coarse histories under upwind advection. The paper's main result is an exact finite-horizon rank law for periodic scalar advection with $P$ parent cells, $r$ children per parent, and horizon $L$: the observation matrix has rank $P + (P-1)\\min\\{L,r-1\\}$. It follows that any continuous centralized encoder needs $(P-1)\\min\\{L,r-1\\}$ extra real coordinates beyond the parent averages, while any product-local encoder needs $\\min\\{L,r-1\\}$ coordinates per parent, and an anchored flux-divergence queue attains the centralized bound. The paper then separates exact observability from stable and long-time relevance: the delayed subcell information enters through a triangular map whose smallest singular values scale like $\\lambda^q$, and for $0<\\lambda<1$ the nonconstant part of the solution is proved to contract toward a perturbation-controlled floor. This matters because reduced, multiscale, and learned models that claim closure on coarse observables must specify hidden information, horizon, conditioning, and dissipation; the scalar upwind hierarchy becomes a solvable benchmark for all of them.","feed_headline":"Coarse upwind averages fail as predictive states","feed_subtitle":"Exact rank law: each new observed layer costs P-1 hidden coordinates, and dissipation later hides them anyway.","key_machinery":"The central object is the finite-horizon observation matrix $O_L$, built from the restriction operator $R$ and the upwind shift $A_\\lambda=(1-\\lambda)I+\\lambda S$; $O_L x$ is the sequence of parent averages from time $0$ to $L$. The rank law follows because the powers $A_\\lambda^t$ span the same row space as $R,RS,\\ldots,RS^L$, and each shift $S^{t+1}$ exposes one new right-collar child layer through the cyclic difference operator $D_t=r(RS^{t+1}-RS^t)$ whose rows sum to zero. The attaining memory is the anchored flux-divergence queue $JB\\Phi^t$, the cyclic differences of the realized right-face fluxes with one coordinate deleted; at saturation its dimension $P-1$ per layer is exactly the rank increment. Conditioning is carried by the triangular collar-to-queue matrix $T_q$ with diagonal $\\lambda,\\ldots,\\lambda^q$, and long-time decay by the per-step contraction rate $\\rho_N=(1-4\\lambda(1-\\lambda)\\sin^2(\\pi/N))^{1/2}$.","core_discovery":"The central discovery is a finite-horizon observation-rank law for the periodic scalar upwind finite-volume hierarchy: for $L\\ge 0$, $P\\ge 2$, $r\\ge 2$ and Courant number $0<\\lambda\\le 1$, the matrix $O_L=(R^\\top,(RA_\\lambda)^\\top,\\ldots,(RA_\\lambda^L)^\\top)^\\top$ has rank $P+(P-1)\\min\\{L,r-1\\}$. Each additional observation exposes one new child-cell layer and contributes $P-1$ independent directions, one fewer than the number of parents, because the spatially uniform flux gauge is invisible to the conservative update. The paper then constructs an anchored flux-divergence queue, the list of cyclic flux differences $J(B\\Phi^0),\\ldots,J(B\\Phi^{q-1})$, that attains this centralized lower bound and, together with the parent averages, forms a minimal autonomous predictive state of dimension $Pr-r+1$ at saturation. Two companion results sharpen the practical meaning: the collar-to-queue map $T_q$ is triangular with diagonal $\\lambda,\\ldots,\\lambda^q$, so its infinity-norm condition number is $((2-\\lambda)/\\lambda)^{q-1}$ and the deepest exposed layer is recoverable only with amplification of order $\\lambda^{-q}$; and for $0<\\lambda<1$ the nonconstant component of the computed solution contracts at rate $\\rho_N$ under separately bounded mean and nonconstant perturbations, so information that is algebraically necessary for short-time closure can be numerically invisible or dissipatively erased later.","pith_inferences":["Repeating the same rank-and-conditioning program for higher-order scalar schemes would test whether a wider exposed collar per step changes the $P-1$ increment; the paper identifies this as a natural next step but does not prove it.","For nonperiodic boundaries the periodic flux gauge is replaced by boundary data, so the $r-1$ coordinate saving at saturation should shrink or vanish; recomputing the rank law with Dirichlet or absorbing interfaces would test this.","The perturbed-decay theorem treats $\\eta$ and $\\eta_m$ as declared per-step error bounds; measuring them for a concrete floating-point implementation would convert the contraction estimate into a certificate for that solver.","A learned closure trained on this hierarchy cannot reproduce all coarse histories at or beyond saturation with fewer than $Pr-r+1$ internal coordinates unless it is discontinuous; this gives a testable lower bound on the internal state dimension of such models."],"forward_implications":["For any continuous centralized predictor on open support, the minimum number of auxiliary coordinates needed to reproduce $L+1$ steps of parent-average history is $(P-1)\\min\\{L,r-1\\}$, and the anchored flux-divergence queue attains it.","At saturation ($L\\ge r-1$), parent averages plus the queue form an autonomous linear predictive state of dimension $Pr-r+1$ that advances by a fixed matrix without revisiting the fine grid.","At small Courant number the deep subcell layers remain algebraically visible but numerically unrecoverable: their singular values fall like $\\lambda^q$, so a declared SVD tolerance records an effective rank below the exact rank.","For $0<\\lambda<1$ hidden subcell differences are first invisible, then become visible after a finite delay, and then contract at the rate $\\rho_N$ toward a perturbation-controlled floor, so exact short-time closure and long-time practical relevance are distinct."],"supporting_citations":[{"why":"gives the conservative flux-form finite-volume update that Proposition 1 telescopes","marker":"[1]"},{"why":"supplies the grid-interface conservation baseline behind one-step reflux messages","marker":"[9]"},{"why":"grounds the unresolved-variable memory viewpoint the introduction contrasts with predictive states","marker":"[12]"},{"why":"adds the complementary memory-flux formalism for coarse non-Markovian dynamics","marker":"[13]"},{"why":"defines observability for discretized PDEs and frames the stable-recovery question","marker":"[42]"},{"why":"provides the singular-value numerical-rank convention used in the effective-rank diagnostic","marker":"[40]"},{"why":"supplies the floating-point accuracy model behind the tolerance threshold and conditioning discussion","marker":"[41]"},{"why":"underwrites the minimal-state/observability perspective behind the minimal autonomous predictive-state dimension","marker":"[44]"}],"fun_headline_variants":["Exact cost of memory: P-1 directions per observation","Dissipative decay hides algebraically visible info","Coarse upwind averages fail, but an anchored queue fixes it","P-1 hidden coordinates per new layer: the rank law"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The sharp lower-bound results assume a continuous encoder on an open support; with discontinuous encoders, or with non-open supports, the stated minimum extra-coordinate counts can be bypassed, and the decay theorem separately assumes a declared split of arithmetic error into mean and nonconstant parts.","fun_headline_variants_meta":{"raw":{"variants":["Exact cost of memory: P-1 directions per observation","Dissipative decay hides algebraically visible info","Coarse upwind averages fail, but an anchored queue fixes it","P-1 hidden coordinates per new layer: the rank law"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000452,"raw_usage":{"total_tokens":2388,"prompt_tokens":1170,"completion_tokens":1218,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":786,"completion_tokens_details":{"reasoning_tokens":1149}},"tokens_in":786,"tokens_out":1218,"duration_ms":12123,"temperature":1.0,"reasoning_tokens":1149,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:30:46.980258+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the singular values (or a rank-revealing factorization) of $O_L$ exactly, for example in rational arithmetic, for $P=8$, $r=6$, $\\lambda=1/2$ across $L=0,\\dots,9$; the theorem is false if the rank ever differs from $P+(P-1)\\min\\{L,5\\}$.","supporting_citations":[{"cited_title":"Cambridge Texts in Applied Mathematics","cited_arxiv_id":null,"evidence_quote":"gives the conservative flux-form finite-volume update that Proposition 1 telescopes"},{"cited_title":"Mathematics of Computation41(164), 321–336 (1983) https: //doi.org/10.1090/S0025-5718-1983-0717689-8","cited_arxiv_id":null,"evidence_quote":"supplies the grid-interface conservation baseline behind one-step reflux messages"},{"cited_title":"SIAM Journal on Numerical Analysis25(3), 586–617 (1988) https://doi.org/10.1137/0725037","cited_arxiv_id":null,"evidence_quote":"defines observability for discretized PDEs and frames the stable-recovery question"},{"cited_title":"Society for Industrial and Applied Mathematics, Philadelphia (1997)","cited_arxiv_id":null,"evidence_quote":"provides the singular-value numerical-rank convention used in the effective-rank diagnostic"}],"review_version":1}