{"id":"c8fd3e5c-deef-4b87-a14f-22df0d4bd9e0","arxiv_id":"2607.15702","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Variational and primal-dual training of multiscale PDE networks have epsilon-uniform error bounds, while strong-residual training classes provably have Rademacher complexity at least 1/(ε√N) and 1/(ε²√N).","lead":"This paper proves error bounds for variational and primal-dual training of neural feature models on multiscale nonlinear PDEs, showing the bounds do not blow up as the microstructure shrinks. It also proves that strong-form residual losses suffer sample-complexity blow-up proportional to 1/ε, giving a theoretical argument for preferring variational formulations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 4.6 is stronger than it appears: Q0∈W^{1,∞}(Ω×Y) forces weak x-differentiability of ∂ξE(·,∇u0):D²u0, generically requiring u0∈W^{3,∞}, beyond the stated W^{2,∞}; Theorem 4.10's scale-robust bound is conditional on this extra regularity.","rationale":"The reader's verdict is CONDITIONAL, and my stress-test agrees. The most load-bearing point is not an internal inconsistency in the conditional statement—the paper explicitly assumes 4.6—but the fact that 4.6 silently upgrades the regularity of the homogenized solution from W^{2,∞} to W^{3,∞}. That makes the practical domain of the primal-dual end-to-end guarantee narrower than the preceding assumptions suggest. The numerical certificate experiments do not test this: they use small m, one-dimensional problems (where flux correctors vanish), and do not check W^{1,∞} regularity of Q0. Theorem 5.4 is on solid ground: the two-point lower bound follows from Khintchine and the averaging lemma, and the numerical scalings support it. No fatal flaw; the conditional result should remain conditional. Thus verdict UNCHANGED.","tokens_in":18943,"tokens_out":13950,"duration_ms":117700,"concrete_test":"Compute Q0 for a concrete two-dimensional monotone flux, e.g. A_j(y,ξ)=a(y_j)ξ_j+γb(y_j)tanh(ξ_j), with a smooth prescribed u0∈W^{2,∞}(Ω) having nonzero third derivatives at some point (e.g. u0(x)=x1^2(1-x1)^2 x2^2(1-x2)^2 times a smooth bump), using the flux corrector E from Wang et al. [2018, Lemma 2.5]. If the distributional ∂xlQ0 contains D³u0 and is not a function, Assumption 4.6 fails despite u0∈W^{2,∞}. Separately, check whether Lemma 4.7's identity can be proved with Q0 only in L∞; if not, restate Theorem 4.10 under u0∈W^{3,∞} or exhibit the required cancellation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 4.10 rests on the exact divergence identity (Lemma 4.7) and the flux approximation (Proposition 4.8). Both require πε = P0 + εQ0 to have a controlled divergence, and Lemma 4.7 needs ∂xi∂xjEεji = 0 after a second differentiation; with Eεji = Eji(x/ε, ∇u0(x)), this computation needs the composition to be twice differentiable. Because Q0i(x,y) = ∂ξkEji(y, ∇u0(x)) ∂xjxku0(x), membership of Q0 in W^{1,∞}(Ω×Y) requires a weak x-derivative of ∂ξE(·,∇u0):D²u0, which contains terms proportional to D³u0. Assumption 3.1 only postulates u0∈W^{2,∞}. No cancellation mechanism is supplied to remove the third-order terms for a general monotone flux and general f, so Assumption 4.6 is not a harmless 'additional flux-corrector regularity' but a hidden higher-regularity hypothesis on the homogenized solution. The paper is honest that 4.6 is not derived from Wang et al. [2018], and the conclusion labels the result conditional; the concern is that the advertised ε-uniform primal-dual certificate (4.42) applies only in this stronger regularity class, which is not verified numerically or analytically. The strong-form lower bound Theorem 5.4 is independent and unaffected.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops non-asymptotic error bounds for variational physics-informed approximation of uniformly monotone, nonlinear, multiscale elliptic equations, using neural feature classes that are linear in trainable coefficients. The primal result (Theorem 2.4) gives an ε-uniform decomposition of the state error into approximation, empirical-quadrature, and projected-gradient terms. Under a quantitative corrector estimate (Assumption 3.1), a two-scale state class yields approximation error O(ε+Φ0²+Φ1²) in arbitrary dimension (Theorem 3.3). The authors then introduce a convex primal–dual gap, prove a scale-independent certificate inequality (Theorem 4.3), and, under an additional flux-corrector regularity assumption (Assumption 4.6), obtain an end-to-end primal–dual bound with a divergence-compatible flux class (Theorem 4.10). In contrast, for strong residuals they prove optimizer-independent lower bounds on empirical Rademacher complexity of order (ε√N)^{-1} and (ε²√N)^{-1} for residual and squared-residual classes, respectively (Theorem 5.4). Numerical experiments test the predicted ε and N scalings, the lower-tail concentration, and the primal–dual certificate in d=1,2,3.","tokens_in":19350,"tokens_out":8222,"duration_ms":71322,"significance":"If the results hold, the paper makes a useful contribution: it separates formulation-level effects from optimization and sampling effects, and identifies a concrete statistical obstruction for strong-form classes while variational and primal–dual formulations remain scale-robust. The lower-bound theorems are clean, optimizer-independent, dimension-independent, and are supported by careful numerical scaling experiments with explicit constants. The upper-bound theory is also refreshingly explicit: no fitted parameters enter the bounds, the empirical-quadrature and optimization constants are given in closed form, and the assumptions are stated rather than hidden. The main caveat is that the central primal–dual upper theorem is conditional on Assumption 4.6, which is stronger than the surrounding regularity hypotheses and is not numerically exercised in its full nonlinear, multidimensional form.","major_comments":[{"comment":"Assumption 4.6 is load-bearing and is stronger than it appears. The field Q0_i = ∂ξ_k E_ji(y,∇u0) ∂_{x_jx_k} u0 must belong to W^{1,∞}(Ω×Y). Its weak x-derivative generically contains ∂ξ E(·,∇u0):D³u0 terms (together with D²u0⊗D²u0 terms). Assumption 3.1 only postulates u0∈W^{2,∞}, and no cancellation mechanism is supplied for general monotone fluxes. Lemma 4.7 differentiates Eε_ji twice in x to obtain the exact divergence identity (4.32); the proof therefore requires this hidden higher regularity of u0. Since Proposition 4.8 and Theorem 4.10 depend on Lemma 4.7, the advertised ε-uniform primal–dual certificate (4.42) is established only in this stronger regularity class. The paper is honest that Assumption 4.6 is not derived from Wang et al. [2018], but the abstract and conclusion present the primal–dual bound as a main contribution. The assumption should be reformulated explicitly as a","section":"§4.4, Eq. (4.30), Lemma 4.7, Proposition 4.8, Theorem 4.10"},{"comment":"The numerical experiments do not exercise the nonlinear, multidimensional flux-corrector mechanism on which Assumption 4.6 and Theorem 4.10 rely. The primal–dual certificate experiment is one-dimensional and uses a linear flux in ∇v (a(x/ε)∇v), with a nonlinear reaction and a scalar corrector. The flux feature class is correspondingly one-dimensional. The nonlinear multidimensional tests in §6.4 concern the strong-form lower bound of Theorem 5.4, not the upper primal–dual closure. Thus the statement that the computations 'validate every computed primal–dual certificate' is accurate, but the strongest positive result of the paper — the full nonlinear multidimensional primal–dual bound with divergence-compatible flux closure — is not numerically tested under the conditions of Assumption 4.6. The authors should either add such an experiment or explicitly state that the numerical validation","section":"§6.6, Eq. (6.9), Table 6"}],"minor_comments":[{"comment":"The text repeatedly says 'Under Theorem 2.1' and 'Under Theorem 2.3' where the intended referents are Assumption 2.1 and Lemma 2.3. This is confusing because the numbering of theorems and assumptions is already dense.","section":"Assumptions 2.1/2.4 and Lemma 2.2"},{"comment":"The statement says the coefficient ball contains approximants 'realizing' the infima in (3.10), but the proof only uses approximants arbitrarily close to the infima. The wording should be 'near-best approximants' or the assumption should be relaxed to approximate realization.","section":"Theorem 3.3, Eq. (3.10)"},{"comment":"The abstract says the lower bounds hold for 'general periodic nonlinear fluxes', but Theorem 5.4 requires a fixed ε-independent trial state v⋆ satisfying the nondegeneracy condition (5.6). Remark 5.5 states this scope, but the abstract's phrasing could be qualified to avoid overstatement.","section":"Abstract and Remark 5.5"}],"recommendation":"major_revision","confidential_remarks":"The lower-bound part of the paper is solid and should be published. My main concern is the upper primal–dual theorem: Assumption 4.6 appears to hide u0∈W^{3,∞} regularity, and the numerical section does not test the nonlinear multidimensional flux-corrector setting. I recommend major revision to clarify, justify, or visibly weaken the role of Assumption 4.6 in the central claim. I saw no issues with citation practice or novelty; the paper is appropriately scoped."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth reading for the strong-form lower bound alone. Theorem 5.4 is an elegant, optimizer-independent argument: for any fixed trial state whose microscopic flux divergence is nondegenerate, the empirical Rademacher complexity of the residual class is Ω(1/(ε√N)) and the squared-loss class Ω(1/(ε²√N)). The proof via two-scale averaging and Hoeffding is clean, and the numerical scalings in d=1,2,3 fit the predicted exponents almost exactly. That is a real contribution to the debate about why strong-form PINNs struggle with oscillatory coefficients.\n\nThe variational upper bounds are also useful. The paper gives an explicit decomposition of the population error into approximation, quadrature, and optimization terms, with ε-independent constants for the non-approximation parts. The two-scale corrector class and the primal-dual gap certificate are well constructed, and the paper is honest about its hypotheses.\n\nThe soft spot is Assumption 4.6. The flux-corrector regularity Q0 ∈ W^{1,∞}(Ω×Y) is not a harmless technical condition. Since Q0 contains ∂ξE(·,∇u0):D²u0, requiring it to have a weak x-gradient effectively forces u0 into W^{3,∞}, which is stronger than the stated W^{2,∞} in Assumption 3.1. The paper explicitly says this regularity is not derived from Wang et al. [2018], and it is not verified numerically. This means Theorem 4.10's end-to-end certificate (4.42) is conditional on a hypothesis that may fail for generic smooth data. The state approximation closure (Theorem 3.3) still stands under the homogenization assumption (3.9), but the advertised scale-robust primal-dual bound should be read as conditional pending verification of Assumption 4.6.\n\nAlso, no code or data is released. The numerical section is detailed enough to reproduce in principle, but not in practice without the exact scripts and seeds. That is a secondary concern.\n\nThe lower-bound theorem is unaffected by these issues. It is a genuine result and independent of the upper theory.\n\nMy take: this deserves a serious referee. The referee should push on Assumption 4.6, asking for a proof that it follows from natural structural conditions, or a numerical check on the homogenized problem, or a revised statement with one derivative less. The paper also needs some clarifications around the coefficient class (linear features vs. all-weights deep networks, which the paper explicitly disclaims). On balance, it is a solid conditional contribution.\n\nI would bring it to a reading group on learning PDEs and cite the lower bound in my own work. For peer review: send it out, with a request for a careful discussion of Assumption 4.6.","headline":"Worth engaging: the strong-form lower bounds are new and likely correct, but the advertised primal-dual certificate rests on an unproven higher-regularity assumption that should be flagged to any referee.","tokens_in":19767,"tokens_out":2208,"would_cite":true,"duration_ms":18653,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["35B27","65N12","65N15","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"For nonlinear multiscale elliptic PDEs, variational and primal–dual learning keep every error constant scale-independent, while strong-form residual classes suffer an intrinsic 1/(ε√N) statistical floor.","keywords":["multiscale elliptic equations","periodic homogenization","uniformly monotone nonlinear PDEs","variational physics-informed learning","primal-dual gap","Rademacher complexity","corrector-enriched trial spaces","non-asymptotic error bounds"],"falsifier":"Construct a periodic nonlinear flux satisfying the monotone structure but whose flux corrector lacks W^{1,∞} regularity, then check numerically whether the primal–dual gap still dominates the state error; if the certificate fails, the extra regularity is the load-bearing point. Alternatively, choose a trial state v⋆ whose microscopic divergence H(y, ∇v⋆) has zero cell average (Θ⋆ = 0) and verify that the Rademacher lower bound disappears; if the ε⁻¹ scaling persists, the nondegeneracy assumption is not the mechanism.","tokens_in":18806,"feed_emoji":"📐","tokens_out":6116,"duration_ms":60942,"temperature":0.7,"pith_summary":"This paper develops a non-asymptotic error theory for learning solutions of uniformly monotone nonlinear multiscale elliptic equations, where coefficients oscillate on a small scale ε. It proves that variational energy losses and a new convex primal–dual physics loss produce error bounds whose sampling, stability, and optimization constants are independent of ε, provided the trial classes include a two-scale corrector term. In contrast, it proves that strong-form residual losses have empirical Rademacher complexity growing at least as 1/(ε√N), and squared strong losses at least as 1/(ε²√N), in every spatial dimension and independently of the optimizer. The practical consequence is a formulation-level diagnosis: variational and primal–dual schemes avoid an intrinsic statistical ill-conditioning that strong-form residual learning suffers when the coefficient oscillates. Numerical experiments in d=1,2,3 confirm the predicted ε- and N-scalings and show the primal–dual gap reliably certifies the state error.","feed_headline":"Strong-form residuals blow up as ε shrinks; variational ones don't","feed_subtitle":"New bounds show variational and primal–dual losses avoid the ε⁻¹ blow-up that strong residuals suffer.","key_machinery":"Two mechanisms carry the argument. The upper bound uses the corrector-enriched two-scale trial class v = U(x) + ε ηε(x) C(x, x/ε), built from the periodic homogenization corrector; its gradient features are ε-uniform because the fast derivative is absorbed by the cutoff and cell derivative. The upper certificate uses the convex primal–dual gap Gε(v,p) = Jε(v) + ∫[F*(x,x/ε,p) + G*(x, f+∇·p)] dx, whose population value bounds the state error from above via Fenchel–Young and equals zero at the exact pair (uε, pε). The lower bound uses a fixed smooth trial state v⋆: the chain rule expresses its strong residual as -(1/ε)H(x/ε, ∇v⋆) plus bounded terms, and the nondegeneracy Θ⋆ > 0 of the microscop","core_discovery":"The paper's central claim is a separation between formulations by scale-robustness. For the variational energy formulation, the state error decomposes into an approximation term, an empirical quadrature term, and a projected-gradient term, with all non-approximation constants uniform in ε; a corrector-enriched two-scale state class closes the approximation term with a bound C(ε + Φ₀² + Φ₁²) in arbitrary dimension. For a new convex primal–dual loss, the population value is a computable upper certificate for the state error, and with an additional flux-corrector regularity assumption, a divergence-compatible two-scale flux class closes the flux approximation, giving a certified state–flux boun","pith_inferences":["Extension (editorial): the same chain-rule amplification should appear in time-dependent or stochastic multiscale problems wherever a fast variable enters through a derivative, so the formulation-level separation is likely not specific to elliptic equations.","Extension (editorial): the theory suggests a testable architectural principle—strong-form classes can avoid the blow-up only by being coefficient-adapted so that the microscopic divergence has vanishing cell average; quantifying the minimal approximation complexity of such adapted classes is a natural next step.","Extension (editorial): because the primal–dual certificate is convex and projected gradient converges globally, the bound (4.42) could be turned into an adaptive sampling algorithm that grows N until the empirical gap drops below a target accuracy, using the certificate as a stopping rule.","Extension (editorial): the ε⁻¹ and ε⁻² exponents might persist for non-monotone fluxes or higher-order operators where the chain-rule amplification changes; testing those generalizations would clarify how universal the strong-form obstruction is."],"forward_implications":["If the bounds hold, variational and primal–dual methods for multiscale elliptic problems inherit ε-uniform sample complexity: the number of samples and optimization iterations needed for a target accuracy does not grow as the microstructure refines.","The primal–dual gap provides a computable, dimension-independent certificate for the state error that practitioners can monitor during training; numerical results show it stays on the safe side of the diagonal with margins near machine precision.","Corrector-enriched trial classes—adding ε ηε(x) C(x, x/ε) features—reduce energy and H¹ errors as ε→0, so a small amount of homogenization-aware architecture delivers the benefit that would otherwise require many more blind coefficients.","Any strong-form class containing an ε-independent nondegenerate witness state inherits the 1/(ε√N) and 1/(ε²√N) lower bounds, making the obstruction optimizer-independent and dimension-independent.","The formulation-level separation suggests that reporting only training error for strong-form PINNs on oscillatory coefficients is misleading; the finite-sample statistical floor will appear as ε shrinks no matter the optimizer."],"fun_headline_variants":["Variational loss avoids ε⁻¹ blow-up; strong residual doesn't","Non-asymptotic bounds: primal-dual loss certifies state error uniformly in ε","Strong-form residuals scale as 1/(ε√N); variational ones stay uniform","Scale-robust PDE learning: ε-uniform bounds for variational and primal-dual losses"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The certified primal–dual error bound rests on an extra smoothness condition on the flux corrector—its leading correction field must have bounded first derivatives in both the spatial and periodic variables—which the paper explicitly says is not implied by the homogenization theory it cites; if that smoothness fails, the certificate in the end-to-end bound does not follow, even though the state approximation theory may remain valid.","fun_headline_variants_meta":{"raw":{"variants":["Variational loss avoids ε⁻¹ blow-up; strong residual doesn't","Non-asymptotic bounds: primal-dual loss certifies state error uniformly in ε","Strong-form residuals scale as 1/(ε√N); variational ones stay uniform","Scale-robust PDE learning: ε-uniform bounds for variational and primal-dual losses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000316,"raw_usage":{"total_tokens":1686,"prompt_tokens":865,"completion_tokens":821,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":734}},"tokens_in":609,"tokens_out":821,"duration_ms":7351,"temperature":1.0,"reasoning_tokens":734,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T22:32:41.453949+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a periodic nonlinear flux satisfying the monotone structure but whose flux corrector lacks W^{1,∞} regularity, then check numerically whether the primal–dual gap still dominates the state error; if the certificate fails, the extra regularity is the load-bearing point. Alternatively, choose a trial state v⋆ whose microscopic divergence H(y, ∇v⋆) has zero cell average (Θ⋆ = 0) and verify that the Rademacher lower bound disappears; if the ε⁻¹ scaling persists, the nondegeneracy assumption is not the mechanism.","supporting_citations":[],"review_version":1}