{"id":"762827d5-b221-4a63-ac23-ca6a72b8952b","arxiv_id":"2607.07778","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":8.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Any width-m two-layer piecewise-linear network with arbitrary weights that fits n noisy labels below the noise floor has Lip ≳ ε sqrt(n/(m log(m n d/ε))) with high probability on the sphere or Gaussian.","lead":"Two-layer ReLU networks with unrestricted weights that fit noisy labels must have Lipschitz constant at least roughly sqrt(n/m), up to a log. This settles the Bubeck–Li–Nagaraj conjecture for piecewise-linear activations and shows depth two is the critical depth for the law of robustness.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly identifies rigidity as the novel deterministic ingredient and correctly notes that the paper itself proves the d=2 failure. Because the manuscript already delimits the geometric hypothesis, supplies the full proofs, and matches the lower bound at the over-parameterized endpoint, the argument for the stated theorem (up to one log) is internally consistent and complete. No stronger load-bearing concern appears; the remaining open questions (log-free bound, general activations at moderate width) are explicitly isolated and do not undermine Theorem 1.2. The verdict ACCEPT with high confidence therefore stands.","tokens_in":34722,"tokens_out":460,"duration_ms":5787,"concrete_test":"Independently re-derive the one-sided derivative jump identity in the proof of Lemma 3.1 (ball) and Lemma 3.2 (sphere) for a two-kink configuration with non-parallel hyperplanes; confirm that |α_{j}| (resp. |α_{j}|√(1-t_{j}^{2})) remains ≤2L at a generic point of the first hyperplane. If the identity holds, the load-bearing step is secure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of Theorem 1.2 rests on the rigidity lemmas (3.1–3.2) that convert an L-Lipschitz bound into coefficient bounds on the canonical kink representation, enabling a function-space entropy estimate independent of the original (unbounded) weights. The manuscript already isolates the precise geometric condition under which this holds (generic points on distinct hyperplanes exist only for d≥3 on the sphere; Proposition 3.3 exhibits an explicit counter-example on S^{1}) and supplies a complete, self-contained proof of the entropy bound, the finite-class concentration argument, and the matching O(1)-Lipschitz construction at width 2n. No hidden cancellation, measurability gap, or regime in which the stated hypotheses fail was found that would invalidate the theorem as written. The single logarithmic factor is flagged throughout and is not claimed to be removable by the present argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proves the Bubeck–Li–Nagaraj law of robustness for two-layer networks with continuous piecewise-linear activations (including ReLU) and completely unrestricted weights, up to a single logarithmic factor. For data uniform on S^{d-1} (d≥3) or N(0,I_d/d), labels in [-1,1] with noise level σ^{2}>0, any width-m network that fits ε below the noise floor must satisfy Lip_D(f)≥c_{0}ε√(n/(m̄ log(C_{0} m̄ n d/ε))) with high probability (Theorem 1.2), where m̄=(K-1)m+1. The argument replaces parameter-space covering by a function-space covering, using a rigidity lemma (Lemmas 3.1–3.2) that bounds each canonical kink coefficient by a multiple of Lip(f). Realized-kink-count, simultaneous-width, structured-architecture, and vector-output variants are derived; a matching O(1)-Lipschitz two-layer ReLU interpolant at width 2n is given (Proposition 10.1); and the log-free case is reduced to an open multiplier estimate (Conjecture 9.1).","tokens_in":34940,"tokens_out":1054,"duration_ms":24946,"significance":"The result settles the natural remaining boundary case of the BLN conjecture: depth-two piecewise-linear networks with unbounded weights, the regime in which Bubeck–Sellke’s universal law required a polynomial-parameter hypothesis known to be necessary already at depth three. The rigidity phenomenon (kinks on distinct hyperplanes cannot cancel at generic points) is a clean geometric contribution that yields a bounded canonical representation without a priori weight bounds. Full self-contained proofs appear in Appendix A, an explicit matching construction is supplied at the overparameterized endpoint m≃n, and numerical scripts are released. The residual logarithm is flagged throughout and partially removed in Section 7 under a stronger sample-size hypothesis; the remaining log-free obstacle is isolated as a single, sharply stated multiplier estimate. These features make the paper a substantial and carefully scoped advance for the theory of robust interpolation.","major_comments":[],"minor_comments":[{"comment":"The relationship between the main text and the supplementary note [13] (reduction of the log-free law to Conjecture 9.1) could be flagged more prominently in the introduction, e.g., a single sentence stating that the note is not required for any theorem proved in the present manuscript.","section":"§1 / §9"},{"comment":"In Lemma 3.6(ii) the bound ∥v∥≤d(B_{0}+2Lm_{0}) is dimension-linear; a short remark that this is absorbed into the entropy logarithm (and cancelled by the isoperimetric factor in Case A of the proof of Theorem 1.2) would help readers who first meet the affine-part estimates.","section":"Lemma 3.6"},{"comment":"Proposition 4.1 states the entropy bound for ε'∈(0,4+8L); the parenthetical that the grid construction works at every scale is correct but easy to miss. A one-line clarification that the displayed closed form is only for notational convenience would remove any ambiguity.","section":"Proposition 4.1"},{"comment":"Section 11 reports path-norm-penalized training runs that stay comfortably above the lower bound but do not saturate √n/m for m≪n. A brief sentence noting that the path-norm penalty is only a proxy (and that the true min-Lipschitz interpolant may be smaller) would prevent over-interpretation of the ratios in the table.","section":"§11"},{"comment":"Minor typographical consistency: the manuscript mixes “Lip_D(f)” and “Lip_{S^{d-1}}(f)”; either is fine, but a single convention after the standing notation list would be cleaner.","section":"§1.1"}],"recommendation":"accept","confidential_remarks":"The manuscript is unusually careful about scope (d=2 failure, smooth activations, depth three, residual log) and ships complete proofs plus code. I see no load-bearing gap. Fit for a strong theory venue is excellent; the residual open multiplier estimate is cleanly isolated and does not undermine the main theorems."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This one actually settles the Bubeck–Li–Nagaraj conjecture for two-layer piecewise-linear nets with arbitrary weights, up to a single log. That is the real news: earlier universal laws needed polynomial parameter bounds, and Bubeck–Sellke already showed those bounds are necessary at depth three. Depth two is therefore the critical case, and Shmalo closes it for ReLU and friends.\n\nWhat is new is the function-space route. Instead of covering weights (impossible when they are unbounded), the paper rewrites every network in a canonical kink form and proves a rigidity lemma: on the ball, and on the sphere for d ≥ 3, each kink coefficient is controlled by the Lipschitz constant of the realized function, because distinct hyperplanes cannot cancel at a generic point. Once the coefficients are bounded, ordinary entropy plus a self-contained finite-class concentration argument finishes the job. The matching O(1)-Lipschitz construction at width 2n is explicit and tight at the over-parameterized end. Full proofs sit in the appendix; the numerical checks are only sanity tests and do not set constants.\n\nSoft spots are the ones the author already flags. Rigidity fails on the circle (explicit counter-example), so the sphere model needs d ≥ 3. The log remains; Section 7 removes the sample size from it under a stronger sample-size hypothesis, but the clean √(n/m) is still open and reduced to one multiplier estimate. General smooth activations are outside the kink mechanism, though small-width projection floors cover every Lipschitz activation. None of these undercut the main theorem as stated.\n\nThis is for people who care about capacity, robustness, and the precise role of depth. The math is careful, the citations are on point, and the remaining open pieces are cleanly isolated. I would send it to referees without hesitation and would cite the rigidity-plus-function-space idea myself.","headline":"Proves the BLN robustness law for unbounded two-layer ReLU (up to one log) via a clean kink-rigidity argument that replaces parameter covering.","tokens_in":35489,"tokens_out":487,"would_cite":true,"duration_ms":6941,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T07","68Q32","60F10","60B20","60B15"],"pacs":[],"model":"grok-4.5","headline":"Any two-layer network with piecewise-linear activation and arbitrary weights that fits n noisy labels must have Lipschitz constant at least order sqrt(n/m), up to one log factor.","keywords":["law of robustness","two-layer neural networks","ReLU networks","arbitrary weights","Lipschitz interpolation","metric entropy","isoperimetry","kink rigidity"],"falsifier":"Exhibit a continuous piecewise-linear two-layer network of width m, with arbitrary real weights, that fits generic sphere or Gaussian data epsilon below the noise floor while keeping Lipschitz constant o(epsilon sqrt(n/(m log(m n d)))) in dimension d greater than or equal to 3.","tokens_in":35629,"feed_emoji":"⚖️","tokens_out":721,"duration_ms":13822,"temperature":0.7,"pith_summary":"The paper settles a conjecture of Bubeck, Li and Nagaraj for continuous piecewise-linear activations, including ReLU: on generic high-dimensional data, a width-m two-layer network that interpolates noisy labels well below the noise floor cannot be very smooth. Its Lipschitz constant is forced to grow like epsilon times the square root of n over m, times a single logarithmic factor, even when every weight is allowed to be arbitrarily large. The argument works by rewriting every realized network into a canonical kink form and proving that distinct kink hyperplanes cannot cancel, so each kink coefficient is controlled by the overall Lipschitz constant. That rigidity turns an unbounded-parameter class into a function class of controlled entropy, after which standard concentration finishes the proof. A matching explicit construction at width 2n shows the bound is essentially tight once there is roughly one neuron per sample.","feed_headline":"Two-layer nets need large Lip constant to fit noisy labels","feed_subtitle":"Even with unbounded weights, piecewise-linear networks obey a sqrt(n/m) robustness floor up to one log","key_machinery":"The rigidity lemma: after rewriting a two-layer piecewise-linear network into canonical ReLU-kink form on the ball or sphere, each kink coefficient is bounded by a multiple of the Euclidean Lipschitz constant of the realized function, because kinks supported on distinct hyperplanes cannot cancel at a generic point of one kink set.","core_discovery":"For data uniform on the sphere (d greater than or equal to 3) or standard Gaussian, labels in [-1,1] with positive noise level, and any continuous piecewise-linear activation, every width-m two-layer network with completely unrestricted real weights that fits epsilon below the noise floor satisfies Lip(f) at least c epsilon times the square root of n over (m-bar times a log of m-bar n d over epsilon), with high probability. The same lower bound holds with m-bar replaced by the number of realized distinct kink hyperplanes plus one.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Two-layer nets need √(n/m) Lip to fit noise, even with free weights","Unbounded two-layer nets still force Lip ≥ cε√(n/(m log))","Kink rigidity yields √(n/m) robustness for any-weight two-layer nets","Piecewise-linear two-layer fits below noise need Lip √(n/m) up to log","Realized kinks bound Lip of unrestricted two-layer nets fitting noise"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The argument needs that kinks on different hyperplanes cannot cancel at a generic point of one kink set; that geometric fact holds on the ball and on spheres of dimension at least three, but fails on the circle.","fun_headline_variants_meta":{"raw":{"variants":["Two-layer nets need √(n/m) Lip to fit noise, even with free weights","Unbounded two-layer nets still force Lip ≥ cε√(n/(m log))","Kink rigidity yields √(n/m) robustness for any-weight two-layer nets","Piecewise-linear two-layer fits below noise need Lip √(n/m) up to log","Realized kinks bound Lip of unrestricted two-layer nets fitting noise"]},"model":"grok-4.5","effort":"low","cost_usd":0.004858,"raw_usage":{"total_tokens":1495,"prompt_tokens":1028,"num_sources_used":0,"completion_tokens":119,"cost_in_usd_ticks":48580000,"prompt_tokens_details":{"text_tokens":1028,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":348,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":1028,"tokens_out":119,"duration_ms":4548,"temperature":1.0,"reasoning_tokens":348,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T18:17:23.700166+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Exhibit a continuous piecewise-linear two-layer network of width m, with arbitrary real weights, that fits generic sphere or Gaussian data epsilon below the noise floor while keeping Lipschitz constant o(epsilon sqrt(n/(m log(m n d)))) in dimension d greater than or equal to 3.","supporting_citations":[],"review_version":1}