{"id":"1bbcb2fc-4626-422d-96f6-72db28861967","arxiv_id":"2412.14695","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"LResNet performs hyperbolic residual connections with a normalized weighted sum (the Lorentzian centroid), avoiding tangent-space mappings and improving efficiency and accuracy in hyperbolic GNNs, graph transformers, and CNNs.","lead":"Researchers introduce LResNet, a way to add residual connections inside hyperbolic neural networks by normalizing a weighted average of two points directly on the curved space, avoiding expensive tangent-space round trips. It is a general plug-in for GNNs, graph transformers, and CNNs that claims faster, more stable training and better accuracy on hierarchical data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prop 4.2's parallel-transport proof is internally inconsistent: the coefficient of x_s in z_s is cosh(α)+(sinh(α)/α)c_v, but the proof sets w_x=cosh(α)+c_v, so the claimed collinearity does not follow; positivity of w_x,w_y is also unverified.","rationale":"The reader's weakest_assumption identifies the right region of the argument: Prop 4.2 gives only spatial proportionality, not equality, and the appendix does not verify w_x,w_y in R+. My stress-test sharpens this into an actual algebraic error in the proof: the claimed choice of w_x cannot produce the stated collinearity because the sinh(α)/α factor is dropped from the c_v term. This is a concrete, checkable flaw in the paper's main theoretical contribution, rather than merely a gap in exposition. It does not, however, undermine the core empirical proposal: Eq 8 is a simple, commutative, numerically stable residual operation, the experiments are extensive, and the parallel-transport example in Theorem 4.3 even shows a case where positive LResNet weights reproduce the previous output exactly. The appropriate verdict remains CONDITIONAL: the method and experiments are credible, but the theoretical 'derives previous methods' claim needs a corrected proof, explicit positivity verification, and a clear statement of whether equality requires the optional scaling of Eq 10. Since this matches the reader's CONDITIONAL verdict, no verdict change is needed; I record partial agreement because my concern is a more specific internal inconsistency than the reader's general gap.","tokens_in":21091,"tokens_out":17758,"duration_ms":158771,"concrete_test":"Re-derive Prop 4.2(a) symbolically from Eq 5, then test the corrected collinearity condition numerically: for K=-1, sample random pairs x,y in L^{-1,n}, compute z = x⊕_P y, and solve for r = w_y/w_x > 0 such that the Klein coordinate of z is a scalar multiple of (w_x x_s + w_y y_s)/(w_x x_t + w_y y_t). Also evaluate the appendix's candidate weights w_x = cosh(α)+c_v, w_y = (sinh(α)/α)c_u (and their corrected versions) and check whether they satisfy the proportionality equation and are nonnegative. If no positive solution exists, or if the candidate weights fail the equation, Proposition 4.2 is false or unproved, and the claim that LResNet derives previous methods must be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central theoretical claim is that LResNet 'theoretically derives' the parallel-transport, tangent-space, and space-addition residual methods (Section 4.1, Prop 4.2). The load-bearing step is Prop 4.2(a), which asserts that for parallel-transport output z there exist weights w_x,w_y in R+ such that the LResNet output lies on the geodesic from o to z. As written, the proof is algebraically inconsistent: from z = cosh(α)x + sinh(α)/α (c_u y' + c_v x'), the space component is z_s = [cosh(α) + (sinh(α)/α)c_v] x_s + (sinh(α)/α) c_u y_s, but the proof instead sets w_x = cosh(α)+c_v and w_y = (sinh(α)/α)c_u. Unless c_v = 0 or α = 0, these weights do not make z_s proportional to w_x x_s + w_y y_s. The appendix also never proves w_x, w_y ≥ 0; c_v is negative in general, so positivity is a substantive gap, not a formality. Separately, even a valid spatial-proportionality result only places m on the same Klein-model ray as z; without applying the optional scaling of Eq 10 it does not give equality of outputs, so 'derive previous methods' overstates what Prop 4.2 actually establishes. The residual formula Eq 8 may still be useful and empirically effective, but the theoretical-universality claim currently rests on an unproven, possibly false, collinearity-plus-positivity assertion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes LResNet, a residual connection defined directly on the Lorentz hyperboloid by normalizing a weighted Euclidean combination of two hyperbolic vectors: x⊕_L f(x) = (w_x x + w_y f(x))/sqrt(-K|...|). It argues that this operation is commutative, numerically stable, and avoids tangent-space mappings, and that by suitable weights it can reproduce (or at least match along geodesics from the origin) the parallel-transport, tangent-space, and space-addition residual methods used in prior hyperbolic networks. Experiments apply LResNet to GNNs, graph transformers, and CNNs on graph and vision benchmarks, reporting accuracy gains over baselines and large speedups. The central theoretical claim is Proposition 4.2, which asserts that the LResNet output lies on the geodesic from the origin to the output of each prior method.","tokens_in":21475,"tokens_out":22167,"duration_ms":144270,"significance":"If the theoretical claim can be made rigorous, LResNet is a valuable unification: it is an O(n), commutative residual operation on the hyperboloid that subsumes several existing residual-connection designs, and the paper provides broad empirical validation across architectures and datasets plus a public code release. The experimental evidence is extensive and the speed advantage is concrete. However, the theoretical-universality claim is load-bearing for the paper's framing, and the current appendix proof contains algebraic inconsistencies, an incorrect logarithmic-map expression, and an unverified positivity condition. The core formula may still be useful empirically, but the 'derives previous methods' assertion is not established as written.","major_comments":[{"comment":"The proof of Proposition 4.2(a) is algebraically inconsistent. From z = cosh(α)x + (sinh(α)/α)(c_u y' + c_v x'), the space component is [cosh(α) + (sinh(α)/α)c_v] x_s + (sinh(α)/α)c_u y_s, so the stated choice w_x = cosh(α) + c_v and w_y = (sinh(α)/α)c_u yields collinearity only if sinh(α)/α = 1 or c_v = 0. The correct candidate would be w_x = cosh(α) + (sinh(α)/α)c_v; as printed, the conclusion that z_s is proportional to w_x x_s + w_y y_s does not follow.","section":"§4.1, Prop. 4.2(a), Appendix A"},{"comment":"The appendix writes log_o(y) as c_u(y + y_t sqrt(-K)o), but Eq. (4) with u = o gives log_o(y) = c_u(y - y_t sqrt(-K)o) = c_u[0, y_s]^T. With the plus sign the 'tangent' vector is not in T_o L, so the parallel-transport step P_{o→x} is applied to a vector outside the tangent space. This also makes the proof of Theorem 4.3 inconsistent with its own worked example: following the proof's definitions for x = [3,2,-2]^T and y = [3,2,2]^T does not produce z = [9,8,-4]^T.","section":"Eq. (4), Prop. 4.2(a), Thm. 4.3 proof"},{"comment":"Even after repairing the algebra, Proposition 4.2 asserts only that m lies on the geodesic from o to z, not that m = z. The text then concludes that LResNet 'can theoretically derive previous methods' and has at least their representative power; this requires either exact equality, which might be obtained through the optional scaling of Eq. (10), or a formal argument that ray-collinearity with the origin preserves expressive power across subsequent layers. Neither is provided. The proof also never verifies w_x, w_y ∈ R+: since c_v is generically negative, positivity is a substantive condition rather than a formality.","section":"Prop. 4.2, §4.1, Eq. (10)"},{"comment":"The tangent-space case is not actually proved: the text says 'one can check' and then writes c_1 = cosh^{-1}(-x_t sqrt(-K)) / sqrt(x_t^2 K - 1), which is not real for K < 0 and mismatches the c_u used in part (a). A complete proof with correct coefficients and positivity verification is needed for this case as well.","section":"Prop. 4.2(b), Appendix A"}],"minor_comments":[{"comment":"The Klein-model isometry is written as φ_K(x) = x_t/x_s; it should be φ_K(x) = x_s/x_t, since collinearity is expressed through equality of space-over-time ratios.","section":"§3, Appendix A"},{"comment":"The proof uses ||·||_L ambiguously: the first equality treats -K||w_x x + w_y y||_L^2 as the signed Lorentzian inner product, while the lemma statement uses the absolute value. The proof should explicitly note that w_x x + w_y y is future timelike, so the absolute value is redundant; with that clarification the inequality appears correct.","section":"Lemma 4.1, Appendix A"},{"comment":"The denominator sqrt(-K | ||w_x x + w_y f(x)||_L |) contains a redundant absolute value because the norm already includes one; define the denominator unambiguously.","section":"Eq. (8)"},{"comment":"There are several typos and one incomplete citation: '[53?]' should be a proper reference, 'Sqirrel' should be 'Squirrel', and 'detials', 'Riemmanian', and 'ODD-detection' should be corrected.","section":"Related Works and throughout"},{"comment":"The claim of 'over 2000 times speedup' holds only for the right-hand column (4096/100,000); the left column shows roughly a 14x speedup. Please state this qualification explicitly.","section":"Table 6, §5.4"}],"recommendation":"major_revision","confidential_remarks":"The empirical portion of the paper is solid and the proposed residual operation is simple enough to be useful, but the theoretical section as printed is not reliable: the appendix proof of Proposition 4.2 contains multiple algebraic and definitional errors. I recommend major revision rather than rejection because the errors look repairable: the correct coefficient is w_x = cosh(α) + (sinh(α)/α)c_v, and positivity appears to follow from |c_v| < α, although the authors must supply the argument. If the repair fails, the 'theoretically derive previous methods' claim should be weakened to a statement about geodesic rays or removed. The editor may also want to confirm the novelty relative to the Lorentzian centroid literature, since the operation is a weighted generalization of the centroid."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the residual formula is simple and likely useful, and the experiments are solid; but the paper's headline theoretical claim—that LResNet 'derives' previous residual methods—is not supported by the proof as written. Proposition 4.2(a) has a genuine algebraic gap: the space component of the parallel-transport output is z_s = [cosh(α)+(sinh(α)/α)c_v] x_s + (sinh(α)/α)c_u y_s, yet the proof sets w_x = cosh(α)+c_v. Unless α=0 or c_v=0, the claimed collinearity doesn't follow. The positivity of the constructed weights is also never verified; c_v is negative in general, so w_x may not lie in R+. This is load-bearing, because the abstract and Section 4.2 tell the reader the method 'theoretically derives' previous methods.\n\nWhat is genuinely new: using the weighted Lorentzian centroid as a residual connection. Equation (8) is elegant, commutative, avoids tangent-space round-trips, and Lemma 4.1 (the stability bound) is actually correct, though the appendix proof is needlessly hard to follow. The experiments are extensive and transparently reported: GNNs, graph transformers, CNNs, OOD robustness, runtime, and over-smoothing. Gains are modest but consistent, and the speedup of the residual operation is real.\n\nThe main overreach is the 'derivation' claim. Even a repaired Proposition 4.2 would only place the LResNet output on the same Klein-model ray as the previous method's output—same direction, not same point. Adding the optional scaling (Eq 10) could slide along the ray to match norms, but the paper doesn't explicitly make that argument. The correct statement is 'same geodesic ray up to scaling,' not 'derives previous methods.'\n\nThe empirical method stands on its own. For a fresh submission I'd ask for a corrected Proposition 4.2 and a softened theoretical claim before accepting; the residual block itself is worth publishing. This is a solid paper with one load-bearing theoretical flaw, not a hollow one.","headline":"Simple, useful residual block for Lorentz hyperbolic nets, but the 'derives previous methods' claim needs a repaired proof and a softer statement.","tokens_in":21955,"tokens_out":4859,"would_cite":true,"duration_ms":34856,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LResNet replaces tangent-space round trips with one normalized sum, giving hyperbolic networks a commutative, stable residual connection.","keywords":["Residual connections","Hyperbolic neural networks","Lorentz model","Lorentzian centroid","Numerical stability","Graph neural networks","Computer vision"],"falsifier":"Take any pair of hyperboloid points and each previous residual method's output, and solve Eq. (8) for the weights that make LResNet's output exactly equal to that output; if no positive weights achieve this, the strong 'derives previous methods' reading is ruled out, consistent with the proof's weaker ray-alignment promise. In addition, if trained LResNet weights in the published experiments converge to negative values, the positivity premise behind the stability and derivation arguments would be violated.","tokens_in":20888,"feed_emoji":"📐","tokens_out":12766,"duration_ms":91057,"temperature":0.7,"pith_summary":"In hyperbolic space, the Euclidean trick of adding a layer's input to its output does not work, because the sum can leave the curved manifold. Existing hyperbolic residuals therefore map to a tangent space, add or transport there, and map back, which this paper argues is slow, non-commutative, numerically unstable, and error-prone. LResNet replaces that round trip with a single normalized weighted sum on the Lorentz hyperboloid: the Euclidean combination $w_x x + w_y f(x)$ is divided by its Lorentzian magnitude so the result stays on the manifold. The authors prove this operation is commutative and stable, and they claim that, with suitable weights, it reproduces the geodesic-ray direction of every prior hyperbolic residual construction, giving it at least their representational power. If the claim holds, any Lorentz-model hyperbolic network gets a drop-in residual block that is faster, simpler, and more stable than existing options, with experiments on graphs, graph Transformers, and images supporting that conclusion.","feed_headline":"One normalized sum replaces tangent-space residuals","feed_subtitle":"Hyperbolic residual nets add directly on the hyperboloid, avoiding the slow and unstable tangent-space detour.","key_machinery":"The central object is the weighted Lorentzian centroid, written as a normalized weighted sum in Eq. (8). For two hyperboloid points $x$ and $f(x)$ and positive scalar weights $w_x$, $w_y$, the Euclidean combination $w_x x + w_y f(x)$ is renormalized by its Lorentzian magnitude, which projects it back onto the hyperboloid; this generalizes the Lorentzian centroid used for aggregation, with weights and a curvature-dependent denominator. The operation does the work of a residual connection entirely on the manifold in $O(n)$ time, and because only the ratio $w_x/w_y$ matters, one weight can be fixed to a positive constant while the other is trained. An optional scaling step slides the output along a Klein-model geodesic to control its Euclidean norm, preserving the ray-alignment property used in the expressiveness argument.","core_discovery":"The authors' central claim is that the operation in Eq. (8), $x\\oplus_L f(x) = (w_x x + w_y f(x))/\\sqrt{-K|\\|w_x x + w_y f(x)\\|_L|}$, is a valid residual connection on the Lorentz hyperboloid: the normalizing denominator forces the output back onto the manifold, so no tangent-space or exponential-map round trip is required. They prove (Lemma 4.1) that the denominator is bounded below by $\\sqrt{w_x^2 + w_y^2}$, ruling out division-by-zero blowups, and they show the operation is commutative because it is a normalized weighted sum. On representational power, they argue in Proposition 4.2 that for each earlier residual construction — parallel transport, tangent-space addition, and space-like addition — there are nonnegative weights making the LResNet output lie on the same geodesic ray (in the Klein model, the same straight line through the origin) as that construction's output, so LResNet can match the expressive power of all of them. Empirically, LResNet used as a drop-in residual block outperforms the previous residual methods on node classification, link prediction, and image classification, and it avoids the NaN failures that parallel transport suffers in deep networks. The authors position LResNet as a generally applicable residual module for any Lorentz-model hyperbolic network.","pith_inferences":["Editorial inference: The ray-alignment proof does not imply pointwise equality, so if downstream layers depend on exact position or off-ray details, LResNet's outputs may differ from the methods it is said to derive; a layer-by-layer output comparison would clarify this.","Editorial inference: Because the normalized weighted sum is commutative and $O(n)$, it is a natural candidate for a general hyperbolic aggregation or attention operator, not only a residual connection.","Editorial inference: The speed advantage suggests the dominant cost of prior methods is the log/exp/parallel-transport round trip; at higher dimensions this makes LResNet the practical route to large-scale hyperbolic embeddings, a regime not directly benchmarked here.","Editorial inference: The stability lemma depends on positive weights, and the implementation takes absolute values in practice; monitoring the sign of trained weights in the reported Transformer and vision experiments would test whether the theoretical positivity condition is actually met in trained models."],"forward_implications":["Any Lorentz-model hyperbolic network can add residual connections by normalization alone, avoiding tangent-space and parallel-transport computations that the paper reports are over 2,000 times slower at scale.","Since LResNet reproduces the spatial ray of previous residual outputs, it can be substituted into existing GNNs, CNNs, and graph Transformers with at least the expressive power of the prior residual methods.","Deep GNNs using LResNet continue to gain from extra layers rather than degrading, while the parallel-transport baseline fails with NaN values at 16 layers or more.","The optional norm-scaling step allows control of embedding magnitude, which matters in vision models where norm is tied to classification confidence.","The same construction applies to any hyperbolic layer type on the Lorentz model, not only the convolution, GNN, and Transformer layers tested."],"supporting_citations":[{"why":"Defines the Lorentzian centroid that Eq. (8) generalizes with weights and curvature normalization.","marker":"[22]"},{"why":"Introduces the parallel-transport and tangent-space addition operations that are the main baselines and the targets of Proposition 4.2.","marker":"[3]"},{"why":"Proposes the parallel-transport residual connection and supplies the vision baselines and setup for the image experiments.","marker":"[39]"},{"why":"Introduces the space-like dimension addition residual method and the fully hyperbolic CNN architecture used as a baseline and adaptation target.","marker":"[2]"},{"why":"Provides the fully hyperbolic GNN base model and the linear and aggregation layers on which LResNet is tested.","marker":"[5]"},{"why":"Is the analogous Riemannian residual network whose parallel-transport construction LResNet claims to subsume.","marker":"[17]"},{"why":"Defines the Euclidean residual connection $x+f(x)$ that LResNet generalizes to hyperbolic space.","marker":"[16]"},{"why":"Shows that earlier midpoint computations in the Klein and Poincaré models equal the Lorentzian centroid, supporting the claim that LResNet extends across hyperbolic models.","marker":"[35]"}],"fun_headline_variants":["LResNet: direct residuals on the Lorentz hyperboloid","Skip the tangent space: residual via weighted centroid","One normalized sum unifies hyperbolic residual methods","Stable Lorentz residual: no exponential map needed","Residual connections without leaving the hyperboloid"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument's load-bearing premise is that sharing a ray from the origin is enough for one residual output to count as reproducing another; the proof only establishes that ray-sharing, not equality of outputs, and it does not verify that the required weights are positive.","fun_headline_variants_meta":{"raw":{"variants":["LResNet: direct residuals on the Lorentz hyperboloid","Skip the tangent space: residual via weighted centroid","One normalized sum unifies hyperbolic residual methods","Stable Lorentz residual: no exponential map needed","Residual connections without leaving the hyperboloid"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000273,"raw_usage":{"total_tokens":1668,"prompt_tokens":1013,"completion_tokens":655,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":582}},"tokens_in":629,"tokens_out":655,"duration_ms":6703,"temperature":1.0,"reasoning_tokens":582,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:00:24.722043+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any pair of hyperboloid points and each previous residual method's output, and solve Eq. (8) for the weights that make LResNet's output exactly equal to that output; if no positive weights achieve this, the strong 'derives previous methods' reading is ruled out, consistent with the proof's weaker ray-alignment promise. In addition, if trained LResNet weights in the published experiments converge to negative values, the positivity premise behind the stability and derivation arguments would be violated.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Lorentzian centroid that Eq. (8) generalizes with weights and curvature normalization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the parallel-transport and tangent-space addition operations that are the main baselines and the targets of Proposition 4.2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the space-like dimension addition residual method and the fully hyperbolic CNN architecture used as a baseline and adaptation target."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the analogous Riemannian residual network whose parallel-transport construction LResNet claims to subsume."}],"review_version":1}