{"id":"a00c40b4-581a-4242-a093-1e574e90351a","arxiv_id":"2608.04060","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"SJEPA learns latent predictive states whose transitions are compact symbolic laws with a regularized neural residual, and demonstrates simpler, less divergent pendulum dynamics than post-hoc symbolic fitting.","lead":"This paper introduces SJEPA, a variant of joint-embedding predictive architectures in which the latent transition is written as a compact symbolic equation plus a small neural correction, instead of a fully opaque neural map. It shows in controlled pendulum experiments that jointly learning the representation and the symbolic dynamics yields simpler and more stable latent laws than fitting symbols after the fact, at the cost of some predictive accuracy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The cross-model rollout comparison rests on an affine decoder that explains only R2=0.293 of SJEPA's test latents; the claimed improvement may reflect decoder misfit rather than superior latent dynamics.","rationale":"Reader's weakest assumption is exactly this affine-probe issue; I agree it is the most load-bearing. The central claim rests on the physical-state rollout comparison, which is the only cross-model metric that connects learned latent dynamics to the observed system. Since SJEPA deliberately searches over coordinates, there is no reason its coordinates must remain linearly aligned with q,p; indeed Table 12 shows they are not. The affine decoder is therefore a metric chosen for convenience, not a property of the representation. A nonlinear decoder could change the ranking. The paper's own caveat about interpretation ('not evidence that operator compression preserves a simple affine correspondence') does not repair the metric used for the headline claim. Other concerns—small seed count, hand-chosen grammar, no released code—are secondary and acknowledged, and the theoretical propositions are sound; they do not independently threaten the central empirical claim as directly as the decoder issue. Since the reader already made acceptance conditional on this issue, the verdict remains CONDITIONAL; no change is warranted unless the proposed test fails, in which case the empirical support for the central claim would be materially weakened.","tokens_in":34253,"tokens_out":6223,"duration_ms":76968,"concrete_test":"Recompute Table 4 using a nonlinear state decoder, e.g., an MLP with two hidden layers (or a Gaussian-process regression), fitted only on training latents to (q,p), for all Experiment 1 models, and report decoder R2 on held-out states alongside rollout MSE. Match decoder capacity across models and test at least two capacities. If SJEPA's test/OOD rollout advantage over post-hoc symbolic shrinks or reverses under the better-fitting decoder, then the headline improvement is an artifact of affine alignment; if the advantage persists and the decoder R2 is comparable across models, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that joint operator compression yields simpler and more adequate latent dynamics—is supported by Table 4's physical-state rollout MSE. Those numbers are computed in Appendix D.3 by fitting an affine map from training latents to (q,p) and applying it to recursive latent rollouts. Table 12 shows the test affine-probe R2 for SJEPA is 0.293±0.154, versus 0.814±0.063 for Neural JEPA / post-hoc symbolic coordinates. Thus the decoder used for the cross-model comparison is far poorer in SJEPA's coordinate system. Because the affine decoder is not a valid projection of the learned coordinates, the comparison in Table 4 (test rollout 0.467 vs 1.229; OOD 3.435 vs 9.262) conflates latent dynamics quality with how well q,p happen to be linearly decodable from each representation. The paper acknowledges the lower alignment (Section 8.1, Table 12) but continues to use affine-aligned rollout as the head-to-head metric. If SJEPA's coordinates are related to the physical state nonlinearly, the reported rollout errors are potentially dominated by decoder misspecification rather than by the learned transition; the claimed 'adequately predictive' property is therefore not cleanly established in observable coordinates.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"SJEPA introduces a joint-embedding predictive architecture in which the latent transition is a hybrid of a compact symbolic law and a regularized neural correction. The paper's central principle is operator compression: among informative, non-collapsed predictive representations, select coordinates whose induced dynamics have minimal symbolic-plus-correction complexity subject to predictive adequacy. The theory formalizes induced-dynamics complexity, proves that predictive coordinates and the symbolic/correction decomposition are non-identifiable from prediction loss alone, identifies a collapse shortcut created by unconstrained operator compression, derives the optimal correction under quadratic regularization, and bounds rollout error under uniform one-step adequacy. Two controlled pendulum experiments are reported: Experiment 1 compares jointly learned symbolic coordinates against post-hoc symbolic regression on frozen Neural JEPA coordinates, reporting lower symbolic complexity (mean Ω 4.67 vs 26.0) and lower physical-state rollout MSE; Experiment 2 studies grammar misspecification and the effect of correction regularization on symbolic–neural allocation. The paper is careful to state limitations and to separate predictive accuracy from parsimony claims.","tokens_in":34508,"tokens_out":4528,"duration_ms":55560,"significance":"If the central empirical comparison were clean, SJEPA would be a valuable contribution: it turns transition complexity into an explicit learning signal inside a reconstruction-free JEPA, with clearly stated propositions and proofs that are elementary but appropriate. The controlled experimental design, disjoint train/validation/test/OOD splits, per-seed reporting, and explicit acknowledgement of scope are strengths. The paper is also unusually honest about the trade-off between symbolic parsimony and raw predictive accuracy. However, the main empirical claim of lower physical-state rollout error currently rests on an affine decoder that is nearly uninformative for the jointly learned coordinates, so the headline comparison is not yet established in observable coordinates. This makes the significance conditional on a revised evaluation.","major_comments":[{"comment":"The central empirical comparison—joint SJEPA reducing physical-state rollout MSE from 1.229 to 0.467 relative to post-hoc symbolic regression—is confounded by the affine alignment map used for evaluation. Table 12 reports test affine-probe R² = 0.293±0.154 for SJEPA, versus 0.814±0.063 for Neural JEPA/post-hoc coordinates. Because the affine decoder explains only about 29% of the variance of SJEPA's test latents, the reported physical-state rollout errors for SJEPA are dominated by decoder misspecification rather than by the quality of the learned latent transition. The manuscript acknowledges this lower alignment in Section 8.1 but continues to use affine-aligned rollout as the head-to-head metric in Table 4 and in the abstract/conclusion claims of 'lower physical-state rollout error.' This is load-bearing because the physical-state rollout comparison is the quantitative basis for the claim that joint learning discovers more adequate dynamics, not merely simpler latent equations. Please either (a) evaluate rollout error through a matched-capacity nonlinear decoder trained to minimize physical-state MSE for all models, (b) use a coordinate-invariant evaluation such as mapping rolled-out latents back to observation space through the known forward observation model, or (c) reframe the empirical claim as concerning complexity of induced latent dynamics only, treating physical-state rollout numbers as illustrative and clearly subordinate to the affine-probe caveat. As written, the claimed cross-model rollout improvement is not a clean comparison of latent dynamics quality.","section":"Section 8.1, Table 4, Appendix D.3, Table 12"}],"minor_comments":[{"comment":"The sentence discussing the two coupled questions ends with a stray '13?' marker that appears to be a leftover reference callout; please remove it.","section":"Section 9, first paragraph"},{"comment":"Equation references such as '(Eq.64)' and '(Eq.81)' mix textual labels with equation numbers; please use one consistent numbering convention throughout.","section":"Section 8.2 and Appendix D"},{"comment":"The entries 'Posner et al. 2026a' and 'Posner et al. 2026b' have identical titles and identical URLs; if they are the same work, list it once, and if they are different works, provide distinct titles or identifiers.","section":"References"},{"comment":"The terminology for the unregularized one-step condition is inconsistent: Table 3 calls it 'one-step collapse diagnostic' while the text and Table 5 refer to a 'no-RIB' condition. Please align the terminology.","section":"Section 8.1, Table 3 and Table 5"},{"comment":"Because latent MSE is coordinate-dependent, consider adding a footnote or a 'within-coordinate' qualifier directly in the table header to discourage cross-model comparison of this column.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a single-author preprint with a number of self-citations to closely related JEPA works; those citations appear relevant and do not raise concern. The main issue is the empirical evaluation metric described in the major comment. If the authors replace or corroborate the affine-aligned rollout comparison with an evaluation that is valid for nonlinearly transformed coordinates, the paper's contribution would be publishable. I do not see a need to question the theoretical propositions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Yongchao's SJEPA paper is worth your time if you care about making latent dynamics interpretable. The central idea—using the complexity of the induced transition as an explicit criterion for representation learning, with a guard against collapse—is genuinely new and is backed by clean, simple theory. The collapse-shortcut proposition is a nice observation, and the hybrid symbolic-neural decomposition is well motivated. The experiments are carefully done and honestly reported; Experiment 2 in particular is a nice controlled study of grammar misspecification. The paper doesn't oversell: it says Neural JEPA remains more accurate and that the method trades predictive power for interpretability.\n\nThe soft spot is the headline empirical comparison. The physical-state rollout numbers in Table 4 are computed through an affine map fitted from latents to (q,p). For SJEPA, that affine probe has test R2=0.293, versus 0.814 for the Neural JEPA/post-hoc coordinates (Table 12). So the decoder used for the comparison is much worse in SJEPA's coordinate system. The paper acknowledges this low alignment but still presents the rollout improvement over post-hoc as a main result. That is a real confound: the error could come from decoder misspecification rather than from better latent dynamics. The right fix is to either use a more flexible decoder, evaluate in a coordinate-invariant manner, or demote the affine-rollout numbers to secondary while making the latent-space and symbolic-complexity comparisons primary. This isn't fatal, but it needs to be addressed before the rollout claim is taken at face value.\n\nThe other limitations are scope and reproducibility: one pendulum, three seeds, a hand-chosen grammar, and no code or data release. The paper itself flags the narrow validation, which is to its credit, but it does limit how much we can infer.\n\nOverall, this is a solid, honest paper with a novel idea and correct theory. I'd send it to peer review, but the experiments need the decoder issue fixed and ideally code/data released. I'd bring it to a reading group and would cite the collapse analysis in my own work.","headline":"SJEPA has a genuinely new idea and sound theory, but its headline rollout comparison is confounded by a poorly fitting affine decoder and needs revision before the empirical claims are convincing.","tokens_in":35038,"tokens_out":2774,"would_cite":true,"duration_ms":29498,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Joint symbolic–neural learning discovers latent dynamics 5.6 times simpler than post-hoc symbolic regression.","keywords":["joint-embedding predictive architecture","symbolic regression","latent dynamics","operator compression","representation collapse","hybrid symbolic-neural model","governing equation discovery","self-supervised learning"],"falsifier":"Run SJEPA and the post-hoc symbolic baseline on a system with known simple latent dynamics under an observation map where the simple coordinates are nonlinearly mixed into observations, then evaluate rollouts through an invertible nonlinear decoder rather than an affine probe; if SJEPA's complexity and rollout advantages vanish under the correct alignment, the operator-compression claim would be refuted. Alternatively, a direct search for the collapsed solution in the unconstrained one-step objective, checking whether every seed reaches a near-constant representation with identity dynamics, would confirm or refute the collapse-shortcut prediction.","tokens_in":34004,"feed_emoji":"🧮","tokens_out":6564,"duration_ms":60898,"temperature":0.7,"pith_summary":"The paper proposes SJEPA, a reconstruction-free joint-embedding predictive architecture whose latent transition is a hybrid of a compact symbolic law and a regularized neural correction. Its central claim is that operator compression—minimizing the complexity of the symbolic-plus-correction transition subject to predictive adequacy and non-collapse constraints—selects latent coordinates whose induced dynamics are simpler and remain adequately predictive. In a controlled pendulum task, jointly learning representation and symbolic dynamics reduced mean weighted symbolic complexity from 26.0 to 4.67 and physical-state test rollout error from 1.229 to 0.467 relative to fitting symbols post hoc to frozen predictive coordinates. The paper also shows that without representation constraints, unconstrained operator compression drives the representation to collapse to a near-constant state, making the transition trivially simple. A second experiment shows that correction regularization keeps a representable symbolic mechanism intact under grammar misspecification while directing the neural component to residual dynamics.","feed_headline":"Joint symbolic learning cuts latent-dynamics complexity 5.6-fold","feed_subtitle":"SJEPA shows that compressing the transition, not the representation, yields simpler and less divergent rollouts.","key_machinery":"The load-bearing object is operator compression realized through the induced-dynamics complexity functional. Given encoders, $C_{\\mathrm{dyn}}(E_\\theta,E_{\\bar\\theta}) = \\inf_{E,\\alpha,\\phi}\\{\\Omega(E)+\\lambda_c R_{\\mathrm{corr}}(\\phi) : L_{\\mathrm{pred}} \\leq \\delta_{\\mathrm{pred}}\\}$ asks how simple the latent transition can be once coordinates are fixed; the outer problem $\\min_{(\\theta,\\bar\\theta)\\in\\Theta_{\\mathrm{repr}}} C_{\\mathrm{dyn}}$ selects coordinates whose dynamics are easiest to describe. The hybrid predictor $H=F+c_\\phi$ combines a typed symbolic expression with structural complexity $\\Omega(E)$ and a neural correction regularized by $R_{\\mathrm{corr}}(\\phi)=\\mathbb{E}[\\|c_\\phi\\|_2^2]+\\tau_\\phi\\|\\phi\\|_2^2$, which controls how much predictive structure is delegated to the correction. Four propositions carry the argument: predictive coordinates are non-identifiable under bijections; unconstrained compression has a collapsed global minimizer; the symbolic–neural decomposition is non-identifiable; and correction regularization shrinks the correction toward the conditional-mean residual by the factor $(1+\\lambda)^{-1}$. The practical objective is a scalarized constrained problem training one-step prediction, rollout, representation quality, symbolic parsimony, and correction control jointly.","core_discovery":"The paper's core discovery is that the complexity of the induced transition operator can serve as a criterion for learning predictive representations: among informative, non-collapsed coordinate systems, SJEPA selects coordinates whose evolution admits the simplest adequate symbolic law. Formally, for fixed encoders the induced-dynamics complexity $C_{\\mathrm{dyn}}(E_\\theta, E_{\\bar\\theta})$ is the minimum symbolic-plus-correction complexity needed to meet a predictive tolerance, and representation learning minimizes this quantity over admissible encoder pairs. The hybrid predictor $H_{E,\\alpha,\\phi}(Z_C,\\varepsilon)=F_{E,\\alpha}(Z_C,\\varepsilon)+c_\\phi(Z_C,\\varepsilon)$ acts on the transition operator rather than partitioning the representation. Predictive coordinates are shown to be non-identifiable up to arbitrary bijections, so operator complexity supplies the extra criterion; unconstrained compression is shown to admit a degenerate global solution where both encoders map everything to one point and an identity transition predicts perfectly. The empirical results support the operator-compression claim: joint learning discovers oscillator-like equations across seeds, lowers rollout divergence, and the unregularized diagnostic reaches the predicted collapse shortcut.","pith_inferences":["The paper does not claim that operator complexity is a general model-selection principle, but the same constrained objective could plausibly regularize latent-dynamics models beyond JEPA whenever inspectable equations are desired.","An untested extension: replacing the affine alignment probe in Experiment 1 with a nonlinear invertible decoder would show whether the rollout advantage of SJEPA persists under a fairer coordinate comparison.","The non-identifiability result implies that reported symbolic complexity is convention-dependent; publishing the grammar, coefficient threshold, and latent normalization alongside the discovered equation would make complexity comparisons reproducible across methods.","The action-conditioned planning interface is provided but not evaluated; a natural next step is a control task where the discovered law reveals which latent directions are action-sensitive."],"forward_implications":["If SJEPA's principle holds, latent coordinates learned under operator compression will expose inspectable governing equations rather than black-box maps, with comparable or better recursive stability than post-hoc symbolic fits.","The collapse shortcut implies that any method that adds an equation-simplicity penalty to a JEPA-style objective must pair it with explicit non-collapse constraints; otherwise the simplest 'law' is the identity on a near-constant representation.","Correction regularization yields a controllable symbolic–neural allocation: the symbolic component retains representable mechanisms and the neural component tracks residual dynamics, with the trade-off visible in the normalized correction-energy ratio.","A complete symbolic grammar can remove the need for a neural correction in controlled settings, recovering the true vector field to within about 7% coefficient error.","Long-horizon accuracy is a separate requirement from one-step adequacy: rollout error grows through the learned transition's Lipschitz constant, so compactness alone does not guarantee non-divergence."],"supporting_citations":[{"why":"Establishes the joint coordinate–equation discovery precedent that SJEPA extends to a reconstruction-free JEPA context–target setting.","marker":"Champion et al., 2019"},{"why":"Provides the I-JEPA latent-prediction architecture whose opaque neural transition SJEPA replaces with a symbolic-plus-correction predictor.","marker":"Assran et al., 2023"},{"why":"Supplies the sparse-identification library machinery used for the differentiable sparse-library implementation of the symbolic law.","marker":"Brunton et al., 2016b"},{"why":"Supplies the VICReg-style variance/invariance/covariance regularizer used to enforce non-collapse in Experiment 1.","marker":"Bardes et al., 2022"},{"why":"Represents the alternative of distilling symbolic models from trained neural predictors, the post-hoc style SJEPA is compared against.","marker":"Cranmer et al., 2020"},{"why":"Recent joint latent-coordinate and symbolic-dynamics recovery with identification guarantees, the closest contemporary baseline for coordinate search.","marker":"Muratore & Mathis, 2026"},{"why":"Frames JEPA as predicting representations rather than reconstructing observations, the architectural premise SJEPA builds on.","marker":"LeCun, 2022"}],"fun_headline_variants":["SJEPA: simplest adequate latent dynamics via hybrid symbolic fit","5.6x simpler transitions with symbolic-neural co-learning","Elegant dynamics: SJEPA compresses operators, not reps","SJEPA: symbolic laws for latent rollouts, collapse averted"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main load-bearing premise is that the affine map fitted from training latents to the physical state is a fair way to compare rollout quality across models, even though for the jointly learned SJEPA coordinates that affine probe explains only about 29% of test variance; if that alignment assumption fails, the reported rollout improvement over post-hoc fitting is not a clean comparison of latent-dynamics quality.","fun_headline_variants_meta":{"raw":{"variants":["SJEPA: simplest adequate latent dynamics via hybrid symbolic fit","5.6x simpler transitions with symbolic-neural co-learning","Elegant dynamics: SJEPA compresses operators, not reps","SJEPA: symbolic laws for latent rollouts, collapse averted"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1284,"prompt_tokens":1006,"completion_tokens":278,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":204}},"tokens_in":622,"tokens_out":278,"duration_ms":3639,"temperature":1.0,"reasoning_tokens":204,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T00:45:22.682702+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run SJEPA and the post-hoc symbolic baseline on a system with known simple latent dynamics under an observation map where the simple coordinates are nonlinearly mixed into observations, then evaluate rollouts through an invertible nonlinear decoder rather than an affine probe; if SJEPA's complexity and rollout advantages vanish under the correct alignment, the operator-compression claim would be refuted. Alternatively, a direct search for the collapsed solution in the unconstrained one-step objective, checking whether every seed reaches a near-constant representation with identity dynamics, would confirm or refute the collapse-shortcut prediction.","supporting_citations":[],"review_version":1}