Pith. sign in

REVIEW 3 major objections 4 minor 39 references

This paper claims that raising the multi-horizon consistency weight λ can push a passive-video latent transition into a near-contractive band, but the same knob fails to contract action-conditioned and natural-video domains.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 12:09 UTC pith:77PKGU4W

load-bearing objection A reproducible diagnostic study showing soft consistency contracts passive-video latent dynamics under an expansion proxy but not action-conditioned ones; the 'unifying' noise law is circular, but the core domain split is worth engaging. the 3 major comments →

arxiv 2607.21645 v1 pith:77PKGU4W submitted 2026-07-21 cs.LG cs.AI

Multi-Horizon Consistency as Geometry: When Latent Dynamics Contract, and When They Do Not

classification cs.LG cs.AI
keywords multi-horizon consistencylatent dynamicsexpansion proxyLipschitz contractionstochastic forcingworld modelslong-horizon predictionMoving-MNIST
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks what a common training term—multi-horizon latent consistency—actually does to the geometry of a learned latent transition. Using a finite-sample expansion proxy, L20,q95, it finds that on passively observed Moving-MNIST, raising the weight λ from 0 to 0.8 drops the proxy from about 4.96 to 1.01 and roughly halves horizon-20 error, with four of six seeds crossing the non-expansive line L=1. The same sweep does not produce population L<1 on action-conditioned control tasks or on KTH natural video, even when error improves. A stochastic-forcing law, L20 ≈ 1.23 + 1.82η, places control domains on the same curve at calibrated noise levels. The paper's usable claim is that soft consistency is a domain-limited geometry dial, not a universal contraction mechanism.

Core claim

The paper's central claim is that the multi-horizon consistency weight λ operates as a diagnosable geometry control: on passively observed Moving-MNIST, raising λ from 0 to 0.8 moves the 95th-percentile 20-step chord expansion L20,q95 from 4.96±2.01 to 1.01±0.06 and roughly halves horizon-20 error (0.365 to 0.177), with four of six seeds crossing below L=1. On action-conditioned Pendulum-v1 and CartPole-v1 and on real KTH video, the same sweep tightens the proxy or improves error without any population-level L<1 crossing. The authors read this as a passive/active boundary: soft multi-horizon agreement can act like implicit spectral regularization on a nearly deterministic, low-curvature late

What carries the argument

The central object is the multi-horizon latent consistency loss: the squared distance between free-run and teacher-forced latents at horizons {1,3,5,10,15,20}, weighted by λ and added to reconstruction and one-step losses. The measurement carrying the argument is the expansion proxy L20,q95, the 95th percentile of 20-step chord ratios ||f(20)(zi) − f(20)(zj)|| / ||zi − zj|| over held-out latent pairs; crossing the reference line L=1 is the paper's operational definition of a near-contractive band. A second mechanism, stochastic forcing, injects noise η during training and yields the linear expansion law that unifies the passive and control regimes.

Load-bearing premise

The load-bearing premise is that the measured 95th-percentile chord ratio L20,q95 faithfully reflects the expansion geometry that actually matters for long-horizon prediction; if the validation latents do not cover the rollout manifold, the chords are too long to be in the linear regime, or action sequences are not matched across pairs, the observed crossing below L=1 could be a sampling artifact rather than a property of the learned transition.

What would settle it

Rerun the Moving-MNIST λ sweep with chord pairs restricted to small separations (true linearization regime) and with action sequences explicitly matched; if L20,q95 no longer drops below 1 at λ=0.8, the threshold is an artifact of long-chord sampling. Separately, estimate the mean spectral radius of the Jacobian over many latents: if it stays above 1 while L20,q95<1, the proxy is not tracking true contraction.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • λ can be used as a geometry dial for passive-video latent models: monitoring L20,q95 gives a measurable operating point near 1, and paired seed tests at critical settings are needed before claiming a threshold.
  • The same λ should not be exported to action-conditioned control or natural video: the loss can improve long-horizon error there while L20,q95 remains above 1.
  • The stochastic-forcing law L20 ≈ 1.23 + 1.82η at λ=0.8 gives a portability rule: calibrate an effective noise level for a new domain and its expansion point is predicted, provided the domain falls on the same curve.
  • The negative latent-MPC result means contraction and planning utility are separable: L<1 on passive video is neither necessary nor sufficient for better control returns.
  • The associational mediation path λ→L→E (r̂≈0.94 on MMNIST) indicates geometry and error move together in the passive regime, but causal language is not licensed.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves an obvious next experiment implicit: randomize λ (or η) across training runs within a fixed architecture and re-estimate the mediation, which would turn the associational path into an intervention-based estimate.
  • The scaling results suggest a crisp testable extension: at d=16, MMNIST is already contractive at λ=0, so lowering latent dimension may substitute for the consistency loss; a joint sweep of d and λ would map how much of the threshold is due to the knob versus the representation's capacity.
  • The KTH exception implies the passive/active boundary may actually be an entropy/dimensionality boundary; a test would be an action-conditioned human-motion video, which should behave like control (no L<1) if actions are the operative factor, or like passive video if appearance statistics dominate.
  • If the forcing-law slope changes with architecture or residual scale, then the law is a property of the specific transition class; checking the same η sweep under a transformer or stochastic transition would tell whether the linear relation is universal.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies the effect of a multi-horizon latent consistency weight λ on the geometry of learned latent transitions. Using a finite-sample expansion proxy L20,q95 (the 95th percentile of 20-step chord-ratio distributions over 240 validation-pool pairs), the authors report that on Moving-MNIST raising λ from 0 to 0.8 reduces L20 from 4.96±2.01 to 1.01±0.06, with 4/6 seeds crossing L<1, and roughly halves horizon-20 error. On action-conditioned Pendulum-v1, CartPole-v1, and on KTH Actions, the same λ sweep tightens spectra or error but does not produce population L<1. To unify these domains, the authors inject stochastic forcing η during training on MMNIST, fit a linear law L20 ≈ 1.23+1.82η, and place Pendulum and CartPole on this curve via calibrated effective noise η_eff. Defensive experiments include architectural baselines, exogenous-digit stress, WorldTest-style scoring, latent MPC, latent-dimension scaling, and joint (λ,η) slices at λ∈{0.4,1.2}. The paper is framed honestly as a diagnostic study with explicit caveats about mediation being associational and L20 not being a certified Lipschitz constant.

Significance. If the central observation is robust, the paper makes a useful empirical contribution: practitioners often treat λ as a regularization knob, but evidence about whether it contracts or only reshapes latent geometry is scarce. The MMNIST threshold result, with per-seed tables, paired tests, and locked CSVs, is a commendable level of transparency and reproducibility for a diagnostic study. The negative results on control domains and KTH are also valuable because they prevent over-exporting contraction from passive video to action-conditioned settings. The stochastic-forcing law is a potentially elegant unifying picture, but, as detailed below, its current formulation is circular and therefore does not yet carry the claimed unification weight. The paper's explicit limitations section is thorough, and the authors avoid strong causal language. Overall, the empirical split between passive and active domains is a plausible and interesting phenomenon, but two load-bearing issues—the unverified coverage condition for L20,q95 and the circular calibration of η_eff—need to be addressed before the central claims can be considered established.

major comments (3)
  1. [§4.4, §6.1, Appendix B.6] The central contraction claim depends on L20,q95 measuring expansion on the rollout manifold, but Appendix B.6 states the statistic is informative only if validation latents cover that manifold. The paper never reports an overlap or drift statistic between teacher-forced validation latents and states visited by free-run f^k. If free-run rollouts drift into high-gain regions, L20,q95 can cross 1 without the learned transition being contractive where predictions are made. This is not a minor caveat: it is the load-bearing support condition for the MMNIST L<1 claim. Please compute L20,q95 on states actually visited by free-run rollouts, or report a quantitative overlap measure (e.g., nearest-neighbor distance from rollout states to the validation pool) for each domain.
  2. [§6.5, Figure 2, Appendix I] The stochastic-forcing law is fitted to MMNIST (L20 ≈ 1.23 + 1.82η), and η_eff for Pendulum, CartPole, and KTH is then solved from the observed L20 of those domains. The resulting “predicted” L-values (≈2.2 and 3.1) are identities by construction, not independent predictions. KTH is placed at η_eff ≈ 0 by construction even though its L20 remains expansive, which is treated as a residual rather than as a test of the law. To support the “one law” claim, the law must be tested out-of-sample: fit η_eff from an independent property of each domain (e.g., action noise or policy entropy) before seeing L20, or use a held-out domain for prediction. As written, the unification is a consistency check, not a validated law.
  3. [§4.4, §6.1, Table 1] The confirmatory statistics at λ=0.8 (paired t p=0.005, Wilcoxon p=0.031) are computed at an operating point selected post hoc from the λ sweep as the point where mean L20 first approaches 1 with multiple seeds below 1. No correction is made for this selection, so the reported p-values overstate confirmatory evidence. The MMNIST threshold should be presented as an exploratory finding, or the p-values should be adjusted for the number of λ values examined; a pre-registered or split-sample confirmatory design would be stronger. Without this, the “critical pair” framing is misleading.
minor comments (4)
  1. [Abstract] Typo: “doesnotproduce” should read “does not produce.” Also, the notation L20,q95 is typeset inconsistently across the paper (e.g., “L 20,q95” in some places).
  2. [§6.5] The label “predicted L-hat” for Pendulum/CartPole is confusing because those values are back-solved from the same observed L20 data used in the fit. Use “implied” or “calibrated” rather than “predicted” unless an out-of-sample procedure is added.
  3. [§10, Limitations] The sample sizes are small and the limitations section acknowledges this, but the paper would benefit from stating explicitly that the n=6 critical pair is the only seed configuration for the headline L<1 crossing; at other λ, n=3. This is not a blocker, but it affects the strength of the threshold claim.
  4. [References] Reference [26] (Srivastava et al., 2015) is cited for the GRU residual transition, but the original paper describes LSTMs; please clarify the architectural lineage or cite a more directly relevant source.

Circularity Check

1 steps flagged

The primary MMNIST/control split is empirical and self-contained, but the stochastic-forcing 'unification' predicts control-domain L20 values by reading off a curve fitted to MMNIST at 'calibrated' eta_eff values, so those predictions reduce to the fitted inputs.

specific steps
  1. fitted input called prediction [Section 6.5, 'Stochastic forcing: a linear expansion law'; Figure 2]
    "Mean L20,q95 increases approximately linearly, L20 ≈ 1.23 + 1.82η ... calibrated effective noise levels place Pendulum (ηeff ≈ 0.53) and CartPole (ηeff ≈ 1.0) at predicted L20 ≈ 2.2 and 3.1, consistent with population L>1 in those domains."

    The linear law is estimated on the same paper's MMNIST runs at λ=0.8. The control-domain L values are then 'predicted' by plugging ηeff into that same fitted line. No independent estimator of ηeff (e.g., action-space variance measured before seeing L20) is specified; ηeff is merely 'calibrated.' For CartPole, the reported L20≈3.17 at λ=0.8 inverts to ηeff≈1.07, and the fitted line gives ≈3.1, so the 'prediction' is the identity Lhat = fittedLaw(calibrate(observed L20)). The Pendulum number is not the inverse of the reported 1.74, so either the calibration is unstated or the prediction is inconsistent; in neither case is this an independent test. The 'unification' claim therefore rests on placing points back on their own fit.

full rationale

The paper's primary empirical claims are not circular: the MMNIST L20 drop (4.96→1.01), E20 drop, and the absence of population L<1 on Pendulum/CartPole/KTH are direct measurements from held-out validation pools, with paired tests and seed tables. The mediation analysis is explicitly associational (λ not randomized), so it does not overclaim causation. There is no load-bearing self-citation chain: references to Asadi et al., WorldTest, and prior world-model work are external. Appendix B's coverage conditions (validation latents covering the rollout manifold, small chord scale, matched actions) are acknowledged but not verified for free-run states; that is a validity threat to the L20 proxy, not a definitional circularity. The one circular step I can exhibit is the stochastic-forcing 'law': a linear fit to MMNIST is used to 'predict' Pendulum/CartPole L20 via ηeff values that are only said to be calibrated, so the predicted values are not independent of the fit. This affects a headline secondary claim (the unifying η-law), not the core passive/active threshold; hence partial circularity rather than score 8.

Axiom & Free-Parameter Ledger

6 free parameters · 4 axioms · 1 invented entities

The headline statistics rest on trusting the L20 proxy and the post-hoc lambda choice; the unifying forcing law additionally depends on eta_eff calibrated to the observed values, which makes the unification an identity rather than an independent check.

free parameters (6)
  • lambda operating point = 0.8
    Selected post hoc as the MMNIST operating point where mean L20 first sits near 1 with multiple seeds below 1 (Section 4.5).
  • a and b in L20(eta) law = a=1.23, b=1.82
    Least-squares fit to MMNIST stochastic-forcing sweep (Section 6.5).
  • eta_eff for Pendulum, CartPole, KTH = ~0.53, ~1.0, ~0
    Calibrated to match observed L20 values; the resulting placement on the law is by construction (Section 6.5).
  • sigma* noise scale = 0.5
    Hand-chosen scale for stochastic forcing (Appendix C).
  • residual scale alpha = 0.5
    Fixed GRU residual transition scale; directly controls how much the transition can expand or contract (Section 4.1).
  • consistency horizon set K = {1,3,5,10,15,20}
    Hand-selected set of horizons for the consistency loss; determines what 'multi-horizon' means (Section 4.2).
axioms (4)
  • standard math Theorem 1 horizon-error recurrence (Lipschitz error accumulation)
    Used to motivate the L=1 reference line; standard Lipschitz contraction argument proved in Appendix A.
  • domain assumption L20,q95 proxy validity conditions (Lipschitz, C1, small delta, matched actions)
    The proxy is informative only if validation latents cover the test manifold, chords are in the linearization regime, and action sequences are matched (Appendix B.6).
  • domain assumption GRU residual architecture is representative of world models
    All main results use one GRU-residual transition; generalization to Transformer/RSSM/JEPA-style predictors is unknown (Section 12).
  • ad hoc to paper Paired tests at post-hoc-selected lambda=0.8 are treated as confirmatory
    The critical pair (0,0.8) was selected after inspecting the sweep, so the reported p-values are not valid confirmatory statistics under a pre-registered grid (Sections 4.5 and 6.1).
invented entities (1)
  • eta_eff (calibrated effective noise) no independent evidence
    purpose: Places action-conditioned and natural-video domains on the MMNIST L20(eta) curve.
    eta_eff is defined by inverting the fitted curve to match observed L20 values, so it has no out-of-sample predictive handle (Section 6.5).

pith-pipeline@v1.3.0-alltime-deepseek · 17275 in / 10988 out tokens · 113988 ms · 2026-08-01T12:09:31.296082+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Multi-Horizon Consistency as Geometry: When Latent Dynamics Contract, and When They Do Not." pith.science (2026). https://pith.science/paper/77PKGU4W

@misc{pith2026260721645,
  author       = {Pith},
  title        = {Pith review of: Multi-Horizon Consistency as Geometry: When Latent Dynamics Contract, and When They Do Not},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/77PKGU4W}},
  note         = {Machine review of arXiv:2607.21645}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Multi-horizon latent consistency is a common training knob in video predictors and world models, but practitioners rarely know what it does to transition geometry. We treat lambda, the weight on multi-step latent agreement, as a diagnostic control and measure an empirical expansion proxy L20,q95 together with horizon-20 prediction error E20. On Moving-MNIST (n=6 seeds at the critical pair), raising lambda from 0 to 0.8 cuts L20 from 4.96 +/- 2.01 to 1.01 +/- 0.06 (paired t p=0.005, Wilcoxon p=0.031) and halves E20 (0.365 to 0.177, paired t p=1.1e-13). Four of six seeds cross L<1 at lambda=0.8. The same loss does not produce population L<1 on action-conditioned Pendulum-v1 or CartPole-v1, nor on KTH Actions video, even when E20 improves. An associational mediation analysis on MMNIST gives r-hat=0.94 (95% CI [0.88, 1.00], n=27, B=2000); lambda was not randomized. Defensive checks (architectural baselines, exogenous stress, WorldTest, MPC, scaling) mostly support a narrow claim: soft consistency can push passive video toward a near-contractive band, and that band is domain-limited. A stochastic-forcing law L20 ~ 1.23 + 1.82 eta at lambda=0.8 (bootstrap slope CI [1.73, 1.92], R^2=0.96) unifies control domains on the same curve via calibrated eta_eff. Complete joint slices at lambda in {0.4, 1.2} (30/30 cells, 5 eta x 3 seeds) show comparable linear L20(eta) slopes (~1.69 and ~2.00); we do not fit a continuous (lambda, eta) surface. We do not report DreamerV3 or TD-MPC2 returns.

Figures

Figures reproduced from arXiv: 2607.21645 by Aadi Joshi, Kavya Bhand.

Figure 1
Figure 1. Figure 1: Left: MMNIST L20 and E20 vs. λ. Right: passive MMNIST vs. action-conditioned Pendulum (mean ± std; dashed line L=1). 6.2 Pendulum and CartPole: tightening without L<1 On Pendulum-v1, L20 decreases from 2.51±0.17 at λ=0 (n=5) to 1.74±0.13 at λ=0.8 (n=3), then stays near 1.78 at λ=1.2. No seed crosses L<1. E20 is already near 10−4 , so error is a weak phase detector. Action-protocol ablations at λ=0.8 (rando… view at source ↗
Figure 2
Figure 2. Figure 2: Stochastic forcing at λ=0.8: mean L20,q95 vs. η on MMNIST (n=5 seeds per point; SEM bars). Dotted line: least-squares fit. Vertical dotted lines: calibrated ηeff for Pendulum, CartPole, and KTH. We do not observe a mean-level L=1 upcrossing on this grid (Lˆ 20(η=0)=1.10, per-η 95% CI [0.95, 1.26]), but calibrated effective noise levels place Pendulum (ηeff≈0.53) and CartPole (ηeff≈1.0) at predicted Lˆ 20 ≈… view at source ↗
Figure 3
Figure 3. Figure 3: Associational mediation DAG used for the path-ratio analysis. Solid arrows: [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Left: architectural baselines at λ=0.8. Right: Pendulum latent-MPC returns (n=3 seeds, 20 episodes each). 9 Discussion The usable claim is narrow: soft consistency contracts passive-video geometry under the L20 proxy, and that effect does not automatically transfer to action-conditioned or natural-video settings. 9.1 Passive versus active contraction Moving-MNIST is passive: transitions are learned from un… view at source ↗
Figure 5
Figure 5. Figure 5: overlays those slices on the primary λ=0.8 curve. 0.0 0.2 0.4 0.6 0.8 1.0 1.2 η 1.0 1.5 2.0 2.5 3.0 L20, q95 (m ean ± std) Joint (λ, η) slices vs. primary curve λ = 0.4 (joint) λ = 1.2 (joint) λ = 0.8 (primary) [PITH_FULL_IMAGE:figures/full_fig_p019_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: s3 proxy (L only) and Pendulum sweep. L Jacobian probe Short (6-epoch) MMNIST models at λ ∈ {0, 0.8}; mean one-step Jacobian spectral radius via power iteration on ∂zt+1/∂zt (scripts/jacobian proxy eval.py). At λ=0: L20,q95=1.295, mean spectral radius 1.285. At λ=0.8: L20,q95=1.069, mean spectral radius 1.018. Both drop with λ, consistent with Appendix B; neither certifies sup ∥Jf ∥ < 1. Source: jacobian p… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 12 linked inside Pith

  1. [1]

    DIAMOND-LoL: Enforcing Lieb-Robinson locality in diffusion world models for long-horizon consistency.OpenReview preprint, 2026

    Anonymous. DIAMOND-LoL: Enforcing Lieb-Robinson locality in diffusion world models for long-horizon consistency.OpenReview preprint, 2026. OpenReview id zBzG4Eze2j

  2. [2]

    Kavosh Asadi, Dipendra Misra, and Michael L. Littman. Lipschitz continuity in model-based reinforcement learning. InInternational Conference on Machine Learning, pages 264–273, 2018. URLhttps://proceedings.mlr.press/v80/asadi18a.html

  3. [3]

    Self-supervised learning from images with a joint-embedding predictive architecture

    Mahmoud Assran, Quentin Duval, Ishan Misra, Yann LeCun, Piotr Bojanowski, Armand Joulin, Michael Rabbat, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 15619–15629, 2023

  4. [4]

    V-JEPA 2: Self-supervised video models enable understanding, prediction and planning.arXiv preprint arXiv:2506.09985, 2025

    Mahmoud Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, 13 Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li...

  5. [5]

    Springer, Berlin, 2nd edition, 2006

    Andrea Bacciotti and Lionel Rosier.Liapunov Functions and Stability in Control Theory. Springer, Berlin, 2nd edition, 2006

  6. [6]

    Revisiting feature prediction for learning visual representations from video.Transactions on Machine Learning Research, 2024

    Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, and Nicolas Ballas. Revisiting feature prediction for learning visual representations from video.Transactions on Machine Learning Research, 2024. URL https://openreview.net/ forum?id=QaCCuDfBk2

  7. [7]

    Baron and David A

    Reuben M. Baron and David A. Kenny. The moderator–mediator variable distinction in social psychological research: Conceptual, strategic, and statistical considerations.Journal of Personality and Social Psychology, 51(6):1173–1182, 1986

  8. [8]

    Antisymmetric RNN: A dynamical system view on recurrent neural networks

    Bo Chang, Minmin Chen, Yi Lu, Rong Kan, Wei Chu, Hongxia Zhou, and Lei Bai. Antisymmetric RNN: A dynamical system view on recurrent neural networks. InInternational Conference on Learn- ing Representations, 2019. URLhttps://openreview.net/forum?id=rJg6Jh0qKm

  9. [9]

    ATM: Action-consistency transfer matrix for diagnosing and improving latent world models.arXiv preprint arXiv:2606.09028, 2026

    Jiaheng Chen. ATM: Action-consistency transfer matrix for diagnosing and improving latent world models.arXiv preprint arXiv:2606.09028, 2026

  10. [10]

    DeepMDP: Learning continuous latent space models for representation learning

    Carles Gelada, Saurabh Kumar, Jacob Buckman, Jonathan Berant, and Ofir Nachum. DeepMDP: Learning continuous latent space models for representation learning. InInternational Conference on Machine Learning, pages 2170–2179, 2019. URL https://proceedings.mlr.press/ v97/gelada19a.html

  11. [11]

    World models

    David Ha and J ¨urgen Schmidhuber. World models. InAdvances in Neural Informa- tion Processing Systems, volume 31, 2018. URL https://papers.nips.cc/paper/ 7512-world-models

  12. [12]

    Learning latent dynamics for planning from pixels

    Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Learning latent dynamics for planning from pixels. InInternational Conference on Machine Learning, pages 2555–2565,

  13. [13]

    Mastering atari with discrete world models

    Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba. Mastering atari with discrete world models. InInternational Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=0oabwyZbOu

  14. [14]

    Mastering diverse domains through world models.arXiv preprint arXiv:2301.04104, 2023

    Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models.arXiv preprint arXiv:2301.04104, 2023

  15. [15]

    TD-MPC2: Scalable, robust world models for continuous control

    Nicklas Hansen, Xiaolong Wang, and Hao Su. TD-MPC2: Scalable, robust world models for continuous control. InInternational Conference on Learning Representations, 2025. URL https: //openreview.net/forum?id=Oxh5CstDJU

  16. [16]

    Kairos: A native world model stack for physical AI.arXiv preprint arXiv:2606.16533, 2026

    Kairos Team. Kairos: A native world model stack for physical AI.arXiv preprint arXiv:2606.16533, 2026

  17. [17]

    Contrastive learning of structured world models

    Thomas Kipf, Elise van der Kolk, Max Welling, and Herke van Hoof. Contrastive learning of structured world models. InInternational Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=H1gax6VtDB

  18. [18]

    A path towards autonomous machine intelligence.OpenReview, 2022

    Yann LeCun. A path towards autonomous machine intelligence.OpenReview, 2022. URL https://openreview.net/forum?id=BZ5a1r-kVsf. 14

  19. [19]

    Toward consistent world models with multi-token prediction and latent semantic enhancement.arXiv preprint arXiv:2604.06155, 2026

    Yuxuan Liu et al. Toward consistent world models with multi-token prediction and latent semantic enhancement.arXiv preprint arXiv:2604.06155, 2026. ACL 2026 long paper

  20. [20]

    Stable recurrent models

    John Miller and Moritz Hardt. Stable recurrent models. InInternational Conference on Learn- ing Representations, 2019. URL https://openreview.net/forum?id=Hygxb2CqKm. arXiv:1805.10369

  21. [21]

    Spectral normalization for generative adversarial networks

    Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. InInternational Conference on Learning Representations, 2018. URLhttps://openreview.net/forum?id=B1QRgziT-

  22. [22]

    Preston Hess, Franz Louis Cesista, Andrii Zahorodnii, Jeremy Bernstein, and Phillip Isola

    Laker Newhouse, R. Preston Hess, Franz Louis Cesista, Andrii Zahorodnii, Jeremy Bernstein, and Phillip Isola. Training transformers with enforced lipschitz bounds. InInternational Con- ference on Learning Representations, 2026. URL https://openreview.net/forum?id= GKr4GKrn6o. arXiv:2507.13338

  23. [23]

    Causal inference in statistics: An overview.Statistics Surveys, 3:96–146, 2009

    Judea Pearl. Causal inference in statistics: An overview.Statistics Surveys, 3:96–146, 2009

  24. [24]

    Predictive objectives discard exogenous control-relevant features: A controlled mechanistic study.arXiv preprint arXiv:2606.30068, 2026

    Ayan Pendharkar. Predictive objectives discard exogenous control-relevant features: A controlled mechanistic study.arXiv preprint arXiv:2606.30068, 2026

  25. [25]

    Is the future compatible? diagnosing dynamic consistency in world action models.arXiv preprint arXiv:2605.07514, 2026

    Bo-Kai Ruan, Teng-Fang Hsiao, Ling Lo, and Hong-Han Shuai. Is the future compatible? diagnosing dynamic consistency in world action models.arXiv preprint arXiv:2605.07514, 2026

  26. [26]

    Unsupervised learning of video representations using LSTMs

    Nitish Srivastava, Elman Mansimov, and Ruslan Salakhutdinov. Unsupervised learning of video representations using LSTMs. InInternational Conference on Machine Learning, pages 843–852,

  27. [27]

    Model regularization for stable sample rollouts

    Erik Talvitie. Model regularization for stable sample rollouts. InConference on Uncertainty in Artificial Intelligence, 2014

  28. [28]

    Deepmind control suite.arXiv preprint arXiv:1801.00690, 2018

    Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy Lillicrap, and Martin Riedmiller. Deepmind control suite.arXiv preprint arXiv:1801.00690, 2018

  29. [29]

    Tenenbaum, Sebastian Josef V ollmer, Kevin Ellis, and Zenna Tavares

    Archana Warrier, Thanh Dat Nguyen, Michelangelo Naim, Moksh Jain, Yichao Liang, Karen Schroeder, Cambridge Yang, Joshua B. Tenenbaum, Sebastian Josef V ollmer, Kevin Ellis, and Zenna Tavares. Benchmarking world-model learning. InInternational Conference on Learn- ing Representations, 2026. URL https://openreview.net/forum?id=HuNIgYhBoy. arXiv:2510.19788

  30. [30]

    Beyond the next step: Variable-length latent world models for long-horizon planning.arXiv preprint arXiv:2606.21775, 2026

    Yifan Zhang et al. Beyond the next step: Variable-length latent world models for long-horizon planning.arXiv preprint arXiv:2606.21775, 2026. A Proof of Theorem 1 We prove the bound by induction on h. Let zt denote a reference trajectory and ˆzt the learned rollout, with et =∥ˆzt −z t∥. Assume ∥ˆzt+1 −f(ˆzt)∥ ≤ϵand ∥f(x)−f(y)∥ ≤L∥x−y∥ for all x, yon the r...

  31. [33]

    Chord scale.Pairs are not dominated by near-duplicate points (ratios numerically unstable) nor by extremely long chords that leave the linearization regime; our implementation clamps ∥zi −z j∥ ≥10−8 and samples uniformly over the pool

  32. [34]

    Action alignment.For action-conditioned rollouts, both trajectories must see the same action sequence (random, zero, or matched protocols in Appendix E); otherwise ratios mix transition geometry with action noise

  33. [35]

    Differentiability.Jacobian interpretation requires C1 transitions; ReLU/tanh GRU maps are piecewise smooth, so power-iteration probes are local rather than global certificates. B.7 Empirical alignment with Jacobian probe Short (six-epoch) MMNIST runs at λ∈ {0,0.8} give: L20,q95 = 1.295→1.069 and mean one-step spectral radius 1.285→1.018 (jacobian proxy.js...

  34. [36]

    2.Seed table.Per-seed values at the critical operating point, not only means

    Proxy definition.Exact formula for the expansion statistic (here L20,q95), pair count, action protocol, and validation pool. 2.Seed table.Per-seed values at the critical operating point, not only means. 3.Paired tests.When claiming aλthreshold, report paired tests on the same seeds

  35. [37]

    Domain boundary.At least one action-conditioned or natural-video negative, so L<1 is not over-exported

  36. [38]

    Planning readout (optional but clarifying).A frozen-checkpoint planner score, even if negative, to separate geometry from control utility

  37. [39]

    Our main text and appendices are written to satisfy (1)–(6) for the MMNIST critical pair and the η-law primary arm

    Compute and incompleteness.GPU hours and which appendix tables remain partial (DMC 11/12; Pendulum seed extras). Our main text and appendices are written to satisfy (1)–(6) for the MMNIST critical pair and the η-law primary arm. 21 R Extended discussion of theL=1reference line The choice ofL=1as a reference line is motivated by Theorem 1: under a global L...

  38. [2015]

    URLhttps://proceedings.mlr.press/v37/srivastava15.html

  39. [2019]

    URLhttps://proceedings.mlr.press/v97/hafner19a.html