REVIEW 4 major objections 5 minor 44 references
Video Interpolation and Prediction with Unsupervised Landmarks
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Unsupervised 2D Gaussian landmarks let a video model predict over 100 frames ahead while preserving foreground structure.
desk verdict A plausible but over-claimed extension of unsupervised landmark learning; the 100-frame prediction claim rests on qualitative evidence while the quantitative curves stop at 30 frames. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Gaussian landmark pose state together with Cholesky parameterization. Each of K parts is represented by a 2D Gaussian fitted to a soft activation map; the Gaussian's mean gives the part's location and its covariance gives its spread and orientation. To manipulate or predict these Gaussians, the covariance is written as $\Sigma_k = L_k L_k^T$ with $L_k$ lower-triangular, yielding five scalars per landmark (two mean coordinates, two Cholesky diagonal entries, one off-diagonal entry). Interpolating or predicting in this space always produces a valid covariance when mapped back through $L L^T$. The decoder then turns the Gaussians back into heatmaps and uses spatially-adaptive normalization (SPADE) to render the frame, while the temporal model is an LSTM operating on residuals of the state vector. This machinery is what lets the method move, interpolate, and extrapolate pose without ever producing an invalid or non-interpretable latent state.
What would settle it
Take a synthetic video of a rigid ellipse rotating in place about its center, so the landmark means stay fixed and only the covariance matrices rotate. If the model's interpolation between first and last frame does not pass through the true intermediate orientation (measured by LPIPS or pixel error at the midpoint), then linear interpolation in Cholesky space does not track this non-linear deformation, and the claim that Gaussian landmark space is a sufficient pose state for general motion fails.
Extended reading notes
Core claim
On its own terms, the central claim is that a factored pose-appearance representation—K 2D Gaussian landmarks (mean $\mu_k$ and covariance $\Sigma_k$) plus per-landmark appearance vectors—is a stable and sufficient state for video dynamics. The landmarks are learned self-supervised through image reconstruction with color jitter, thin-plate-spline warping, and temporal frame sampling as perturbations, so the encoder must localize the same semantic parts across appearance and deformation changes. Interpolation is done by linear interpolation in the coordinates $(\mu_k, L_k)$, where $L_k$ is the Cholesky factor of $\Sigma_k$, which guarantees the reconstructed covariance $L_k L_k^T$ stays positive definite. Extrapolation is done by an LSTM that predicts residuals to $(\mu_k, L_k)$ at each time step, a formulation the paper argues is key to stable long-range predictions. The paper demonstrates this on sign-language, robot-pushing, and action videos, reporting that the predicted sequences remain structurally intact for roughly a hundred frames.
Load-bearing premise
The load-bearing premise is that each frame's future is fully described by K Gaussian landmark positions and their per-part appearance vectors, so any motion or appearance change that this representation cannot express—moving backgrounds, occlusions, novel clothing or lighting—will be dropped and the prediction will drift.
Editorial extensions
If this is right
- Long-range prediction cost scales with the number of landmarks, not with image resolution, since the LSTM operates on a few dozen scalars per frame.
- No keypoint or pose annotations are needed; the same pipeline transfers to a new object class as long as rough object-level crops and video data are available.
- Interpolation in pose space yields predictable, editable motion paths between two keyframes, which matters for video editing and animation.
- Residual prediction in Cholesky space keeps covariance matrices valid indefinitely, so the representation cannot drift into an invalid state even over long rollouts.
- On the tested datasets, the method matches or beats stochastic adversarial baselines on a perceptual metric after roughly 15 predicted frames, while better preserving foreground structure.
Reading between the lines
- The paper leaves background motion largely unmodeled; a natural extension is to add a global background latent or a background flow field to the state, which would address the reported failure on scenes with moving cameras.
- Because the state is just a few Gaussians and appearance vectors, the same Cholesky-residual recipe could be applied to other compact pose parameterizations, such as 3D keypoints or articulated body models, whenever a differentiable decoder exists.
- The reported 100-frame stability suggests the LSTM learns local, nearly linear deformations; an autoregressive model with the same residual state (e.g., a Transformer over time) might extend the range further, which the paper does not test.
- The method's linear interpolation in Cholesky space is one of several valid covariant interpolation schemes; the paper notes Wasserstein barycenters as an alternative, so a head-to-head comparison on rotating objects would clarify which parameterization tracks true motion best.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an unsupervised landmark-based video interpolation and prediction method. An encoder factorizes each frame into K 2D Gaussian landmark "pose" parameters and per-landmark appearance vectors; a decoder reconstructs the frame from these factors using SPADE-like normalization. For interpolation, the authors linearly interpolate the Cholesky-decomposed Gaussian parameters. For extrapolation, an LSTM predicts residual updates to the means and Cholesky factors. The method is evaluated on BBC Pose, BAIR, and KTH. The authors report improved landmark accuracy over Lorenz et al. on BBC, LPIPS/SSIM/PSNR curves on BAIR and KTH up to 30-35 predicted frames, and qualitative long-range (100-frame) predictions on BBC.
Significance. If substantiated, the work is a useful step toward interpretable latent-space video dynamics: it offers a principled Cholesky parameterization that guarantees valid covariances under interpolation and extrapolation, and it demonstrates a clean way to combine unsupervised landmarks with a learned temporal model. The reported BBC landmark improvement (75.0±0.9 vs. Lorenz's 74.5) and the qualitative stability of motion structure are creditable. However, the central long-range claim is not currently backed by quantitative evidence at the claimed horizon, and the paper's own limitations acknowledge degradation around 100 frames and poor handling of background and novel appearances.
major comments (4)
- [Abstract/§1 Contribution 1; §3.1, Fig. 3] The headline claim that the method 'can interpolate and extrapolate over 100 frames into the future while maintaining the structure of the moving foreground object' is not supported by the quantitative evaluation: the BAIR and KTH curves in Figs. 5 and 7 stop at 30 and 35 predicted frames, respectively, and the only 100-frame evidence is qualitative BBC Pose examples (Fig. 3 and Appendix C) that use training-set signer appearances. Please report quantitative long-horizon metrics (e.g., LPIPS, SSIM, and PSNR at 60, 80, and 100 frames on BBC or a suitable dataset, with per-sequence variance), or revise the claim to the horizon actually evaluated.
- [§5 Limitations] The limitations paragraph states that temporal prediction is stable 'for a little over 100 frames, after which it starts to degrade' and that the background is 'largely unhandled' and novel appearances are not rendered faithfully. This undermines the 100-frame structure-preservation claim as stated: the Fig. 4 BAIR example shows the red object disappearing at frame 24, and the Fig. 3 second row shows attire mismatch. Please define precisely what structure is preserved, restrict the central claim accordingly, and quantify the onset of degradation rather than relying on a qualitative 'little over 100 frames.'
- [§3.2; Figs. 5–7] No error bars, standard deviations, or repeated-seed statistics are reported for the BAIR and KTH LPIPS, SSIM, and PSNR curves, so the assertion that the method becomes competitive after 15 frames is not shown to be outside noise. Please include variance over at least three training runs (as in Table 1 for BBC) and report the number of test sequences averaged.
- [§3.2, Appendix E, Fig. 13] The interpolation evaluation is limited to SSIM against SuperSlomo on BAIR, with no error bars, and the qualitative trajectory comparison in Fig. 6 does not include a quantitative motion or structural metric. Please add a quantitative interpolation evaluation (e.g., LPIPS and a trajectory error metric) over the full interpolation span for all methods, with variance estimates.
minor comments (5)
- [Appendix A] The BAIR LSTM is described as trained with a '10 input 0 future setup (never conditions on its own output during training)'; clarify whether this is purely open-loop training and how it is reconciled with the residual-prediction formulation in §2.4.
- [Table 1] The sentence 'Our implementation outperforms that of [18]' appears to rely on best-of-3 values (75.7/76.1), while the mean for Ours (Lorenz), 74.2, is below Lorenz's reported 74.5; please report the comparison for means with standard errors.
- [Fig. 8 caption] The caption contains a typo: 'may not reproduce the foreground and background aas accurately' should read 'as accurately.'
- [Fig. 6 caption] The caption notes that VideoFlow frames do not strictly correspond to labeled time steps; please align the time axes or use a controlled reimplementation for a fair comparison.
- [§3.1 and Appendix C] Section 3.1 says the first sequence features held-out frames from the training set while Appendix C calls these 'easier predictions' because the model has access to visually similar frames from the same sequence; clarify the distinction between train-set and validation-set evaluation throughout.
Circularity Check
No significant circularity: the video prediction and interpolation claims rest on external held-out evaluations and a pose-dynamics LSTM trained on ground-truth pose trajectories.
full rationale
The paper's derivation chain is self-contained and non-circular. Pose and appearance encoders are trained by image reconstruction on individual frames (Eqs. 2-4), with the pose represented as 2D Gaussian landmarks. Interpolation is an explicit linear operation in Cholesky parameter space (Eq. 7), and extrapolation is performed by an LSTM trained to predict residual perturbations of these pose parameters from ground-truth pose sequences extracted from training videos. The prediction target (future frames or landmarks) is not used to define the learned state, and no fitted parameter is renamed as a prediction. Quantitative evaluation uses external benchmarks: BBC supervised keypoint annotations with a fitted linear regressor, and LPIPS/SSIM/PSNR comparisons against SAVP, SVG-LP, DRNet, and SuperSlomo on held-out BAIR and KTH splits. The only self-referential aspect is that the landmark representation is learned by reconstruction and then reused for dynamics, which is standard representation learning rather than circularity. Self-citations to prior NVIDIA works (SDC-Net, cycle-consistency video interpolation) appear only in related work and are not load-bearing for the central claims. The skepticism about the 100-frame claim is a question of evidence strength and quantitative support, not a reduction of the derivation to its inputs; the paper's own limitations section even concedes degradation past roughly 100 frames and poor background handling. Those are correctness or evaluation concerns, not circularity.
Assumptions & free parameters
free parameters (3)
- Number of landmarks K =
40 (BBC), 30 (BAIR, KTH)
- Reconstruction loss weights lambda_vgg, lambda_MSE =
not reported
- LSTM training horizon =
10+10 (BBC), 10+0 (BAIR), 10+10 (KTH)
assumptions (4)
- standard math Cholesky decomposition L with positive diagonal uniquely parameterizes a positive definite covariance Sigma = L L^T, and LL^T of any real L remains positive semidefinite.
- domain assumption Each moving part's activation map can be faithfully approximated by a single 2D Gaussian.
- domain assumption The K landmark means and covariances plus K appearance vectors constitute a sufficient state for predicting future frames.
- domain assumption Temporal dynamics of landmarks are learnable by an LSTM operating on residuals.
Cite this review
Pith. "Pith review of Video Interpolation and Prediction with Unsupervised Landmarks." pith.science (2026). https://pith.science/paper/TQDSPDUZ
@misc{pith2026190902749,
author = {Pith},
title = {Pith review of: Video Interpolation and Prediction with Unsupervised Landmarks},
year = {2026},
howpublished = {\url{https://pith.science/paper/TQDSPDUZ}},
note = {Machine review of arXiv:1909.02749}
}
read the original abstract
Prediction and interpolation for long-range video data involves the complex task of modeling motion trajectories for each visible object, occlusions and dis-occlusions, as well as appearance changes due to viewpoint and lighting. Optical flow based techniques generalize but are suitable only for short temporal ranges. Many methods opt to project the video frames to a low dimensional latent space, achieving long-range predictions. However, these latent representations are often non-interpretable, and therefore difficult to manipulate. This work poses video prediction and interpolation as unsupervised latent structure inference followed by a temporal prediction in this latent space. The latent representations capture foreground semantics without explicit supervision such as keypoints or poses. Further, as each landmark can be mapped to a coordinate indicating where a semantic part is positioned, we can reliably interpolate within the coordinate domain to achieve predictable motion interpolation. Given an image decoder capable of mapping these landmarks back to the image domain, we are able to achieve high-quality long-range video interpolation and extrapolation by operating on the landmark representation space.
Figures
Figures from the paper (16 more)
Reference graph
Works this paper leans on
-
[1]
M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine. Stochastic variational video prediction. arXiv preprint arXiv:1710.11252, 2017
arXiv 2017
-
[2]
J. Charles, T. Pfister, D. Magee, D. Hogg, and A. Zisserman. Domain adaptation for upper body pose tracking in signed TV broadcasts. In British Machine Vision Conference, 2013
work page 2013
-
[3]
Y . Chen, T. T. Georgiou, and A. Tannenbaum. Optimal transport for gaussian mixture models. IEEE Access, 7:6269–6278, 2019
work page 2019
-
[4]
E. Denton and R. Fergus. Stochastic video generation with a learned prior. arXiv preprint arXiv:1802.07687, 2018
arXiv 2018
-
[5]
E. L. Denton et al. Unsupervised learning of disentangled representations from video. In Advances in neural information processing systems , pages 4414–4423, 2017
work page 2017
-
[6]
L. Dinh, J. Sohl-Dickstein, and S. Bengio. Density estimation using real nvp. arXiv preprint arXiv:1605.08803, 2016
arXiv 2016
-
[7]
C. Finn, I. Goodfellow, and S. Levine. Unsupervised learning for physical interaction through video prediction. In Advances in neural information processing systems , pages 64–72, 2016
work page 2016
-
[8]
S. Hochreiter and J. Schmidhuber. Lstm can solve hard long time lag problems. In Advances in neural information processing systems , pages 473–479, 1997
work page 1997
Show all 44 references
-
[9]
Jakab, A
T. Jakab, A. Gupta, H. Bilen, and A. Vedaldi. Unsupervised learning of object landmarks through conditional image generation. In Advances in Neural Information Processing Systems, 2018
2018
-
[10]
Jiang, D
H. Jiang, D. Sun, V . Jampani, M.-H. Yang, E. Learned-Miller, and J. Kautz. Super slomo: High quality estimation of multiple intermediate frames for video interpolation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 9000–9008, 2018
2018
-
[11]
Kanazawa, D
A. Kanazawa, D. W. Jacobs, and M. Chandraker. Warpnet: Weakly supervised matching for single-view reconstruction. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016
2016
-
[12]
D. P. Kingma and P. Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. In Advances in Neural Information Processing Systems , pages 10215–10224, 2018
2018
-
[13]
D. P. Kingma and M. Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013
2013 arXiv
-
[14]
Kumar, M
M. Kumar, M. Babaeizadeh, D. Erhan, C. Finn, S. Levine, L. Dinh, and D. Kingma. Videoflow: A flow-based generative model for video. arXiv preprint arXiv:1903.01434, 2019
1903 arXiv
-
[15]
Laptev, B
I. Laptev, B. Caputo, et al. Recognizing human actions: a local svm approach. pages 32–36. IEEE, 2004
2004
-
[16]
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine. Stochastic adversarial video prediction. arXiv preprint arXiv:1804.01523, 2018
2018 arXiv
-
[17]
Z. Liu, R. A. Yeh, X. Tang, Y . Liu, and A. Agarwala. Video frame synthesis using deep voxel flow. In Proceedings of the IEEE International Conference on Computer Vision , 2017
2017
-
[18]
Lorenz, L
D. Lorenz, L. Bereska, T. Milbich, and B. Ommer. Unsupervised part-based disentangling of object shape and appearance. In CVPR, 2019
2019
-
[19]
Madyastha, V
V . Madyastha, V . Ravindra, S. Mallikarjunan, and A. Goyal. Extended kalman filter vs. error state kalman filter for aircraft attitude estimation. 08 2011
2011
-
[20]
Mathieu, C
M. Mathieu, C. Couprie, and Y . LeCun. Deep multi-scale video prediction beyond mean square error. arXiv preprint arXiv:1511.05440, 2015
2015 arXiv
-
[21]
Miyato, T
T. Miyato, T. Kataoka, M. Koyama, and Y . Yoshida. Spectral normalization for generative adversarial networks. In International Conference on Learning Representations (ICLR), 2018
2018
-
[22]
A. Paliwal. Super-slomo. https://github.com/avinashpaliwal/Super-SloMo
-
[23]
Park, M.-Y
T. Park, M.-Y . Liu, T.-C. Wang, and J.-Y . Zhu. Semantic image synthesis with spatially- adaptive normalization. In CVPR, 2019
2019
-
[24]
Pfister, J
T. Pfister, J. Charles, and A. Zisserman. Flowing convnets for human pose estimation in videos. In Proceedings of the IEEE International Conference on Computer Vision , pages 1913–1921, 2015
1913
-
[25]
Pottorff, J
R. Pottorff, J. Nielsen, and D. Wingate. Video extrapolation with an invertible linear embed- ding. arXiv preprint arXiv:1903.00133, 2019
1903 arXiv
-
[26]
F. A. Reda, G. Liu, K. J. Shih, R. Kirby, J. Barker, D. Tarjan, A. Tao, and B. Catanzaro. Sdc- net: Video prediction using spatially-displaced convolution. In Proceedings of the European Conference on Computer Vision (ECCV), pages 718–733, 2018
2018
-
[27]
F. A. Reda, D. Sun, A. Dundar, M. Shoeybi, G. Liu, K. J. Shih, A. Tao, J. Kautz, and B. Catanzaro. Unsupervised video interpolation using cycle consistency. arXiv preprint arXiv:1906.05928, 2019. 10
1906 arXiv
-
[28]
S. E. Reed, Y . Zhang, Y . Zhang, and H. Lee. Deep visual analogy-making. In Advances in neural information processing systems, pages 1252–1260, 2015
2015
-
[29]
Ronneberger, P
O. Ronneberger, P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical im- age segmentation. In International Conference on Medical image computing and computer- assisted intervention, 2015
2015
-
[30]
Schuldt, I
C. Schuldt, I. Laptev, and B. Caputo. Recognizing human actions: a local svm approach. In Proceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004. , volume 3, pages 32–36. IEEE, 2004
2004
-
[31]
Siarohin, S
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe. Animating arbitrary objects via deep motion transfer. arXiv preprint arXiv:1812.08861, 2018
2018 arXiv
-
[32]
Suwajanakorn, N
S. Suwajanakorn, N. Snavely, J. J. Tompson, and M. Norouzi. Discovery of latent 3d keypoints via end-to-end geometric reasoning. In Advances in Neural Information Processing Systems , pages 2059–2070, 2018
2018
-
[33]
Thewlis, H
J. Thewlis, H. Bilen, and A. Vedaldi. Unsupervised learning of object frames by dense equiv- ariant image labelling. In Advances in Neural Information Processing Systems, pages 844–855, 2017
2017
-
[34]
Thewlis, H
J. Thewlis, H. Bilen, and A. Vedaldi. Unsupervised learning of object landmarks by factorized spatial embeddings. In International Conference on Computer Vision (ICCV) , 2017
2017
-
[35]
Tulyakov, M.-Y
S. Tulyakov, M.-Y . Liu, X. Yang, and J. Kautz. Mocogan: Decomposing motion and content for video generation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1526–1535, 2018
2018
-
[36]
Unterthiner, S
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly. To- wards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018
2018 arXiv
-
[37]
Villegas, J
R. Villegas, J. Yang, Y . Zou, S. Sohn, X. Lin, and H. Lee. Learning to generate long-term future via hierarchical prediction. In Proceedings of the 34th International Conference on Machine Learning-V olume 70, pages 3560–3569. JMLR. org, 2017
2017
-
[38]
V ondrick and A
C. V ondrick and A. Torralba. Generating the future with adversarial transformers. InProceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 1020–1028, 2017
2017
-
[39]
Walker, A
J. Walker, A. Gupta, and M. Hebert. Dense optical flow prediction from a static image. In Proceedings of the IEEE International Conference on Computer Vision , 2015
2015
-
[40]
Wichers, R
N. Wichers, R. Villegas, D. Erhan, and H. Lee. Hierarchical long-term video prediction without supervision. arXiv preprint arXiv:1806.04768, 2018
2018 arXiv
-
[41]
T. Xue, J. Wu, K. Bouman, and B. Freeman. Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks. In Advances in Neural Information Processing Systems, 2016
2016
-
[42]
R. Zhang. https://github.com/richzhang/perceptualsimilarity. https://github.com/ richzhang/PerceptualSimilarity
-
[43]
Zhang, P
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018
2018
-
[44]
Zhang, Y
Y . Zhang, Y . Guo, Y . Jin, Y . Luo, Z. He, and H. Lee. Unsupervised discovery of object land- marks as structural representations. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 11 A Implementation Details The overall architecture consis...
2013
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.