Pith. sign in

REVIEW 2 cited by

Towards an Interpretable Latent Space in Structured Models for Video Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.07713 v1 pith:CRPGWDNK submitted 2021-07-16 cs.LG

classification cs.LG
keywords modelobjectspacelatentkipfphysicalpositionsprediction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We focus on the task of future frame prediction in video governed by underlying physical dynamics. We work with models which are object-centric, i.e., explicitly work with object representations, and propagate a loss in the latent space. Specifically, our research builds on recent work by Kipf et al. \cite{kipf&al20}, which predicts the next state via contrastive learning of object interactions in a latent space using a Graph Neural Network. We argue that injecting explicit inductive bias in the model, in form of general physical laws, can help not only make the model more interpretable, but also improve the overall prediction of model. As a natural by-product, our model can learn feature maps which closely resemble actual object positions in the image, without having any explicit supervision about the object positions at the training time. In comparison with earlier works \cite{jaques&al20}, which assume a complete knowledge of the dynamics governing the motion in the form of a physics engine, we rely only on the knowledge of general physical laws, such as, world consists of objects, which have position and velocity. We propose an additional decoder based loss in the pixel space, imposed in a curriculum manner, to further refine the latent space predictions. Experiments in multiple different settings demonstrate that while Kipf et al. model is effective at capturing object interactions, our model can be significantly more effective at localising objects, resulting in improved performance in 3 out of 4 domains that we experiment with. Additionally, our model can learn highly intrepretable feature maps, resembling actual object positions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    Latent transitions in models like Dreamer are biased toward dense regions, creating attractors that hide true dynamics discrepancies and cause epistemic uncertainty to be unreliable while overestimating rewards.

  2. Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models

    cs.LG 2026-04 conditional novelty 6.0 of 10

    Latent world-model rollouts exhibit an attractor effect that makes ensemble uncertainty underestimate compounding rollout error and overestimate rewards.

Pith tools