Pith. sign in

REVIEW 2 cited by

S-HR-VQVAE: Sequential Hierarchical Residual Learning Vector Quantized Variational Autoencoder for Video Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.06701 v3 pith:A3M2Z3NS submitted 2023-07-13 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords predictionlearningmodelnovels-hr-vqvaevideoast-pmautoencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We address the video prediction task by putting forth a novel model that combines (i) a novel hierarchical residual learning vector quantized variational autoencoder (HR-VQVAE), and (ii) a novel autoregressive spatiotemporal predictive model (AST-PM). We refer to this approach as a sequential hierarchical residual learning vector quantized variational autoencoder (S-HR-VQVAE). By leveraging the intrinsic capabilities of HR-VQVAE at modeling still images with a parsimonious representation, combined with the AST-PM's ability to handle spatiotemporal information, S-HR-VQVAE can better deal with major challenges in video prediction. These include learning spatiotemporal information, handling high dimensional data, combating blurry prediction, and implicit modeling of physical characteristics. Extensive experimental results on four challenging tasks, namely KTH Human Action, TrafficBJ, Human3.6M, and Kitti, demonstrate that our model compares favorably against state-of-the-art video prediction techniques both in quantitative and qualitative evaluations despite a much smaller model size. Finally, we boost S-HR-VQVAE by proposing a novel training method to jointly estimate the HR-VQVAE and AST-PM parameters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MambaVideo for Discrete Video Tokenization with Channel-Split Quantization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A Mamba-based hierarchical video tokenizer with channel-split quantization achieves state-of-the-art reconstruction and generation scores while preserving token count.

  2. Scaling Image Tokenizers with Grouped Spherical Quantization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GSQ combines spherical codebook initialization, normalized lookup, and group-wise latent decomposition to achieve strong reconstruction quality at 16x spatial downsampling in far fewer training steps than prior tokenizers.

Pith tools