Pith. sign in

REVIEW 15 cited by

VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.18903 v3 pith:NDORMW4V submitted 2025-06-23 cs.CV

VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory

classification cs.CV
keywords sceneviewsmemorypastvideovmemcoherenceconsistent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We propose a novel memory module for building video generators capable of interactively exploring environments. Previous approaches have achieved similar results either by out-painting 2D views of a scene while incrementally reconstructing its 3D geometry-which quickly accumulates errors-or by using video generators with a short context window, which struggle to maintain scene coherence over the long term. To address these limitations, we introduce Surfel-Indexed View Memory (VMem), a memory module that remembers past views by indexing them geometrically based on the 3D surface elements (surfels) they have observed. VMem enables efficient retrieval of the most relevant past views when generating new ones. By focusing only on these relevant views, our method produces consistent explorations of imagined environments at a fraction of the computational cost required to use all past views as context. We evaluate our approach on challenging long-term scene synthesis benchmarks and demonstrate superior performance compared to existing methods in maintaining scene coherence and camera control.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MemLearner: Learning to Query Context memory for Video World Models

    cs.CV 2026-06 unverdicted novelty 7.0

    MemLearner introduces a learning-based adaptive context query method using query tokens in video world models to improve long-term scene consistency over rule-based retrieval.

  2. PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

    cs.CV 2026-06 unverdicted novelty 6.0

    PermaVid disentangles spatial context into semantic appearance and geometric structure via multi-modal memory banks and edit-aware updates to maintain long-term consistency in video generation after edits.

  3. Latent Spatial Memory for Video World Models

    cs.CV 2026-06 unverdicted novelty 6.0

    Mirage stores and queries 3D scene information in diffusion latent space via depth-guided lifting and warping, yielding 10.57× faster generation and 55× smaller memory than explicit RGB point-cloud baselines while rea...

  4. MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data

    cs.CV 2026-06 unverdicted novelty 6.0

    MetaWorld scales multi-agent video world models from single-view videos using monocular decomposition into ego-motion and trajectories, subject-aware generation, and cross-attention alignment for consistency.

  5. Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    Robust Dreamer uses Latent Gaussian Memory anchored to diffusion latents and Deviation Learning with a Dynamic Deviation Archive to reduce drift in long-horizon action-controlled image-to-video generation, reporting S...

  6. E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control

    cs.CV 2026-05 unverdicted novelty 6.0

    E³C is a video diffusion model that disentangles persistent 3D scene structure via point-cloud memory from human dynamics via ego-exo pose controls for improved egocentric video generation on the Nymeria dataset.

  7. WorldKV: Efficient World Memory with World Retrieval and Compression

    cs.CV 2026-05 unverdicted novelty 6.0

    WorldKV enables persistent world memory in autoregressive video diffusion models by selectively retrieving and compressing KV-cache chunks, matching full-cache fidelity at roughly twice the throughput without training.

  8. I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation

    cs.CV 2026-03 conditional novelty 6.0

    An implicit 3D-aware memory mechanism, I3DM, improves revisit consistency and camera control in video scene generation by retrieving historical frames with NVS features and injecting 3D-aligned conditioned latents.

  9. Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

    cs.CV 2026-03 unverdicted novelty 6.0

    VEGA-3D extracts intermediate spatiotemporal features from a pretrained video diffusion model and fuses them into MLLMs to improve geometric and embodied reasoning without explicit 3D supervision.

  10. Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

    cs.CV 2026-03 conditional novelty 6.0

    Fusing intermediate-layer, mid-noise features of a video diffusion model with semantic tokens via learned gating improves MLLMs on 3D scene understanding benchmarks.

  11. UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models

    cs.CV 2026-02 conditional novelty 6.0

    A video-generation world model that warps positional encodings of memory frames to target viewpoints achieves state-of-the-art long-term consistency and camera control.

  12. CustomX: Unified Character, Action, and Scene Customization in Video World Models

    cs.CV 2025-12 conditional novelty 6.0

    AniX generates controllable videos of a user-supplied character performing typed actions inside a user-supplied 3D scene by fine-tuning a pre-trained video generator on small locomotion datasets.

  13. DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory

    cs.CV 2026-05 unverdicted novelty 5.0

    DecMem proposes a decoupled memory system using sparse global and anchored local components to enable consistent minute-long controllable video generation in world models.

  14. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

    cs.CV 2026-04 unverdicted novelty 4.0

    HY-World 2.0 generates and reconstructs high-fidelity navigable 3D Gaussian Splatting worlds from text, images, or videos via upgraded panorama, planning, expansion, and composition modules, with released code claimin...

  15. Evolution of Video Generative Foundations

    cs.CV 2026-04 unverdicted novelty 2.0

    This survey traces video generation technology from GANs to diffusion models and then to autoregressive and multimodal approaches while analyzing principles, strengths, and future trends.