Pith. sign in

REVIEW 15 cited by

LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.17242 v2 pith:NSTA57FV submitted 2024-10-22 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords lvsmsynthesisviewnovelmethodsmodelpreviousapproach
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose the Large View Synthesis Model (LVSM), a novel transformer-based approach for scalable and generalizable novel view synthesis from sparse-view inputs. We introduce two architectures: (1) an encoder-decoder LVSM, which encodes input image tokens into a fixed number of 1D latent tokens, functioning as a fully learned scene representation, and decodes novel-view images from them; and (2) a decoder-only LVSM, which directly maps input images to novel-view outputs, completely eliminating intermediate scene representations. Both models bypass the 3D inductive biases used in previous methods -- from 3D representations (e.g., NeRF, 3DGS) to network designs (e.g., epipolar projections, plane sweeps) -- addressing novel view synthesis with a fully data-driven approach. While the encoder-decoder model offers faster inference due to its independent latent representation, the decoder-only LVSM achieves superior quality, scalability, and zero-shot generalization, outperforming previous state-of-the-art methods by 1.5 to 3.5 dB PSNR. Comprehensive evaluations across multiple datasets demonstrate that both LVSM variants achieve state-of-the-art novel view synthesis quality. Notably, our models surpass all previous methods even with reduced computational resources (1-2 GPUs). Please see our website for more details: https://haian-jin.github.io/projects/LVSM/ .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time

    cs.CV 2025-06 conditional novelty 7.0 of 10

    4D-LRM is a transformer that maps sparse posed frames scattered across time to a cloud of 4D Gaussians and renders any query view at any query time in under 1.5 seconds.

  2. Wonderland: Navigating 3D Scenes from a Single Image

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A feed-forward pipeline reconstructs 3D Gaussian scenes from single images by regressing 3DGS directly from camera-conditioned video diffusion latents.

  3. InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A single-image feed-forward Gaussian splatting method that samples supports from predicted depth and decodes Gaussian attributes implicitly, improving cross-dataset large-baseline novel view synthesis.

  4. NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Anchoring Gaussian centers to predicted raymaps and jointly optimizing RGB, raymap, and camera losses with a dual-frequency curriculum suppresses pose drift and improves pose-free 3D reconstruction on long sequences.

  5. Real-Time Human Reconstruction and Animation using Feed-Forward Gaussian Splatting

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    A feed-forward transformer predicts SMPL-X vertex-aligned 3D Gaussians in a canonical T-pose, enabling real-time animation by linear blend skinning without per-frame network inference.

  6. LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows

    cs.CV 2026-04 conditional novelty 6.0 of 10

    Scaling transformer context with sparse attention and 3D-aware block routing improves feed-forward 3D reconstruction and inverse rendering, closing much of the quality gap with dense-view optimization.

  7. I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    An implicit 3D-aware memory mechanism, I3DM, improves revisit consistency and camera control in video scene generation by retrieving historical frames with NVS features and injecting 3D-aligned conditioned latents.

  8. ILV: Iterative Latent Volumes for Fast and Accurate Sparse-View CT Reconstruction

    cs.CV 2026-03 conditional novelty 6.0 of 10

    ILV recovers fine anatomical detail in sparse-view CBCT by iteratively updating an explicit 3D latent volume with multi-view X-ray features and a learned prior, outperforming prior feed-forward and optimization method...

  9. OmniNWM: Omniscient Driving Navigation World Models

    cs.CV 2025-10 conditional novelty 6.0 of 10

    OmniNWM jointly generates long panoramic multi-modal driving videos, controls them precisely via normalized Plücker ray-maps, and derives dense driving rewards from generated 3D occupancy.

  10. iLRM: An Iterative Large 3D Reconstruction Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    iLRM reconstructs 3D Gaussian scenes from multiple photos through iterative refinement of viewpoint tokens, achieving higher quality and speed than prior feed-forward models.

  11. RayZer: A Self-supervised Large View Synthesis Model

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A self-supervised transformer model predicts camera poses and scene features from unposed images and renders novel views, reaching performance on par with pose-supervised baselines.

  12. MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 700K-scene procedural, non-semantic synthetic dataset improves large reconstruction models by 1.2 to 1.8 dB PSNR when combined with real data.

  13. SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A cross-view attention module and hybrid training recipe let a frozen text-to-video model generate synchronized multi-camera videos from arbitrary viewpoints given only a text prompt and camera poses.

  14. OpenWorldLib: A Unified Codebase and Definition of Advanced World Models

    cs.CV 2026-04 unverdicted novelty 4.0 of 10

    OpenWorldLib offers a standardized codebase and definition for world models that combine perception, interaction, and memory to understand and predict the world.

  15. DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multi-view conditioning framework that improves controllable novel view synthesis and 3D reconstruction by injecting fused 3D latents into frozen image and video diffusion models.

Pith tools