Pith. sign in

REVIEW 2 cited by

Scene Representation Transformer: Geometry-Free Novel View Synthesis Through Set-Latent Scene Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.13152 v3 pith:PEOZZFUL submitted 2021-11-25 cs.CV cs.AIcs.GRcs.LGcs.RO

classification cs.CVcs.AIcs.GRcs.LGcs.RO
keywords scenerepresentationnovelimagestransformerrepresentationsviewsinteractive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A classical problem in computer vision is to infer a 3D scene representation from few images that can be used to render novel views at interactive rates. Previous work focuses on reconstructing pre-defined 3D representations, e.g. textured meshes, or implicit representations, e.g. radiance fields, and often requires input images with precise camera poses and long processing times for each novel scene. In this work, we propose the Scene Representation Transformer (SRT), a method which processes posed or unposed RGB images of a new area, infers a "set-latent scene representation", and synthesises novel views, all in a single feed-forward pass. To calculate the scene representation, we propose a generalization of the Vision Transformer to sets of images, enabling global information integration, and hence 3D reasoning. An efficient decoder transformer parameterizes the light field by attending into the scene representation to render novel views. Learning is supervised end-to-end by minimizing a novel-view reconstruction error. We show that this method outperforms recent baselines in terms of PSNR and speed on synthetic datasets, including a new dataset created for the paper. Further, we demonstrate that SRT scales to support interactive visualization and semantic segmentation of real-world outdoor environments using Street View imagery.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis

    cs.CV 2026-03 accept novelty 6.0 of 10

    Initializing a highway encoder-decoder NVS network from VGGT 3D-aware features yields 31.4 PSNR on RealEstate10k with real-time decoding and optional unposed inputs.

  2. Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning

    cs.RO 2025-08 conditional novelty 4.0 of 10

    The thesis demonstrates that combining implicit 3D scene representations with LLM-based reasoning, using text as an interface, yields strong performance on robotic perception and spatial language tasks.

Pith tools