Pith. sign in

REVIEW 5 cited by

FlowCam: Training Generalizable 3D Radiance Fields without Camera Poses via Pixel-Aligned Scene Flow

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00180 v1 pith:X27DBQV5 submitted 2023-05-31 cs.CV cs.AIcs.GRcs.LG

classification cs.CVcs.AIcs.GRcs.LG
keywords scenecameraflowposesvideoestimationmethodneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reconstruction of 3D neural fields from posed images has emerged as a promising method for self-supervised representation learning. The key challenge preventing the deployment of these 3D scene learners on large-scale video data is their dependence on precise camera poses from structure-from-motion, which is prohibitively expensive to run at scale. We propose a method that jointly reconstructs camera poses and 3D neural scene representations online and in a single forward pass. We estimate poses by first lifting frame-to-frame optical flow to 3D scene flow via differentiable rendering, preserving locality and shift-equivariance of the image processing backbone. SE(3) camera pose estimation is then performed via a weighted least-squares fit to the scene flow field. This formulation enables us to jointly supervise pose estimation and a generalizable neural scene representation via re-rendering the input video, and thus, train end-to-end and fully self-supervised on real-world video datasets. We demonstrate that our method performs robustly on diverse, real-world video, notably on sequences traditionally challenging to optimization-based pose estimation techniques.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DynSUP: Dynamic Gaussian Splatting from An Unposed Image Pair

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A pose-free two-image pipeline decomposes a dynamic scene into rigid objects and fits per-Gaussian SE(3) motions to synthesize novel views of moving scenes.

  2. RegGS: Unposed Sparse Views Gaussian Splatting with 3DGS Registration

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RegGS aligns locally generated 3D Gaussian maps using a Sinkhorn-approximated mixture Wasserstein distance, improving pose estimation and novel view synthesis from sparse unposed views.

  3. SelfSplat: Pose-Free and 3D Prior-Free Generalizable 3D Gaussian Splatting

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SelfSplat jointly predicts depth, camera poses and 3D Gaussians from unposed image triplets, and outperforms prior pose-free baselines on RealEstate10K, ACID and DL3DV.

  4. Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    EgoMono4D estimates depth, camera intrinsics and poses from unlabeled egocentric videos in a single feed-forward pass, reconstructing dense per-frame point clouds better than baseline methods on in-domain and zero-sho...

  5. SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Proposal-guided multi-view projection with iterative VLM selection achieves 55.6% and 53.2% Acc@0.25 on ScanRefer and Nr3D, a new zero-shot 3D visual grounding state of the art.

Pith tools