Pith. sign in

REVIEW 28 cited by

PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.12024 v2 pith:PZJOYWCM submitted 2023-11-20 cs.CV

classification cs.CV
keywords pf-lrmlargemodelreconstructioncameraobjectposepose-free
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose a Pose-Free Large Reconstruction Model (PF-LRM) for reconstructing a 3D object from a few unposed images even with little visual overlap, while simultaneously estimating the relative camera poses in ~1.3 seconds on a single A100 GPU. PF-LRM is a highly scalable method utilizing the self-attention blocks to exchange information between 3D object tokens and 2D image tokens; we predict a coarse point cloud for each view, and then use a differentiable Perspective-n-Point (PnP) solver to obtain camera poses. When trained on a huge amount of multi-view posed data of ~1M objects, PF-LRM shows strong cross-dataset generalization ability, and outperforms baseline methods by a large margin in terms of pose prediction accuracy and 3D reconstruction quality on various unseen evaluation datasets. We also demonstrate our model's applicability in downstream text/image-to-3D task with fast feed-forward inference. Our project website is at: https://totoro97.github.io/pf-lrm .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    CORGI reconstructs high-fidelity, animatable 3D dogs from a single in-the-wild image via canonical orbital generation, deformable 3DGS anchored to D-SMAL, and self-supervised generative repair, without 3D supervision.

  2. DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos

    cs.GR 2025-06 conditional novelty 7.0 of 10

    A single feed-forward transformer predicts per-pixel deformable 3D Gaussians with dense scene flow from a posed monocular video, enabling real-time dynamic view synthesis and 3D tracking.

  3. Can Generative Video Models Help Pose Estimation?

    cs.CV 2024-12 conditional novelty 7.0 of 10

    Generating intermediate frames with a video model and feeding them to DUSt3R improves pairwise pose estimation for low-overlap images, when a medoid-based self-consistency score selects the best generated video.

  4. DynSUP: Dynamic Gaussian Splatting from An Unposed Image Pair

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A pose-free two-image pipeline decomposes a dynamic scene into rigid objects and fits per-Gaussian SE(3) motions to synthesize novel views of moving scenes.

  5. GeoWorldAD: Geometry World Action Model for Autonomous Driving

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Grounding an autonomous-driving action model in ego-aligned multi-scale 3D geometry and latent future-geometry tokens improves NAVSIM closed-loop PDMS/EPDMS over prior geometry- and world-model-based planners.

  6. Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A diffusion model guided by hand-object interaction and geometric cues reconstructs 3D hand-held object geometry from monocular RGB images.

  7. MagicHOI: Leveraging 3D Priors for Accurate Hand-object Reconstruction from Short Monocular Video Clips

    cs.CV 2025-08 conditional novelty 6.0 of 10

    MagicHOI integrates a novel view synthesis diffusion prior with a visibility-aware weighting strategy to reconstruct accurate hand-object 3D shapes from short monocular videos with partial object visibility.

  8. DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DualMat is a dual-path diffusion model combining an albedo-optimized pretrained latent path with a material-specialized compact latent path, using feature distillation and rectified flow to estimate PBR materials from...

  9. Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Ultra3D speeds up sparse-voxel 3D generation by generating a coarse mesh with compact VecSet latents, then refining voxel features with part-localized attention.

  10. MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MMIG-Bench is a unified benchmark of 4,850 prompts and 1,750 reference images with a three-level evaluation suite, including the VQA-based Aspect Matching Score that correlates with human ratings.

  11. X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography

    eess.IV 2025-05 conditional novelty 6.0 of 10

    A large transformer with fixed-voxel Gaussian splatting reconstructs CT volumes from 6-10 X-ray projections in under a second, substantially beating prior sparse-view methods in simulation.

  12. RayZer: A Self-supervised Large View Synthesis Model

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A self-supervised transformer model predicts camera poses and scene features from unposed images and renders novel views, reaching performance on par with pose-supervised baselines.

  13. LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fields

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A feed-forward transformer reconstructs shape, PBR materials, and view-dependent radiance from 3 to 6 posed images in under a second, rivaling slower optimization-based inverse rendering.

  14. Matrix3D: Large Photogrammetry Model All-in-One

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A single multi-modal diffusion transformer trained with masked learning performs pose estimation, depth prediction, and novel view synthesis in one model, reporting SOTA pose and NVS numbers.

  15. UVRM: A Scalable 3D Reconstruction Model from Unposed Videos

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A transformer-based model reconstructs 3D objects from unposed monocular videos, trained with score distillation and iterative diffusion-based pseudo-view augmentation.

  16. Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Dora-VAE uses sharp-edge sampling plus dual cross-attention to match XCube-level reconstruction with 1,280 latent codes; Dora-bench adds complexity tiers and a sharp normal error metric.

  17. Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A finetuned pose-conditioned diffusion model, Drive-1-to-3, synthesizes photorealistic novel views of real vehicles from a single image on Waymo and other driving datasets, beating prior methods on FID and LPIPS.

  18. MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 700K-scene procedural, non-semantic synthetic dataset improves large reconstruction models by 1.2 to 1.8 dB PSNR when combined with real data.

  19. You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale

    cs.CV 2024-12 reject novelty 6.0 of 10

    See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...

  20. Turbo3D: Ultra-fast Text-to-3D Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A text-to-3D system that generates Gaussian splatting assets in 0.35 seconds through dual-teacher distillation and latent-space reconstruction, with quality on par with slower baselines.

  21. SelfSplat: Pose-Free and 3D Prior-Free Generalizable 3D Gaussian Splatting

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SelfSplat jointly predicts depth, camera poses and 3D Gaussians from unposed image triplets, and outperforms prior pose-free baselines on RealEstate10K, ACID and DL3DV.

  22. Generalizable Single-view Object Pose Estimation by Two-side Generating and Matching

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Two-side generation and matching of intermediate views with a score distillation loss improves single-reference object pose estimation under large viewpoint changes.

  23. LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Fine-tuning LLaMA-3.1-8B on an OBJ-as-text dataset lets one chat model both answer questions and generate simple 3D meshes, with no vocabulary expansion.

  24. A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Introduces a multi-role red teaming framework using attacker and jury models that increases attack success rates by up to 7.9% on LLM faithfulness in question-answering tasks.

  25. WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting

    cs.CV 2025-10 conditional novelty 5.0 of 10

    A unified feed-forward network that ingests optional geometric priors and jointly predicts point maps, depth, camera poses, normals, and 3D Gaussians, reporting state-of-the-art results on multiple 3D benchmarks.

  26. ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.

  27. MultiGO: Towards Multi-level Geometry Learning for Monocular 3D Textured Human Reconstruction

    cs.CV 2024-12 conditional novelty 5.0 of 10

    MultiGO combines skeleton, joint, and wrinkle level improvements on a Gaussian-based 3D human reconstruction model and reports SOTA results on CustomHuman and THuman3.0.

  28. ARM: Appearance Reconstruction Model for Relightable 3D Generation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    ARM is a feed-forward model that reconstructs a 3D mesh and PBR texture maps (albedo, roughness, metalness) from sparse-view images, improving texture sharpness and relighting quality over prior single-image-to-3D methods.

Pith tools