REVIEW 28 cited by
PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose a Pose-Free Large Reconstruction Model (PF-LRM) for reconstructing a 3D object from a few unposed images even with little visual overlap, while simultaneously estimating the relative camera poses in ~1.3 seconds on a single A100 GPU. PF-LRM is a highly scalable method utilizing the self-attention blocks to exchange information between 3D object tokens and 2D image tokens; we predict a coarse point cloud for each view, and then use a differentiable Perspective-n-Point (PnP) solver to obtain camera poses. When trained on a huge amount of multi-view posed data of ~1M objects, PF-LRM shows strong cross-dataset generalization ability, and outperforms baseline methods by a large margin in terms of pose prediction accuracy and 3D reconstruction quality on various unseen evaluation datasets. We also demonstrate our model's applicability in downstream text/image-to-3D task with fast feed-forward inference. Our project website is at: https://totoro97.github.io/pf-lrm .
Forward citations
Cited by 28 Pith papers
-
CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild
CORGI reconstructs high-fidelity, animatable 3D dogs from a single in-the-wild image via canonical orbital generation, deformable 3DGS anchored to D-SMAL, and self-supervised generative repair, without 3D supervision.
-
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
A single feed-forward transformer predicts per-pixel deformable 3D Gaussians with dense scene flow from a posed monocular video, enabling real-time dynamic view synthesis and 3D tracking.
-
Can Generative Video Models Help Pose Estimation?
Generating intermediate frames with a video model and feeding them to DUSt3R improves pairwise pose estimation for low-overlap images, when a medoid-based self-consistency score selects the best generated video.
-
DynSUP: Dynamic Gaussian Splatting from An Unposed Image Pair
A pose-free two-image pipeline decomposes a dynamic scene into rigid objects and fits per-Gaussian SE(3) motions to synthesize novel views of moving scenes.
-
GeoWorldAD: Geometry World Action Model for Autonomous Driving
Grounding an autonomous-driving action model in ego-aligned multi-scale 3D geometry and latent future-geometry tokens improves NAVSIM closed-loop PDMS/EPDMS over prior geometry- and world-model-based planners.
-
Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance
A diffusion model guided by hand-object interaction and geometric cues reconstructs 3D hand-held object geometry from monocular RGB images.
-
MagicHOI: Leveraging 3D Priors for Accurate Hand-object Reconstruction from Short Monocular Video Clips
MagicHOI integrates a novel view synthesis diffusion prior with a visibility-aware weighting strategy to reconstruct accurate hand-object 3D shapes from short monocular videos with partial object visibility.
-
DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion
DualMat is a dual-path diffusion model combining an albedo-optimized pretrained latent path with a material-specialized compact latent path, using feature distillation and rectified flow to estimate PBR materials from...
-
Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
Ultra3D speeds up sparse-voxel 3D generation by generating a coarse mesh with compact VecSet latents, then refining voxel features with part-localized attention.
-
MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models
MMIG-Bench is a unified benchmark of 4,850 prompts and 1,750 reference images with a three-level evaluation suite, including the VQA-based Aspect Matching Score that correlates with human ratings.
-
X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography
A large transformer with fixed-voxel Gaussian splatting reconstructs CT volumes from 6-10 X-ray projections in under a second, substantially beating prior sparse-view methods in simulation.
-
RayZer: A Self-supervised Large View Synthesis Model
A self-supervised transformer model predicts camera poses and scene features from unposed images and renders novel views, reaching performance on par with pose-supervised baselines.
-
LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fields
A feed-forward transformer reconstructs shape, PBR materials, and view-dependent radiance from 3 to 6 posed images in under a second, rivaling slower optimization-based inverse rendering.
-
Matrix3D: Large Photogrammetry Model All-in-One
A single multi-modal diffusion transformer trained with masked learning performs pose estimation, depth prediction, and novel view synthesis in one model, reporting SOTA pose and NVS numbers.
-
UVRM: A Scalable 3D Reconstruction Model from Unposed Videos
A transformer-based model reconstructs 3D objects from unposed monocular videos, trained with score distillation and iterative diffusion-based pseudo-view augmentation.
-
Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders
Dora-VAE uses sharp-edge sampling plus dual cross-attention to match XCube-level reconstruction with 1,280 latent codes; Dora-bench adds complexity tiers and a sharp normal error metric.
-
Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
A finetuned pose-conditioned diffusion model, Drive-1-to-3, synthesizes photorealistic novel views of real vehicles from a single image on Waymo and other driving datasets, beating prior methods on FID and LPIPS.
-
MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data
A 700K-scene procedural, non-semantic synthetic dataset improves large reconstruction models by 1.2 to 1.8 dB PSNR when combined with real data.
-
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
See3D proposes a pose-free visual condition for multi-view diffusion trained on web videos, claiming SOTA single- and sparse-view 3D generation, but the evaluation protocol leaks ground-truth information and mixes ben...
-
Turbo3D: Ultra-fast Text-to-3D Generation
A text-to-3D system that generates Gaussian splatting assets in 0.35 seconds through dual-teacher distillation and latent-space reconstruction, with quality on par with slower baselines.
-
SelfSplat: Pose-Free and 3D Prior-Free Generalizable 3D Gaussian Splatting
SelfSplat jointly predicts depth, camera poses and 3D Gaussians from unposed image triplets, and outperforms prior pose-free baselines on RealEstate10K, ACID and DL3DV.
-
Generalizable Single-view Object Pose Estimation by Two-side Generating and Matching
Two-side generation and matching of intermediate views with a score distillation loss improves single-reference object pose estimation under large viewpoint changes.
-
LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models
Fine-tuning LLaMA-3.1-8B on an OBJ-as-text dataset lets one chat model both answer questions and generate simple 3D meshes, with no vocabulary expansion.
-
A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation
Introduces a multi-role red teaming framework using attacker and jury models that increases attack success rates by up to 7.9% on LLM faithfulness in question-answering tasks.
-
WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting
A unified feed-forward network that ingests optional geometric priors and jointly predicts point maps, depth, camera poses, normals, and 3D Gaussians, reporting state-of-the-art results on multiple 3D benchmarks.
-
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.
-
MultiGO: Towards Multi-level Geometry Learning for Monocular 3D Textured Human Reconstruction
MultiGO combines skeleton, joint, and wrinkle level improvements on a Gaussian-based 3D human reconstruction model and reports SOTA results on CustomHuman and THuman3.0.
-
ARM: Appearance Reconstruction Model for Relightable 3D Generation
ARM is a feed-forward model that reconstructs a 3D mesh and PBR texture maps (albedo, roughness, metalness) from sparse-view images, improving texture sharpness and relighting quality over prior single-image-to-3D methods.
Discussion (0). Continue with ORCID to comment.