REVIEW 4 major objections 5 minor 103 references
MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims that a view-conditioned diffusion model equipped with structure-aware inputs, an auxiliary mask-prediction task, and a structure-guided timestep scheduler synthesizes consistent novel views of multi-object indoor scenes…
desk verdict Solid multi-object NVS extension with honest limitations; the indoor-scenes claim outruns the quantitative evidence, which stops at white-background composites. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a re-purposed latent diffusion denoiser with three additions. Structure-aware feature amalgamation VAE-encodes the input image, normalized depth, and instance-mask image and concatenates their latents with the noised target-view latent. An auxiliary head on the final denoising U-Net layer predicts the target-view object mask, supervised with weight $\gamma = 0.1$. The structure-guided timestep sampling scheduler draws $t \sim \mathcal{N}(\mu(s), \sigma)$ with $\sigma = 200$, where $\mu(s)$ decays linearly from 1000 to 500 over training, so early training emphasizes global object placement and later training emphasizes fine-grained geometry and appearance. The paper shows through DDIM rollouts that global placement is fixed at large timesteps while mask boundaries sharpen only at small timesteps, which is the observation the scheduler is designed around.
What would settle it
Take a real indoor dataset with ground-truth novel views and per-instance masks, run the released MOVIS model with DepthFM/SAM estimates, and compute foreground IoU and MASt3R Hit Rate against ground truth; if under realistic clutter and backgrounds these metrics fall to the level of the Zero-1-to-3 baseline, the title-level claim about indoor scenes would be unsupported. A cheaper check is to compare MOVIS outputs using ground-truth depth and masks versus DepthFM/SAM estimates on the same inputs; a large drop would indicate the transfer depends on the estimators rather than on learned structure awareness.
Extended reading notes
Core claim
On the paper's own terms, MOVIS establishes that a view-conditioned diffusion model can move from single-object to multi-object novel view synthesis when it is given structural information and trained to reproduce structure. The central empirical claim is that MOVIS substantially outperforms Zero-1-to-3, ZeroNVS, and Free3D on multi-object synthesis and cross-view consistency. On the paper's C3DFS test set it reports PSNR 17.432 versus 14.811 for the best baseline, foreground IoU 58.1 versus 34.4, and a MASt3R-based Hit Rate of 19.3 versus 4.8, with gains persisting on the Objaverse and Room-Texture generalization sets. The paper attributes these gains to three design choices and isolates them in ablations: removing the scheduler drops IoU from 58.1 to 49.1, removing mask prediction to 54.7, and removing depth input to 57.2.
Load-bearing premise
The load-bearing premise is that training on synthetic white-background composites of three to six furniture objects, with ground-truth depth and masks during training and DepthFM/SAM estimates during inference, transfers to real indoor scenes with backgrounds and clutter; the paper tests SUNRGB-D only qualitatively.
Editorial extensions
If this is right
- A single image of several objects can be turned into a new viewpoint with each object retaining its identity, position, and rough geometry, which is what image-to-3D and scene reconstruction pipelines need as a first stage.
- The three components do independent work: ablations show the scheduler matters most, then mask prediction, then depth input, so future models can adopt components selectively.
- Because the auxiliary head outputs novel-view masks, editing operations like object removal under a new viewpoint follow directly from thresholding the predicted mask.
- The proposed Hit Rate and nearest-matching-distance metrics give a quantitative handle on cross-view consistency that PSNR, SSIM, and LPIPS miss, and can be applied to any novel-view synthesis method.
- Demonstrated generalization to Objaverse, Room-Texture, 3D-FRONT, and SUNRGB-D (the real dataset tested qualitatively) suggests a model trained only on synthetic furniture composites can transfer when background is not the focus.
Reading between the lines
- The paper's own limitation statement concedes that multi-view consistency among synthesized images is not guaranteed and background texture is not modeled; a natural next test is whether adding background modeling or training on real RGB-D scenes with backgrounds removes the synthetic-to-real gap in quantitative metrics.
- The timestep-schedule insight is likely not specific to novel view synthesis: any diffusion task with a coarse-to-fine structure, such as layout-conditioned generation or compositional image editing, could benefit from a mean-shifting noise schedule that first forces global arrangement and later allocates capacity to details.
- The foreground-IoU and MASt3R matching metrics could double as a lightweight proxy for 3D awareness in other single-image generative models, since they detect whether objects move or deform correctly with viewpoint.
- Because inference on real images depends on DepthFM and SAM estimates, the method inherits their failure modes; comparing outputs with ground-truth versus estimated depth and masks would show how much of the reported gap is due to the structure conditions themselves.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MOVIS, a view-conditioned diffusion model for multi-object novel view synthesis (NVS). The method augments a Stable Diffusion backbone with three components: (i) structure-aware input conditioning using the input-view depth map and instance mask, (ii) an auxiliary task that predicts the target-view object mask from the U-Net's final features, and (iii) a structure-guided timestep sampling scheduler whose Gaussian mean decays linearly during training to shift from global placement learning to fine-grained detail recovery. The authors introduce a new synthetic dataset, C3DFS, composed of 3-6 furniture objects on white backgrounds, and propose two cross-view consistency metrics (Hit Rate and Nearest Matching Distance) based on MASt3R matching, in addition to foreground IoU. They compare against Zero-1-to-3, ZeroNVS, and Free3D on C3DFS, Objaverse composites, and Room-Texture, with qualitative results on 3D-FRONT and SUNRGB-D, and report consistent improvements in PSNR/SSIM/LPIPS, IoU, and cross-view consistency.
Significance. If the transfer to real indoor scenes holds, the work is a useful practical contribution: the structure conditioning, auxiliary mask prediction, and curriculum-style timestep sampling are simple and appear to yield large gains on multi-object placement and cross-view consistency relative to strong baselines. The C3DFS dataset with disjoint furniture splits is a valuable benchmark, and the MASt3R-based evaluation protocol addresses a real gap in the NVS literature. The main caveat is that all quantitative evidence is confined to synthetic composites with white backgrounds; the real-world and full-scene claims rest on qualitative demonstrations only. The paper also ships a clear ablation study showing each component contributes, though without error bars or significance tests.
major comments (4)
- [Sec. 4.1 (Datasets), Table 1, and Fig. 4] All quantitative results -- image-level metrics, foreground IoU, Hit Rate, and Dist -- are computed on synthetic composites rendered on white backgrounds (C3DFS, Objaverse, Room-Texture). The real-world SUNRGB-D and synthetic 3D-FRONT appear only in qualitative figures (Fig. 4 and Figs. S.9, S.10). Because the title and abstract claim multi-object NVS for indoor scenes and the introduction emphasizes generalization to realistic datasets, the absence of quantitative evaluation on a full-scene dataset is a load-bearing gap. I request a quantitative evaluation on 3D-FRONT, which has renderable ground truth, and preferably also on SUNRGB-D with an adapted protocol (e.g., SAM-based masks and background-agnostic metrics).
- [Sec. B.3 (Metrics, IoU)] The foreground IoU metric computes the foreground mask by thresholding the generated image as M = IL < 250, explicitly exploiting the white background of the synthetic composites. This metric cannot be applied to real indoor images containing walls, floors, and clutter. Consequently, the placement evidence for the paper's central claim is limited to the synthetic white-background setting. Please add a background-agnostic placement metric, such as per-instance IoU using SAM-estimated masks, and report it on at least one scene-level dataset.
- [Tables 1, 2, and S.5] All numerical results are single-run point estimates with no error bars, confidence intervals, or significance tests. Some improvements are modest (e.g., PSNR 10.014 vs. 9.623 on Room-Texture), so it is important to verify that the reported gains are stable across at least two or three training seeds, or to provide a statistical test over test-set samples, before concluding that MOVIS 'significantly outperforms' the baselines.
- [Sec. 3.3 and Table S.3] The structure-guided scheduler has several hand-chosen hyperparameters (mu_global=1000, mu_local=500, sigma=200, warmup of 4000 steps, decay over 2000 steps, final 6000 steps). The ablation compares only the linear-decay schedule against three alternatives and is performed only on C3DFS. Since the scheduler is a central contribution, please include a small sensitivity analysis (varying sigma and the decay length) and at least one result on a held-out dataset such as Objaverse to show that the choice is not overfit to C3DFS.
minor comments (5)
- [Algorithm 1] Line 13 contains a typo: 'Hits ← −Hits + 1' should be 'Hits ← Hits + 1'.
- [Sec. 4.1] The sentence 'We also evaluate our model on diverse indoor scenes from both the synthetic dataset 3D-FRONT and the real-world dataset SUNRGB-D' is imprecise: the evaluation on these datasets is qualitative only, while the quantitative evaluation is on synthetic composites. Please reword to distinguish these clearly.
- [Sec. 3.2, Eq. (3)] The auxiliary mask loss is applied to latent mask features, as described in Sec. A.3, but the main text does not specify that the supervision is in latent space. Please state this explicitly in the main text to avoid the impression that the loss is computed on decoded masks.
- [Fig. 1 caption] The caption appears to contain the run-together text 'NVSCross View Matching'; please insert a space.
- [Sec. 4.1] The sentence 'This choice stems from the recent advancements in object segmentation [33], while we leave the background modeling for future work' is a significant scope limitation that is deferred to the supplementary material. I recommend stating this limitation prominently in the main text, since it directly bears on the indoor-scene claim.
Circularity Check
No circularity: model outputs are evaluated with external metrics and disjoint furniture splits; the synthetic-only evidence for real indoor scenes is a generalization gap, not a circular step.
full rationale
The paper's central claims are empirical rather than derivational, and I find no step where a claimed prediction reduces by construction to its own inputs. The model (Eqs. 2-3 of Sec. 3.2) is trained with depth and mask conditioning plus an auxiliary target-mask prediction task; these are additional supervision signals, not re-statements of the evaluated outputs. The cross-view consistency metrics (Sec. 4.1 and Sec. B.3) use the external matcher MASt3R on ground-truth and predicted images, with Hit Rate and Nearest Matching Distance computed against ground-truth correspondences, so the evaluation is not defined in terms of the model's own outputs. The structure-guided timestep scheduler (Sec. 3.3) is an ablation-tested training-strategy choice with the hyperparameters (mu_global=1000, mu_local=500, sigma=200) justified by inference visualizations and compared against alternatives on a test set whose furniture instances are disjoint from training; this is hyperparameter selection, not a fitted parameter being renamed as a prediction. Self-citations appear in related work and in dataset curation (e.g., the seven furniture categories 'following previous work [6]' in Sec. B.2), but none is load-bearing for the central NVS claim, and no uniqueness theorem is imported from the authors' prior work. The manuscript's own limitation statements (Sec. C: background texture is not modeled; multi-view consistency is not guaranteed) and the fact that all quantitative results are on white-background synthetic composites, with SUNRGB-D and 3D-FRONT shown only qualitatively, weaken the title-level generalization claim, but this is an evidence/scope gap rather than a circular argument. Accordingly, the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (6)
- Auxiliary mask loss weight gamma =
0.1
- Timestep sampler mean endpoints mu_global and mu_local =
1000 -> 500
- Timestep sampler variance sigma =
200
- Curriculum schedule durations =
warmup 4000, decay 2000, stabilize from 6000 of 8000 steps
- Inference settings =
50 DDIM steps, guidance scale 3.0, 256x256
- Foreground mask threshold beta_th =
250
assumptions (6)
- standard math DDPM/DDIM denoising objective and noise schedule are valid for view-conditioned fine-tuning.
- domain assumption Pre-trained Stable Diffusion and DINOv2 features provide transferable image priors for indoor furniture composition.
- domain assumption Input-view depth and instance mask are sufficient structural conditions for novel-view object placement.
- domain assumption Synthetic C3DFS composites of 3 to 6 furniture items on white backgrounds are a valid proxy for indoor multi-object scenes.
- domain assumption MASt3R correspondences between input and target views are a reliable ground-truth reference for cross-view consistency.
- ad hoc to paper A Gaussian timestep sampler with linearly decaying mean implements a curriculum from global placement to local detail.
Cite this review
Pith. "Pith review of MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes." pith.science (2026). https://pith.science/paper/E6RD57DY
@misc{pith2026241211457,
author = {Pith},
title = {Pith review of: MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes},
year = {2026},
howpublished = {\url{https://pith.science/paper/E6RD57DY}},
note = {Machine review of arXiv:2412.11457}
}
read the original abstract
Repurposing pre-trained diffusion models has been proven to be effective for NVS. However, these methods are mostly limited to a single object; directly applying such methods to compositional multi-object scenarios yields inferior results, especially incorrect object placement and inconsistent shape and appearance under novel views. How to enhance and systematically evaluate the cross-view consistency of such models remains under-explored. To address this issue, we propose MOVIS to enhance the structural awareness of the view-conditioned diffusion model for multi-object NVS in terms of model inputs, auxiliary tasks, and training strategy. First, we inject structure-aware features, including depth and object mask, into the denoising U-Net to enhance the model's comprehension of object instances and their spatial relationships. Second, we introduce an auxiliary task requiring the model to simultaneously predict novel view object masks, further improving the model's capability in differentiating and placing objects. Finally, we conduct an in-depth analysis of the diffusion sampling process and carefully devise a structure-guided timestep sampling scheduler during training, which balances the learning of global object placement and fine-grained detail recovery. To systematically evaluate the plausibility of synthesized images, we propose to assess cross-view consistency and novel view object placement alongside existing image-level NVS metrics. Extensive experiments on challenging synthetic and realistic datasets demonstrate that our method exhibits strong generalization capabilities and produces consistent novel view synthesis, highlighting its potential to guide future 3D-aware multi-object NVS tasks. Our project page is available at https://jason-aplp.github.io/MOVIS/.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Omni3d: A large benchmark and model for 3d object detection in the wild
Garrick Brazil, Abhinav Kumar, Julian Straub, Nikhila Ravi, Justin Johnson, and Georgia Gkioxari. Omni3d: A large benchmark and model for 3d object detection in the wild. In Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
2023
-
[2]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. InConfer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[3]
Scenetex: High-quality texture synthesis for indoor scenes via diffusion priors
Dave Zhenyu Chen, Haoxuan Li, Hsin-Ying Lee, Sergey Tulyakov, and Matthias Nießner. Scenetex: High-quality texture synthesis for indoor scenes via diffusion priors. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[4]
On the importance of noise scheduling for diffu- sion models
Ting Chen. On the importance of noise scheduling for diffu- sion models. arXiv preprint arXiv:2301.10972, 2023. 5
arXiv 2023
-
[5]
Cascade-zero123: One image to highly consistent 3d with self-prompted nearby views
Yabo Chen, Jiemin Fang, Yuyang Huang, Taoran Yi, Xi- aopeng Zhang, Lingxi Xie, Xinggang Wang, Wenrui Dai, Hongkai Xiong, and Qi Tian. Cascade-zero123: One image to highly consistent 3d with self-prompted nearby views. In European Conference on Computer Vision (ECCV), 2024. 2, 3
2024
-
[6]
Single-view 3d scene reconstruc- tion with high-fidelity shape and texture
Yixin Chen, Junfeng Ni, Nan Jiang, Yaowei Zhang, Yixin Zhu, and Siyuan Huang. Single-view 3d scene reconstruc- tion with high-fidelity shape and texture. In International Conference on 3D Vision (3DV), 2024. 3, 14
2024
-
[7]
Comboverse: Compositional 3d assets creation using spatially-aware diffusion guidance
Yongwei Chen, Tengfei Wang, Tong Wu, Xingang Pan, Kui Jia, and Ziwei Liu. Comboverse: Compositional 3d assets creation using spatially-aware diffusion guidance. In Euro- pean Conference on Computer Vision (ECCV), 2024. 3
2024
-
[8]
Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. In European Conference on Computer Vision (ECCV), 2024. 3
2024
Show all 103 references
-
[9]
Blender - a 3D modelling and rendering package
Blender Online Community. Blender - a 3D modelling and rendering package. Blender Foundation, Stichting Blender Foundation, Amsterdam, 2018. 14
2018
-
[10]
Anyskill: Learning open- vocabulary physical skill for interactive agents
Jieming Cui, Tengyu Liu, Nian Liu, Yaodong Yang, Yixin Zhu, and Siyuan Huang. Anyskill: Learning open- vocabulary physical skill for interactive agents. In Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[11]
Objaverse-xl: A universe of 10m+ 3d objects
Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram V oleti, Samir Yitzhak Gadre, et al. Objaverse-xl: A universe of 10m+ 3d objects. Advances in Neural Informa- tion Processing Systems (NeurIPS), 2023. 3
2023
-
[12]
Objaverse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Conference on Com- puter Vision and Pattern Recognition (CVPR), 2023. 2, ...
2023
-
[13]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Informa- tion Processing Systems (NeurIPS), 2021. 5
2021
-
[14]
General- izable 3d scene reconstruction via divide and conquer from a single view
Andreea Dogaru, Mert Özer, and Bernhard Egger. General- izable 3d scene reconstruction via divide and conquer from a single view. arXiv preprint arXiv:2404.03421, 2024. 3
2024 arXiv
-
[15]
3d-front: 3d furnished rooms with layouts and semantics
Huan Fu, Bowen Cai, Lin Gao, Ling-Xiao Zhang, Jiaming Wang, Cao Li, Qixun Zeng, Chengyue Sun, Rongfei Jia, Bin- qiang Zhao, et al. 3d-front: 3d furnished rooms with layouts and semantics. In International Conference on Computer Vi- sion (ICCV), 2021. 2, 6, 16
2021
-
[16]
3d-future: 3d fur- niture shape with texture
Huan Fu, Rongfei Jia, Lin Gao, Mingming Gong, Binqiang Zhao, Steve Maybank, and Dacheng Tao. 3d-future: 3d fur- niture shape with texture. International Journal of Computer Vision (IJCV), 129:3313–3337, 2021. 6, 14, 15, 17
2021
-
[17]
Arnold: A benchmark for language-grounded task learning with continuous states in realistic 3d scenes
Ran Gong, Jiangyong Huang, Yizhou Zhao, Haoran Geng, Xiaofeng Gao, Qingyang Wu, Wensi Ai, Ziheng Zhou, Demetri Terzopoulos, Song-Chun Zhu, et al. Arnold: A benchmark for language-grounded task learning with continuous states in realistic 3d scenes. arXiv preprint arXiv:2304.04...
2023 arXiv
-
[18]
Depthfm: Fast monocular depth estimation with flow matching
Ming Gui, Johannes S Fischer, Ulrich Prestel, Pingchuan Ma, Dmytro Kotovenko, Olga Grebenkova, Stefan Andreas Baumann, Vincent Tao Hu, and Björn Ommer. Depthfm: Fast monocular depth estimation with flow matching. arXiv preprint arXiv:2403.13788, 2024. 14, 15
2024 arXiv
-
[19]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. Advances in Neural Information Processing Systems (NeurIPS), 2020. 2, 3
2020
-
[20]
Text2room: Extracting textured 3d meshes from 2d text-to-image models
Lukas Höllein, Ang Cao, Andrew Owens, Justin Johnson, and Matthias Nießner. Text2room: Extracting textured 3d meshes from 2d text-to-image models. In International Con- ference on Computer Vision (ICCV), 2023. 2
2023
-
[21]
An embodied generalist agent in 3d world
Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An embodied generalist agent in 3d world. In International Conference on Machine Learning (ICML), 2024. 2
2024
-
[22]
Unveiling the mist over 3d vision-language under- standing: Object-centric evaluation with chain-of-analysis
Jiangyong Huang, Baoxiong Jia, Yan Wang, Ziyu Zhu, Xiongkun Linghu, Qing Li, Song-Chun Zhu, and Siyuan Huang. Unveiling the mist over 3d vision-language under- standing: Object-centric evaluation with chain-of-analysis. In Conference on Computer Vision and Pattern Recognition ...
2025
-
[23]
Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion
Zehuan Huang, Hao Wen, Junting Dong, Yaohui Wang, Yangguang Li, Xinyuan Chen, Yan-Pei Cao, Ding Liang, Yu 9 Qiao, Bo Dai, et al. Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion. In Conference on Computer Vision and Pattern Recognition (CVPR...
2024
-
[24]
Roompainter: View-integrated diffusion for consistent indoor scene texturing
Zhipeng Huang, Wangbo Yu, Xinhua Cheng, ChengShu Zhao, Yunyang Ge, Mingyi Guo, Li Yuan, and Yonghong Tian. Roompainter: View-integrated diffusion for consistent indoor scene texturing. arXiv preprint arXiv:2412.16778 ,
-
[25]
Putting nerf on a diet: Semantically consistent few-shot view synthe- sis
Ajay Jain, Matthew Tancik, and Pieter Abbeel. Putting nerf on a diet: Semantically consistent few-shot view synthe- sis. In International Conference on Computer Vision (ICCV),
-
[26]
Zero-shot text-guided object gener- ation with dream fields
Ajay Jain, Ben Mildenhall, Jonathan T Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object gener- ation with dream fields. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
2022
-
[27]
Sceneverse: Scaling 3d vision-language learning for grounded scene understanding
Baoxiong Jia, Yixin Chen, Huangyue Yu, Yan Wang, Xuesong Niu, Tengyu Liu, Qing Li, and Siyuan Huang. Sceneverse: Scaling 3d vision-language learning for grounded scene understanding. In European Conference on Computer Vision (ECCV), 2024. 2
2024
-
[28]
Autonomous character-scene interaction synthesis from text instruction
Nan Jiang, Zimo He, Zi Wang, Hongjie Li, Yixin Chen, Siyuan Huang, and Yixin Zhu. Autonomous character-scene interaction synthesis from text instruction. In SIGGRAPH Asia 2024 Conference Papers, 2024. 2
2024
-
[29]
Scaling up dynamic human-scene interaction model- ing
Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction model- ing. In Conference on Computer Vision and Pattern Recog- nition (CVPR), 2024. 2
2024
-
[30]
Efficient- 3dim: Learning a generalizable single-image novel-view synthesizer in one day
Yifan Jiang, Hao Tang, Jen-Hao Rick Chang, Liangchen Song, Zhangyang Wang, and Liangliang Cao. Efficient- 3dim: Learning a generalizable single-image novel-view synthesizer in one day. arXiv preprint arXiv:2310.03015 ,
-
[31]
Repurpos- ing diffusion-based image generators for monocular depth estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- zger, Rodrigo Caye Daudt, and Konrad Schindler. Repurpos- ing diffusion-based image generators for monocular depth estimation. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 3, 5, 8
2024
-
[32]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics ,
-
[33]
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In International Conference on Computer Vision (ICCV), 2023. 3, 5, 6, 8, 14, 15
2023
-
[34]
Eschernet: A generative model for scalable view synthesis
Xin Kong, Shikun Liu, Xiaoyang Lyu, Marwan Taher, Xi- aojuan Qi, and Andrew J Davison. Eschernet: A generative model for scalable view synthesis. In Conference on Com- puter Vision and Pattern Recognition (CVPR), 2024. 13, 19
2024
-
[35]
Ground- ing image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and Jérôme Revaud. Ground- ing image matching in 3d with mast3r. In European Confer- ence on Computer Vision (ECCV), 2024. 2, 8, 15
2024
-
[36]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In In- ternational Conference on Machine Learning (ICML), 2023. 3
2023
-
[37]
Gendexgrasp: General- izable dexterous grasping
Puhao Li, Tengyu Liu, Yuyang Li, Yiran Geng, Yixin Zhu, Yaodong Yang, and Siyuan Huang. Gendexgrasp: General- izable dexterous grasping. In International Conference on Robotics and Automation (ICRA), 2023. 2
2023
-
[38]
Ag2manip: Learning novel manipulation skills with agent- agnostic visual and action representations
Puhao Li, Tengyu Liu, Yuyang Li, Muzhi Han, Haoran Geng, Shu Wang, Yixin Zhu, Song-Chun Zhu, and Siyuan Huang. Ag2manip: Learning novel manipulation skills with agent- agnostic visual and action representations. In International Conference on Intelligent Robots and Systems (IR...
2024
-
[39]
Flowbothd: History-aware diffuser handling ambiguities in articulated objects manipulation
Yishu Li, Wen Hui Leng, Yiming Fang, Ben Eisner, and David Held. Flowbothd: History-aware diffuser handling ambiguities in articulated objects manipulation. arXiv preprint arXiv:2410.07078, 2024
2024 arXiv
-
[40]
Grasp multi- ple objects with one hand
Yuyang Li, Bo Liu, Yiran Geng, Puhao Li, Yaodong Yang, Yixin Zhu, Tengyu Liu, and Siyuan Huang. Grasp multi- ple objects with one hand. IEEE Robotics and Automation Letters (RA-L), 2024. 2
2024
-
[41]
Magic3d: High-resolution text-to-3d content creation
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2023. 2
2023
-
[42]
Consistent123: One image to highly consistent 3d asset using case-aware diffusion priors
Yukang Lin, Haonan Han, Chaoqun Gong, Zunnan Xu, Yachao Zhang, and Xiu Li. Consistent123: One image to highly consistent 3d asset using case-aware diffusion priors. In Proceedings of the 32nd ACM International Conference on Multimedia, 2024. 2, 3
2024
-
[43]
Multi-modal situated reasoning in 3d scenes.Advances in Neural Information Pro- cessing Systems (NeurIPS), 2024
Xiongkun Linghu, Jiangyong Huang, Xuesong Niu, Xiaojian Ma, Baoxiong Jia, and Siyuan Huang. Multi-modal situated reasoning in 3d scenes.Advances in Neural Information Pro- cessing Systems (NeurIPS), 2024. 2
2024
-
[44]
One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion
Minghua Liu, Chao Xu, Haian Jin, Linghao Chen, Mukund Varma T, Zexiang Xu, and Hao Su. One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion. Advances in Neural Information Processing Systems (NeurIPS), 2023. 2, 3
2023
-
[45]
One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion
Minghua Liu, Ruoxi Shi, Linghao Chen, Zhuoyang Zhang, Chao Xu, Xinyue Wei, Hansheng Chen, Chong Zeng, Ji- ayuan Gu, and Hao Su. One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion. In Conference on Computer Vision and Pattern Re...
2024
-
[46]
Zero-1-to-3: Zero-shot one image to 3d object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3d object. In International Confer- ence on Computer Vision (ICCV), 2023. 1, 2, 3, 4, 6, 13
2023
-
[47]
Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age
Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age. arXiv preprint arXiv:2309.03453, 2023. 2, 3, 19
2023 arXiv
-
[48]
Slotlifter: Slot-guided feature lifting for learning object- 10 centric radiance fields
Yu Liu, Baoxiong Jia, Yixin Chen, and Siyuan Huang. Slotlifter: Slot-guided feature lifting for learning object- 10 centric radiance fields. In European Conference on Com- puter Vision (ECCV), 2024. 3
2024
-
[49]
Building interactable replicas of complex articulated objects via gaussian splatting
Yu Liu, Baoxiong Jia, Ruijie Lu, Junfeng Ni, Song-Chun Zhu, and Siyuan Huang. Building interactable replicas of complex articulated objects via gaussian splatting. arXiv preprint arXiv:2502.19459, 2025. 2
2025 arXiv
-
[50]
Wonder3d: Sin- gle image to 3d using cross-domain diffusion
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al. Wonder3d: Sin- gle image to 3d using cross-domain diffusion. InConference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[51]
Manigaussian: Dynamic gaus- sian splatting for multi-task robotic manipulation
Guanxing Lu, Shiyi Zhang, Ziwei Wang, Changliu Liu, Ji- wen Lu, and Yansong Tang. Manigaussian: Dynamic gaus- sian splatting for multi-task robotic manipulation. In Euro- pean Conference on Computer Vision (ECCV), 2024. 2
2024
-
[52]
Taco: Taming diffusion for in-the-wild video amodal completion
Ruijie Lu, Yixin Chen, Yu Liu, Jiaxiang Tang, Junfeng Ni, Diwen Wan, Gang Zeng, and Siyuan Huang. Taco: Taming diffusion for in-the-wild video amodal completion. arXiv preprint arXiv:2503.12049, 2025. 3
2025 arXiv
-
[53]
Repaint: Inpainting using denoising diffusion probabilistic models
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. In Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[54]
Unsuper- vised discovery of object-centric neural fields.arXiv preprint arXiv:2402.07376, 2024
Rundong Luo, Hong-Xing Yu, and Jiajun Wu. Unsuper- vised discovery of object-centric neural fields.arXiv preprint arXiv:2402.07376, 2024. 2, 6, 7, 16, 17
2024 arXiv
-
[55]
Realfusion: 360deg reconstruction of any object from a single image
Luke Melas-Kyriazi, Iro Laina, Christian Rupprecht, and Andrea Vedaldi. Realfusion: 360deg reconstruction of any object from a single image. In Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2023. 3
2023
-
[56]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European Conference on Computer Vision (ECCV), 2020. 3
2020
-
[57]
Phyrecon: Physically plausible neural scene recon- struction
Junfeng Ni, Yixin Chen, Bohan Jing, Nan Jiang, Bin Wang, Bo Dai, Puhao Li, Yixin Zhu, Song-Chun Zhu, and Siyuan Huang. Phyrecon: Physically plausible neural scene recon- struction. In Advances in Neural Information Processing Systems (NeurIPS), 2024. 3
2024
-
[58]
Decompositional neu- ral scene reconstruction with generative diffusion prior
Junfeng Ni, Yu Liu, Ruijie Lu, Zirui Zhou, Song-Chun Zhu, Yixin Chen, and Siyuan Huang. Decompositional neu- ral scene reconstruction with generative diffusion prior. In Conference on Computer Vision and Pattern Recognition (CVPR), 2025. 3
2025
-
[59]
Total3dunderstanding: Joint lay- out, object pose and mesh reconstruction for indoor scenes from a single image
Yinyu Nie, Xiaoguang Han, Shihui Guo, Yujian Zheng, Jian Chang, and Jian Jun Zhang. Total3dunderstanding: Joint lay- out, object pose and mesh reconstruction for indoor scenes from a single image. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 3
2020
-
[60]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 5, 13
2023 arXiv
-
[61]
pix2gestalt: Amodal segmentation by synthesizing wholes
Ege Ozguroglu, Ruoshi Liu, Dídac Surís, Dian Chen, Achal Dave, Pavel Tokmakov, and Carl V ondrick. pix2gestalt: Amodal segmentation by synthesizing wholes. In Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[62]
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 3
2022 arXiv
-
[63]
Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors.arXiv preprint arXiv:2306.17843,
Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Sko- rokhodov, Peter Wonka, Sergey Tulyakov, et al. Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors.arXiv preprint arXiv:2306.17843,
-
[64]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International Conference on Machine Learning...
2021
-
[65]
Language embedded radiance fields for zero-shot task- oriented grasping
Adam Rashid, Satvik Sharma, Chung Min Kim, Justin Kerr, Lawrence Yunliang Chen, Angjoo Kanazawa, and Ken Gold- berg. Language embedded radiance fields for zero-shot task- oriented grasping. In Conference on Robot Learning (CoRL),
-
[66]
High-resolution image syn- thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image syn- thesis with latent diffusion models. In Conference on Com- puter Vision and Pattern Recognition (CVPR) , 2022. 2, 3, 4
2022
-
[67]
Zeronvs: Zero-shot 360- degree view synthesis from a single image
Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry La- gun, Li Fei-Fei, Deqing Sun, et al. Zeronvs: Zero-shot 360- degree view synthesis from a single image. InConference on Computer Vision and Pattern Recognition (CVPR)...
2024
-
[68]
Zero123++: a single image to consistent multi-view dif- fusion base model
Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. Zero123++: a single image to consistent multi-view dif- fusion base model. arXiv preprint arXiv:2310.15110, 2023. 2, 3, 19
-
[69]
Mvdream: Multi-view diffusion for 3d gen- eration
Yichun Shi, Peng Wang, Jianglong Ye, Long Mai, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d gen- eration. arXiv:2308.16512, 2023. 2, 19
2023 arXiv
-
[70]
Denois- ing diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models. In International Conference on Learning Representations (ICLR), 2020. 5
2020
-
[71]
Sun rgb-d: A rgb-d scene understanding benchmark suite
Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In Conference on Computer Vision and Pattern Recognition (CVPR), 2015. 3, 6, 16, 17
2015
-
[72]
Dreamgaussian: Generative gaussian splatting for effi- cient 3d content creation
Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. Dreamgaussian: Generative gaussian splatting for effi- cient 3d content creation. arXiv preprint arXiv:2309.16653,
-
[73]
Make-it-3d: High-fidelity 3d 11 creation from a single image with diffusion prior
Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, and Dong Chen. Make-it-3d: High-fidelity 3d 11 creation from a single image with diffusion prior. In Inter- national Conference on Computer Vision (ICCV) , 2023. 2, 3
2023
-
[74]
Lgm: Large multi-view gaussian model for high-resolution 3d content creation
Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. In European Conference on Computer Vision (ECCV), 2024. 2
2024
-
[75]
Intex: Interactive text-to-texture syn- thesis via unified depth-aware inpainting
Jiaxiang Tang, Ruijie Lu, Xiaokang Chen, Xiang Wen, Gang Zeng, and Ziwei Liu. Intex: Interactive text-to-texture syn- thesis via unified depth-aware inpainting. arXiv preprint arXiv:2403.11878, 2024. 3
2024 arXiv
-
[76]
Megascenes: Scene-level view synthesis at scale
Joseph Tung, Gene Chou, Ruojin Cai, Guandao Yang, Kai Zhang, Gordon Wetzstein, Bharath Hariharan, and Noah Snavely. Megascenes: Scene-level view synthesis at scale. In European Conference on Computer Vision (ECCV), 2024. 2, 3
2024
-
[77]
Tracking through containers and occlud- ers in the wild
Basile Van Hoorick, Pavel Tokmakov, Simon Stent, Jie Li, and Carl V ondrick. Tracking through containers and occlud- ers in the wild. In Conference on Computer Vision and Pat- tern Recognition (CVPR), 2023. 16
2023
-
[78]
Generative camera dolly: Ex- treme monocular dynamic novel view synthesis
Basile Van Hoorick, Rundi Wu, Ege Ozguroglu, Kyle Sar- gent, Ruoshi Liu, Pavel Tokmakov, Achal Dave, Changxi Zheng, and Carl V ondrick. Generative camera dolly: Ex- treme monocular dynamic novel view synthesis. In Euro- pean Conference on Computer Vision (ECCV), 2024. 3
2024
-
[79]
Imagedream: Image-prompt multi-view diffusion for 3d generation
Peng Wang and Yichun Shi. Imagedream: Image-prompt multi-view diffusion for 3d generation. arXiv preprint arXiv:2312.02201, 2023. 3, 19
2023 arXiv
-
[80]
Barron, Ricardo Martin- Brualla, Noah Snavely, and Thomas Funkhouser
Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul Srini- vasan, Howard Zhou, Jonathan T. Barron, Ricardo Martin- Brualla, Noah Snavely, and Thomas Funkhouser. Ibrnet: Learning multi-view image-based rendering. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
2021
-
[81]
Roomtex: Textur- ing compositional indoor scenes via iterative inpainting
Qi Wang, Ruijie Lu, Xudong Xu, Jingbo Wang, Michael Yu Wang, Bo Dai, Gang Zeng, and Dan Xu. Roomtex: Textur- ing compositional indoor scenes via iterative inpainting. In European Conference on Computer Vision (ECCV), 2024. 2, 3
2024
-
[82]
Dust3r: Geometric 3d vi- sion made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vi- sion made easy. In Conference on Computer Vision and Pat- tern Recognition (CVPR), 2024. 2, 16
2024
-
[83]
Novel view synthesis with diffusion models
Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models. arXiv preprint arXiv:2210.04628, 2022. 4
2022 arXiv
-
[84]
Consistent123: Improve consistency for one image to 3d object synthesis
Haohan Weng, Tianyu Yang, Jianan Wang, Yu Li, Tong Zhang, CL Chen, and Lei Zhang. Consistent123: Improve consistency for one image to 3d object synthesis. arXiv preprint arXiv:2310.08092, 2023. 2, 3
2023 arXiv
-
[85]
Con- vnext v2: Co-designing and scaling convnets with masked autoencoders
Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie. Con- vnext v2: Co-designing and scaling convnets with masked autoencoders. In Conference on Computer Vision and Pat- tern Recognition (CVPR), 2023. 13
2023
-
[86]
Unique3d: High-quality and efficient 3d mesh generation from a single image
Kailu Wu, Fangfu Liu, Zhihan Cai, Runjie Yan, Hanyang Wang, Yating Hu, Yueqi Duan, and Kaisheng Ma. Unique3d: High-quality and efficient 3d mesh generation from a single image. In Advances in Neural Information Processing Sys- tems (NeurIPS), 2024. 3
2024
-
[87]
Structured 3d latents for scalable and versatile 3d gen- eration
Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, and Jiaolong Yang. Structured 3d latents for scalable and versatile 3d gen- eration. arXiv preprint arXiv:2412.01506, 2024. 2
2024 arXiv
-
[88]
Neurallift-360: Lifting an in-the-wild 2d photo to a 3d object with 360deg views
Dejia Xu, Yifan Jiang, Peihao Wang, Zhiwen Fan, Yi Wang, and Zhangyang Wang. Neurallift-360: Lifting an in-the-wild 2d photo to a 3d object with 360deg views. InConference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
2023
-
[89]
Amodal com- pletion via progressive mixed context diffusion
Katherine Xu, Lingzhi Zhang, and Jianbo Shi. Amodal com- pletion via progressive mixed context diffusion. In Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 9099–9109, 2024. 3, 16
2024
-
[90]
Consistnet: Enforcing 3d consistency for multi- view images diffusion
Jiayu Yang, Ziang Cheng, Yunfei Duan, Pan Ji, and Hong- dong Li. Consistnet: Enforcing 3d consistency for multi- view images diffusion. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 19
2024
-
[91]
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 3
2024
-
[92]
Physcene: Physically interactable 3d scene synthe- sis for embodied ai
Yandan Yang, Baoxiong Jia, Peiyuan Zhi, and Siyuan Huang. Physcene: Physically interactable 3d scene synthe- sis for embodied ai. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[93]
Scannet++: A high-fidelity dataset of 3d indoor scenes
Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai. Scannet++: A high-fidelity dataset of 3d indoor scenes. In International Conference on Computer Vi- sion (ICCV), 2023. 8
2023
-
[94]
pixelnerf: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
2021
-
[95]
Amodal ground truth and completion in the wild
Guanqi Zhan, Chuanxia Zheng, Weidi Xie, and Andrew Zis- serman. Amodal ground truth and completion in the wild. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 3, 16
2024
-
[96]
Self-supervised scene de- occlusion
Xiaohang Zhan, Xingang Pan, Bo Dai, Ziwei Liu, Dahua Lin, and Chen Change Loy. Self-supervised scene de- occlusion. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 16
2020
-
[97]
Free3d: Consistent novel view synthesis without 3d representation
Chuanxia Zheng and Andrea Vedaldi. Free3d: Consistent novel view synthesis without 3d representation. In Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[98]
Stereo magnification: Learning view syn- thesis using multiplane images
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view syn- thesis using multiplane images. In SIGGRAPH, 2018. 8
2018
-
[99]
norm patchtokens
Yan Zhu, Yuandong Tian, Dimitris Metaxas, and Piotr Dol- lár. Semantic amodal segmentation. In Conference on Com- puter Vision and Pattern Recognition (CVPR), 2017. 16 12 MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes Supplementary Material A. Model detai...
2017
-
[100]
FXRvdpBKje59HAj3tIVdwRQpPk0=
in Fig. S.17, and on Room-Texture [54] in Fig. S.18. More visualized comparisons with baselines on Room- Texture [54], SUNRGB-D [71] and 3D-FRONT [15] are shown in Fig. S.9. More results on in-the-wild datasets are shown in Fig. S.8.A more complete ablation study on other data...
-
[101]
If an object’s visible mask is exactly its full mask, there exists no occlusion
-
[102]
If an object’s visible mask is more than 70% of its full mask, the object is occluded
-
[103]
Afterward, we segment the predicted view image with ground truth per-object visible mask
If an object’s visible mask is less than 70% of its full mask, the object is heavily occluded. Afterward, we segment the predicted view image with ground truth per-object visible mask. We calculate the spe- cific region’s PSNR, SSIM, and LPIPS metrics as shown in Tab. S.6. It ...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.