REVIEW 3 major objections 5 minor 1 cited by
SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read SplatFlow claims that a single text-conditioned model can generate and edit 3D Gaussian Splatting scenes directly from prompts.
desk verdict A genuinely unified 3DGS generation/editing system with a clever joint latent and strong pose results; evaluation is thin in places, but the core idea holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the multi-view rectified flow model, trained by conditional flow matching on tuples of image latent, depth latent, and Plücker ray coordinates for eight views, concatenated along the channel axis. Plücker ray coordinates encode each pixel's camera ray as a direction vector and a moment vector, so camera poses are generated latents rather than fixed inputs. During sampling, the model stops updating rays at $t_{\mathrm{stop}}=150$, predicts the destination at $t=0$, recovers camera poses from the predicted rays, and re-projects them onto the valid ray manifold; before that stop it inserts the base text-to-image model's velocity field every third step to improve generalization. The Gaussian Splatting Decoder, initialized from the same base decoder with cross-view attention, then maps the latents to pixel-aligned 3D Gaussians, using depth latents and a vision-aided adversarial loss as additional design choices.
What would settle it
Run the same prompts and random seeds with and without that scheduled injection, then compare not only FID and CLIP but also a geometric consistency measure such as the epipolar error between feature matches in generated views computed from the predicted poses. If the injection leaves geometric consistency unchanged or worse while improving FID and CLIP, the claimed mechanism is not doing the work the paper assigns to it.
Extended reading notes
Core claim
The paper's central claim is that a single text-conditioned rectified flow model can generate the complete input needed for 3D Gaussian Splatting: images, depths, and camera poses, all at once from text. It treats these as one joint latent distribution over image, depth, and Plücker-ray-coordinate latents, and it shows that conditioning on any known subset recovers the rest through flow-based inpainting. The same model therefore covers generation, editing, pose estimation, and novel view synthesis, with quantitative results on MVImgNet and DL3DV-7K showing lower FID and higher CLIP scores than the closest baseline.
Load-bearing premise
The load-bearing assumption is that injecting the base text-to-image model's velocity field every third sampling step before $t_{\mathrm{stop}}=150$ improves image quality without breaking the joint consistency of images, depths, and camera poses.
Editorial extensions
If this is right
- Text-to-3DGS synthesis no longer requires per-scene optimization in the base pipeline; a scene splat is obtained after one flow sampling run and one feed-forward decode.
- Object replacement becomes a masked inversion of the rectified flow followed by inpainting, so no separate cross-view editing module is needed.
- Camera pose estimation and novel view synthesis are the same operation from opposite sides: hold the known latents fixed and inpaint the unknown ray or view latents.
- Since the latent space is shared with a large text-to-image base model, prompts from the 2D editing literature can guide 3D edits at sampling time.
- The joint-distribution design naturally extends to any task that can be stated as completing a partial multi-view latent.
Reading between the lines
- A direct test of cross-view geometric consistency (e.g., epipolar error from predicted poses) would tell whether the scheduled injection of the base model's velocity field improves images without paying for it in 3D coherence; the paper only reports FID and CLIP.
- Because the GSDecoder deliberately smooths inconsistencies between generated views, the framework could afford to train the flow model for per-view sharpness and leave consistency repair to the decoder—a trade-off the paper observes but does not quantify.
- The same inpainting machinery could solve relocalization or scene completion with an unknown number of views, since any partial set of image, depth, and ray latents is a valid conditioning signal.
- The depth labels used during training are not sacred: pose estimation improved when depth latents were excluded, so a learned geometric prior might replace external depth supervision entirely.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SplatFlow is a framework for text-conditioned 3D Gaussian Splatting (3DGS) synthesis and editing. It consists of a multi-view rectified flow (RF) model that jointly generates image latents, depth latents, and Plücker ray latents for K views, and a Gaussian Splatting Decoder (GSDecoder) that converts these latents into pixel-aligned 3DGS in a feed-forward manner. The same RF model is also used for training-free 3D object editing, camera pose estimation, and novel view synthesis via inversion and inpainting. The method is evaluated on MVImgNet and DL3DV-7K, comparing generation to Director3D and editing to DGE and MVInpainter.
Significance. If the claims hold, SplatFlow is a valuable step toward unifying 3D scene generation and editing in a single model, avoiding per-scene optimization at generation time (except for the optional SDS++ refinement) and eliminating task-specific modules for editing and inpainting. The design choices—joint modeling of images, depths, and poses; a shared latent space with Stable Diffusion 3; and a two-stage decoder—are well motivated, and the ablations on depth-latent integration, adversarial loss, early stopping, and SD3 guidance are useful. The main weaknesses are in the depth and breadth of validation: the generation comparison uses only one baseline and no uncertainty estimates, the novel-view-synthesis results lack any baseline and are of low absolute quality, and no metric directly measures multi-view geometric consistency, which is the core assumption of the entire pipeline. These gaps currently prevent the paper from fully substantiating its central claims.
major comments (3)
- [Section 4.3, Algorithm 1 lines 4-6, and Appendix D.2] The SD3 guidance step replaces image-latent velocity components with v_phi(Y_ti[:n], ti), a single-image velocity field that is not conditioned on depth latents, ray latents, or the other views. The paper validates this intervention only with FID-10K and CLIPScore (Table 7), which are per-view metrics and cannot detect cross-view geometric inconsistency, double-mapped regions, or pose drift. Appendix D.2 states that the multi-view generated images contain color and shape inconsistencies that the GSDecoder smooths out by blurring, so the risk that SD3 guidance pushes each view toward an independent 2D prior and away from the joint multi-view manifold is not merely hypothetical. I ask for a direct evaluation of multi-view consistency with and without SD3 guidance—for example, epipolar or feature-metric error between generated views, or consistency between generated images and generated camera poses—as this joint consistency is load-bearing for the GSDecoder, editing, and inpainting applications.
- [Section 5.4, Table 4] The novel view synthesis results report PSNR between 14.73 and 18.82 and LPIPS between 0.483 and 0.648. These values are far below what is typically considered usable for image-based novel view synthesis, and there is no baseline method evaluated on the same scenes. Without a nearest-neighbor baseline (e.g., returning the closest input view) or a standard sparse-view NVS method, the claim that SplatFlow 'supports' novel view synthesis cannot be interpreted quantitatively; the current table does not establish that the multi-view RF model adds value over simple view selection for this task.
- [Sections 5.2 and 5.3] The core generation and editing claims are supported by comparisons to only one baseline each (Director3D for generation, DGE for editing), with no error bars, confidence intervals, or statistical tests. FID and CLIP are global metrics that do not measure geometric consistency, and the editing benchmark is built from 100 scenes with GPT-4-generated target captions. To make the comparisons convincing, I recommend reporting per-scene standard deviations or confidence intervals, adding at least one more recent scene-level baseline for generation, and considering a quantitative comparison to MVInpainter (or another multi-view inpainting method) on the editing task.
minor comments (5)
- [Abstract and Section 5.1] The paper describes SplatFlow as avoiding per-scene optimization, but the default generation pipeline includes an optional SDS++ refinement step that takes about 5 minutes per scene. Please clarify whether the 'direct generation' claim applies only to the base model and separately describe the role and cost of SDS++.
- [Section 5.2] The evaluation uses 'the rendered image of the generated 3DGS' but does not specify how the 3DGS is rendered for FID/CLIP computation (e.g., which views, what camera trajectory). This detail is needed for reproducibility.
- [Table 3 and Appendix D.1] The ablation shows that SplatFlow without depth latents performs consistently better than with depth latents for camera pose estimation, yet the main method always uses depth latents in the GSDecoder. Please discuss whether this discrepancy affects the GSDecoder's reliance on generated depth latents and whether the pose estimation task is best formulated with or without depth.
- [Appendix B.3] The editing benchmark relies on GPT-4 to generate target captions, which may introduce a bias in the CLIP-based metrics because CLIP and GPT-4 are trained on overlapping text distributions. Please mention this as a limitation and consider a human evaluation or a hand-curated prompt set.
- [Appendix D] The heading 'Anaylsis and Discussion' contains a typo ('Anaylsis' should be 'Analysis').
Circularity Check
No significant circularity: SplatFlow is an empirical system whose claims are evaluated on held-out benchmarks, and no load-bearing step reduces to its own inputs.
full rationale
SplatFlow's central chain is not circular in the sense that counts here. The multi-view RF model is trained with conditional flow matching on held-out splits of MVImgNet and DL3DV-7K, and its outputs (image/depth/ray latents) are decoded by a separately trained GSDecoder; the text-to-3DGS evaluation is against GT rendered images via FID/CLIP, camera pose estimation is compared against dataset GT poses (Table 3), and novel view synthesis is compared against GT views with PSNR/SSIM/LPIPS (Table 4). No fitted parameter is renamed as a prediction: ablations in Section C.2 (t_stop and SD3 guidance) are empirical design choices, not quantities defined in terms of the evaluated outputs. The paper's own D.2 analysis concedes that multi-view generated images contain color/shape inconsistencies that the GSDecoder smooths, but this is a limitation of consistency evidence, not circular reasoning. Self-citations to VideoRF-Splat [26] and SteerX [80] appear only in the extended related-work discussion and carry none of the paper's load. No uniqueness theorem, ansatz, or definition is imported from the authors' prior work to force a conclusion. The honest finding is therefore no significant circularity.
Assumptions & free parameters
free parameters (5)
- Classifier-free guidance scales for image/depth/ray latents =
7, 5, 1
- Stable Diffusion 3 guidance scale =
3
- Early stopping timestep t_stop =
150 out of 200
- GSDecoder loss weights and vision-aided GAN activation =
w1=1, w2=0.05, GAN after 200K iterations
- Number of sampling steps N =
200
assumptions (5)
- standard math Flow matching losses L_CFM and L_FM have identical gradients, so training with conditional flow matching learns the marginal vector field.
- standard math Pinhole camera model with Plucker ray representation and DLT pose recovery is valid for the datasets.
- domain assumption The frozen SD3 encoder provides a latent space shared by SplatFlow and SD3, and DepthAnythingV2 depth maps are accurate enough to serve as structural supervision.
- domain assumption Eight views with a shared intrinsic matrix are a sufficient representation for real-world scenes in MVImgNet and DL3DV.
- domain assumption RePaint-style resampling of unknown latents converges to a plausible joint multi-view distribution.
Cite this review
Pith. "Pith review of SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis." pith.science (2026). https://pith.science/paper/W2EARPIJ
@misc{pith2026241116443,
author = {Pith},
title = {Pith review of: SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis},
year = {2026},
howpublished = {\url{https://pith.science/paper/W2EARPIJ}},
note = {Machine review of arXiv:2411.16443}
}
read the original abstract
Text-based generation and editing of 3D scenes hold significant potential for streamlining content creation through intuitive user interactions. While recent advances leverage 3D Gaussian Splatting (3DGS) for high-fidelity and real-time rendering, existing methods are often specialized and task-focused, lacking a unified framework for both generation and editing. In this paper, we introduce SplatFlow, a comprehensive framework that addresses this gap by enabling direct 3DGS generation and editing. SplatFlow comprises two main components: a multi-view rectified flow (RF) model and a Gaussian Splatting Decoder (GSDecoder). The multi-view RF model operates in latent space, generating multi-view images, depths, and camera poses simultaneously, conditioned on text prompts, thus addressing challenges like diverse scene scales and complex camera trajectories in real-world settings. Then, the GSDecoder efficiently translates these latent outputs into 3DGS representations through a feed-forward 3DGS method. Leveraging training-free inversion and inpainting techniques, SplatFlow enables seamless 3DGS editing and supports a broad range of 3D tasks-including object editing, novel view synthesis, and camera pose estimation-within a unified framework without requiring additional complex pipelines. We validate SplatFlow's capabilities on the MVImgNet and DL3DV-7K datasets, demonstrating its versatility and effectiveness in various 3D generation, editing, and inpainting-based tasks.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
Flow Straight and Fast in Hilbert Space: Functional Rectified Flow
Functional rectified flow is defined and proved to preserve marginals in separable Hilbert spaces, with functional flow matching and probability-flow ODEs as special cases.
Reference graph
Works this paper leans on
-
[1]
Direct linear transformation from comparator coordinates into object space coordinates in close range photogrammetry
Yousset I Abdel-Aziz. Direct linear transformation from comparator coordinates into object space coordinates in close range photogrammetry. Proc. Amer. Soc. Photogram- metry, 1971, pages 1–19, 1971. 16
1971
-
[2]
Building nor- malizing flows with stochastic interpolants
Michael S Albergo and Eric Vanden-Eijnden. Building nor- malizing flows with stochastic interpolants. arXiv preprint arXiv:2209.15571, 2022. 2, 3
arXiv 2022
-
[3]
Denoising diffusion via image-based rendering
Titas Anciukeviˇcius, Fabian Manhardt, Federico Tombari, and Paul Henderson. Denoising diffusion via image-based rendering. In The Twelfth International Conference on Learning Representations, 2024. 21
2024
-
[4]
Sine: Semantic-driven image-based nerf editing with prior-guided editing field
Chong Bao, Yinda Zhang, Bangbang Yang, Tianxing Fan, Zesong Yang, Hujun Bao, Guofeng Zhang, and Zhaopeng Cui. Sine: Semantic-driven image-based nerf editing with prior-guided editing field. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20919–20929, 2023. 1, 3
2023
-
[5]
Mvinpainter: Learning multi-view consis- tent inpainting to bridge 2d and 3d editing
Chenjie Cao, Chaohui Yu, Yanwei Fu, Fan Wang, and Xi- angyang Xue. Mvinpainter: Learning multi-view consis- tent inpainting to bridge 2d and 3d editing. arXiv preprint arXiv:2408.08000, 2024. 7, 18
arXiv 2024
-
[6]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021. 15
2021
-
[7]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19457–19467, 2024. 2, 4
2024
-
[8]
Lara: Efficient large-baseline radiance fields
Anpei Chen, Haofei Xu, Stefano Esposito, Siyu Tang, and Andreas Geiger. Lara: Efficient large-baseline radiance fields. In European Conference on Computer Vision, pages 338–355. Springer, 2025. 3
2025
Show all 144 references
-
[9]
Pixart- α: Fast training of diffusion trans- former for photorealistic text-to-image synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart- α: Fast training of diffusion trans- former for photorealistic text-to-image synthesis. arXiv preprint arXiv:2310.00426, 2023. 2, 3
-
[10]
Dge: Direct gaussian 3d editing by consistent multi-view editing
Minghao Chen, Iro Laina, and Andrea Vedaldi. Dge: Direct gaussian 3d editing by consistent multi-view editing. arXiv preprint arXiv:2404.18929, 2024. 2, 7, 18
2024 arXiv
-
[11]
Fan- tasia3d: Disentangling geometry and appearance for high- quality text-to-3d content creation
Rui Chen, Yongwei Chen, Ningxin Jiao, and Kui Jia. Fan- tasia3d: Disentangling geometry and appearance for high- quality text-to-3d content creation. In Proceedings of the IEEE/CVF international conference on computer vision , pages 22246–22256, 2023. 19
2023
-
[12]
Gaussianeditor: Swift and controllable 3d editing with gaussian splatting
Yiwen Chen, Zilong Chen, Chi Zhang, Feng Wang, Xiaofeng Yang, Yikai Wang, Zhongang Cai, Lei Yang, Huaping Liu, and Guosheng Lin. Gaussianeditor: Swift and controllable 3d editing with gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...
2024
-
[13]
Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. In European Conference on Computer Vision, pages 370–386. Springer, 2024. 2, 4
2024
-
[14]
Text-to-3d using gaussian splatting
Zilong Chen, Feng Wang, Yikai Wang, and Huaping Liu. Text-to-3d using gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21401–21412, 2024. 2, 3
2024
-
[15]
Diffusion posterior sam- pling for general noisy inverse problems
Hyungjin Chung, Jeongsol Kim, Michael T Mccann, Marc L Klasky, and Jong Chul Ye. Diffusion posterior sam- pling for general noisy inverse problems. arXiv preprint arXiv:2209.14687, 2022. 2, 5
2022 arXiv
-
[16]
Luciddreamer: Domain-free generation of 3d gaussian splatting scenes
Jaeyoung Chung, Suyoung Lee, Hyeongjin Nam, Jaerin Lee, and Kyoung Mu Lee. Luciddreamer: Domain-free generation of 3d gaussian splatting scenes. arXiv preprint arXiv:2311.13384, 2023. 3, 6, 7
2023 arXiv
-
[17]
Diffedit: Diffusion-based seman- tic image editing with mask guidance
Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. Diffedit: Diffusion-based seman- tic image editing with mask guidance. arXiv preprint arXiv:2210.11427, 2022. 2, 5
2022 arXiv
-
[18]
A survey on diffusion mod- els for inverse problems
Giannis Daras, Hyungjin Chung, Chieh-Hsin Lai, Yuki Mit- sufuji, Jong Chul Ye, Peyman Milanfar, Alexandros G Di- makis, and Mauricio Delbracio. A survey on diffusion mod- els for inverse problems. arXiv preprint arXiv:2410.00083,
-
[19]
Objaverse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...
2023
-
[20]
Objaverse-xl: A universe of 10m+ 3d objects
Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram V oleti, Samir Yitzhak Gadre, et al. Objaverse-xl: A universe of 10m+ 3d objects. Advances in Neural Informa- tion Processing Systems, 36, 2024. 2, 3
2024
-
[21]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural informa- tion processing systems, 34:8780–8794, 2021. 2
2021
-
[22]
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883, 2021. 4
2021
-
[23]
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first International Conference on Machin...
-
[24]
Towards practical plug-and-play diffusion models
Hyojun Go, Yunsung Lee, Jin-Young Kim, Seunghyun Lee, Myeongho Jeong, Hyun Seung Lee, and Seungtaek Choi. Towards practical plug-and-play diffusion models. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1962–1971, 2023. 2
1962
-
[25]
Addressing nega- tive transfer in diffusion models
Hyojun Go, Yunsung Lee, Seunghyun Lee, Shinhyeok Oh, Hyeongdon Moon, and Seungtaek Choi. Addressing nega- tive transfer in diffusion models. Advances in Neural Infor- mation Processing Systems, 36, 2024. 2 9
2024
-
[26]
Videorfsplat: Direct scene-level text-to-3d gaussian splatting generation with flexible pose and multi-view joint modeling, 2025
Hyojun Go, Byeongjun Park, Hyelin Nam, Byung-Hoon Kim, Hyungjin Chung, and Changick Kim. Videorfsplat: Direct scene-level text-to-3d gaussian splatting generation with flexible pose and multi-view joint modeling, 2025. 21
2025
-
[27]
Diffusion models as plug-and-play priors
Alexandros Graikos, Nikolay Malkin, Nebojsa Jojic, and Dimitris Samaras. Diffusion models as plug-and-play priors. Advances in Neural Information Processing Systems , 35: 14715–14728, 2022. 2
2022
-
[28]
T3bench: Benchmarking current progress in text-to-3d generation,
Yuze He, Yushi Bai, Matthieu Lin, Wang Zhao, Yubin Hu, Jenny Sheng, Ran Yi, Juanzi Li, and Yong-Jin Liu. T3bench: Benchmarking current progress in text-to-3d generation,
-
[29]
Prompt-to-prompt im- age editing with cross attention control
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt im- age editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022. 2
2022 arXiv
-
[30]
Clipscore: A reference-free evaluation met- ric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation met- ric for image captioning. arXiv preprint arXiv:2104.08718,
-
[31]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bern- hard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. Advances in neural information processing systems, 30, 2017. 6, 18
2017
-
[32]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 16
2022 arXiv
-
[33]
Denoising diffu- sion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 2, 3
2020
-
[34]
Imagen video: High definition video generation with diffusion mod- els
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion mod- els. arXiv preprint arXiv:2210.02303, 2022. 2
-
[35]
Text2room: Extracting textured 3d meshes from 2d text-to-image models
Lukas H¨ollein, Ang Cao, Andrew Owens, Justin Johnson, and Matthias Nießner. Text2room: Extracting textured 3d meshes from 2d text-to-image models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7909–7920, 2023. 3
2023
-
[36]
Instruct 3d-to-3d: Text in- struction guided 3d-to-3d conversion
Hiromichi Kamata, Yuiko Sakuma, Akio Hayakawa, Masato Ishii, and Takuya Narihira. Instruct 3d-to-3d: Text in- struction guided 3d-to-3d conversion. arXiv preprint arXiv:2303.15780, 2023. 1, 3, 7
2023 arXiv
-
[37]
Free-editor: zero-shot text-driven 3d scene editing
Nazmul Karim, Hasan Iqbal, Umar Khalid, Chen Chen, and Jing Hua. Free-editor: zero-shot text-driven 3d scene editing. In European Conference on Computer Vision, pages 436–
-
[38]
Training generative ad- versarial networks with limited data
Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Training generative ad- versarial networks with limited data. Advances in neural information processing systems, 33:12104–12114, 2020. 15
2020
-
[39]
Repurpos- ing diffusion-based image generators for monocular depth estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- zger, Rodrigo Caye Daudt, and Konrad Schindler. Repurpos- ing diffusion-based image generators for monocular depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9492–9502,
-
[40]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1,
-
[41]
Latenteditor: Text driven local editing of 3d scenes
Umar Khalid, Hasan Iqbal, Nazmul Karim, Jing Hua, and Chen Chen. Latenteditor: Text driven local editing of 3d scenes. arXiv preprint arXiv:2312.09313, 2023. 1, 3
2023 arXiv
-
[42]
3dego: 3d editing on the go! In European Conference on Computer Vision, pages 73–89
Umar Khalid, Hasan Iqbal, Azib Farooq, Jing Hua, and Chen Chen. 3dego: 3d editing on the go! In European Conference on Computer Vision, pages 73–89. Springer, 2024. 2, 3
2024
-
[43]
Dif- fusionclip: Text-guided diffusion models for robust image manipulation
Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. Dif- fusionclip: Text-guided diffusion models for robust image manipulation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 2426– 2435, 2022. 2
2022
-
[44]
Adam: A method for stochastic opti- mization
Diederik P Kingma. Adam: A method for stochastic opti- mization. arXiv preprint arXiv:1412.6980, 2014. 16
2014 arXiv
-
[45]
Ensembling off-the-shelf models for gan training
Nupur Kumari, Richard Zhang, Eli Shechtman, and Jun- Yan Zhu. Ensembling off-the-shelf models for gan training. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10651–10662, 2022. 4, 5, 15
2022
-
[46]
Multi-architecture multi-expert diffusion models
Yunsung Lee, JinYoung Kim, Hyojun Go, Myeongho Jeong, Shinhyeok Oh, and Seungtaek Choi. Multi-architecture multi-expert diffusion models. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 13427–13436,
-
[47]
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 6
2024 arXiv
-
[48]
Dreamscene: 3d gaussian-based text-to-3d scene generation via formation pattern sampling
Haoran Li, Haolin Shi, Wenli Zhang, Wenjun Wu, Yong Liao, Lin Wang, Lik-hang Lee, and Pengyuan Zhou. Dreamscene: 3d gaussian-based text-to-3d scene generation via formation pattern sampling. arXiv preprint arXiv:2404.03575, 2024. 6, 7
2024 arXiv
-
[49]
Diffusion-lm improves con- trollable text generation
Xiang Li, John Thickstun, Ishaan Gulrajani, Percy S Liang, and Tatsunori B Hashimoto. Diffusion-lm improves con- trollable text generation. Advances in Neural Information Processing Systems, 35:4328–4343, 2022. 2
2022
-
[50]
Dual3d: Efficient and consistent text-to-3d generation with dual-mode multi-view latent diffusion
Xinyang Li, Zhangyu Lai, Linning Xu, Jianfei Guo, Li- ujuan Cao, Shengchuan Zhang, Bo Dai, and Rongrong Ji. Dual3d: Efficient and consistent text-to-3d generation with dual-mode multi-view latent diffusion. arXiv preprint arXiv:2405.09874, 2024. 16
2024 arXiv
-
[51]
Direc- tor3d: Real-world camera trajectory and 3d scene generation from text
Xinyang Li, Zhangyu Lai, Linning Xu, Yansong Qu, Liujuan Cao, Shengchuan Zhang, Bo Dai, and Rongrong Ji. Direc- tor3d: Real-world camera trajectory and 3d scene generation from text. Advances in neural information processing sys- tems, 2024. 3, 5, 6, 7, 19, 21
2024
-
[52]
Connecting consistency distillation to score distillation for text-to-3d generation
Zongrui Li, Minghui Hu, Qian Zheng, and Xudong Jiang. Connecting consistency distillation to score distillation for text-to-3d generation. In European Conference on Computer Vision, pages 274–291. Springer, 2024. 2
2024
-
[53]
Wonderland: Nav- 10 igating 3d scenes from a single image
Hanwen Liang, Junli Cao, Vidit Goel, Guocheng Qian, Sergei Korolev, Demetri Terzopoulos, Konstantinos N Pla- taniotis, Sergey Tulyakov, and Jian Ren. Wonderland: Nav- 10 igating 3d scenes from a single image. arXiv preprint arXiv:2412.12091, 2024. 21
2024 arXiv
-
[54]
Relpose++: Recovering 6d poses from sparse-view observations
Amy Lin, Jason Y Zhang, Deva Ramanan, and Shubham Tulsiani. Relpose++: Recovering 6d poses from sparse-view observations. arXiv preprint arXiv:2305.04926, 2023. 7, 8
2023 arXiv
-
[55]
Diffsplat: Repurposing image diffusion models for scalable gaussian splat generation.arXiv preprint arXiv:2501.16764, 2025
Chenguo Lin, Panwang Pan, Bangbang Yang, Zeming Li, and Yadong Mu. Diffsplat: Repurposing image diffusion models for scalable gaussian splat generation.arXiv preprint arXiv:2501.16764, 2025. 21
2025 arXiv
-
[56]
Magic3d: High-resolution text-to-3d content creation
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...
2023
-
[57]
Align your gaussians: Text-to-4d with dynamic 3d gaussians and composed diffusion models
Huan Ling, Seung Wook Kim, Antonio Torralba, Sanja Fi- dler, and Karsten Kreis. Align your gaussians: Text-to-4d with dynamic 3d gaussians and composed diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8576–8588, 2024. 3
2024
-
[58]
Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pa...
2024
-
[59]
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maxim- ilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022. 2, 3
2022 arXiv
-
[60]
Flow straight and fast: Learning to generate and transfer data with rectified flow
Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022. 2, 3
2022 arXiv
-
[61]
Humangaussian: Text-driven 3d human generation with gaussian splatting
Xian Liu, Xiaohang Zhan, Jiaxiang Tang, Ying Shan, Gang Zeng, Dahua Lin, Xihui Liu, and Ziwei Liu. Humangaussian: Text-driven 3d human generation with gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6646–6657, 2024. 2
2024
-
[62]
Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age
Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age. arXiv preprint arXiv:2309.03453, 2023. 3
2023 arXiv
-
[63]
Wonder3d: Single image to 3d using cross-domain diffusion
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al. Wonder3d: Single image to 3d using cross-domain diffusion. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Patt...
2024
-
[64]
Decoupled weight decay regularization
I Loshchilov. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 17
2017 arXiv
-
[65]
Repaint: Inpainting using denoising diffusion probabilistic models
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11461–11471, 2022. 2, ...
2022
-
[66]
Sdedit: Guided image synthesis and editing with stochastic differential equations
Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073, 2021. 2, 4, 5, 19
2021 arXiv
-
[67]
Latent-nerf for shape-guided generation of 3d shapes and textures
Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or. Latent-nerf for shape-guided generation of 3d shapes and textures. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12663–12673, 2023. 19
2023
-
[68]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM, 65(1):99–106, 2021. 1, 2
2021
-
[69]
No-reference image quality assessment in the spatial domain
Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. No-reference image quality assessment in the spatial domain. IEEE Transactions on image processing, 21(12): 4695–4708, 2012. 19
2012
-
[70]
completely blind
Anish Mittal, Rajiv Soundararajan, and Alan C Bovik. Mak- ing a “completely blind” image quality analyzer. IEEE Signal processing letters, 20(3):209–212, 2012. 19
2012
-
[71]
Null-text inversion for editing real im- ages using guided diffusion models
Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Null-text inversion for editing real im- ages using guided diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6038–6047, 2023. 2, 5
2023
-
[72]
Instant neural graphics primitives with a mul- tiresolution hash encoding
Thomas M¨uller, Alex Evans, Christoph Schied, and Alexan- der Keller. Instant neural graphics primitives with a mul- tiresolution hash encoding. ACM transactions on graphics (TOG), 41(4):1–15, 2022. 1
2022
-
[73]
Glide: Towards photorealistic image genera- tion and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image genera- tion and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021. 2
2021 arXiv
-
[74]
Point-e: A system for gener- ating 3d point clouds from complex prompts
Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for gener- ating 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 2
2022 arXiv
-
[75]
Gsedit: Efficient text-guided edit- ing of 3d objects via gaussian splatting
Francesco Palandra, Andrea Sanchietti, Daniele Baieri, and Emanuele Rodol `a. Gsedit: Efficient text-guided edit- ing of 3d objects via gaussian splatting. arXiv preprint arXiv:2403.05154, 2024. 2, 3
2024 arXiv
-
[76]
Point-dynrf: Point- based dynamic radiance fields from a monocular video
Byeongjun Park and Changick Kim. Point-dynrf: Point- based dynamic radiance fields from a monocular video. In Proceedings of the IEEE/CVF Winter Conference on Appli- cations of Computer Vision, pages 3171–3181, 2024. 3
2024
-
[77]
Denoising task routing for diffusion models
Byeongjun Park, Sangmin Woo, Hyojun Go, Jin-Young Kim, and Changick Kim. Denoising task routing for diffusion models. arXiv preprint arXiv:2310.07138, 2023. 2
2023 arXiv
-
[78]
Bridging implicit and explicit geometric transformation for single- image view synthesis
Byeongjun Park, Hyojun Go, and Changick Kim. Bridging implicit and explicit geometric transformation for single- image view synthesis. IEEE Transactions on Pattern Analy- sis and Machine Intelligence, 2024. 2
2024
-
[79]
Switch diffusion trans- former: Synergizing denoising tasks with sparse mixture-of- experts
Byeongjun Park, Hyojun Go, Jin-Young Kim, Sangmin Woo, Seokil Ham, and Changick Kim. Switch diffusion trans- former: Synergizing denoising tasks with sparse mixture-of- experts. arXiv preprint arXiv:2403.09176, 2024. 2
2024 arXiv
-
[80]
Steerx: Creating 11 any camera-free 3d and 4d scenes with geometric steering
Byeongjun Park, Hyojun Go, Hyelin Nam, Byung-Hoon Kim, Hyungjin Chung, and Changick Kim. Steerx: Creating 11 any camera-free 3d and 4d scenes with geometric steering. arXiv preprint arXiv:2503.12024, 2025. 21
2025 arXiv
-
[81]
Ed-nerf: Efficient text-guided editing of 3d scene using latent space nerf
Jangho Park, Gihyun Kwon, and Jong Chul Ye. Ed-nerf: Efficient text-guided editing of 3d scene using latent space nerf. arXiv preprint arXiv:2310.02712, 2023. 1, 3
2023 arXiv
-
[82]
On aliased resizing and surprising subtleties in gan evaluation
Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu. On aliased resizing and surprising subtleties in gan evaluation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11410–11420, 2022. 17
2022
-
[83]
Analytisch-geometrische Entwicklungen
Julius Pl¨ucker. Analytisch-geometrische Entwicklungen. GD Baedeker, 1828. 5
-
[84]
Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 2, 3
2023 arXiv
-
[85]
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 2, 19
2022 arXiv
-
[86]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...
2021
-
[87]
Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In Proceedings of the IEEE/CVF international conference on computer vi...
2021
-
[88]
Texture: Text-guided texturing of 3d shapes
Elad Richardson, Gal Metzer, Yuval Alaluf, Raja Giryes, and Daniel Cohen-Or. Texture: Text-guided texturing of 3d shapes. In ACM SIGGRAPH 2023 conference proceedings, pages 1–11, 2023. 3
2023
-
[89]
Geometry-free view synthesis: Transformers and no 3d pri- ors
Robin Rombach, Patrick Esser, and Bj ¨orn Ommer. Geometry-free view synthesis: Transformers and no 3d pri- ors. In Proceedings of the IEEE/CVF International Confer- ence on Computer Vision, pages 14356–14366, 2021. 2
2021
-
[90]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 4, 15
2022
-
[91]
A theoretical justification for image inpainting using denoising diffusion probabilistic models
Litu Rout, Advait Parulekar, Constantine Caramanis, and Sanjay Shakkottai. A theoretical justification for image inpainting using denoising diffusion probabilistic models. arXiv preprint arXiv:2302.01217, 2023. 2, 5
2023 arXiv
-
[92]
Solving linear inverse problems provably via posterior sampling with latent diffusion models
Litu Rout, Negin Raoof, Giannis Daras, Constantine Cara- manis, Alex Dimakis, and Sanjay Shakkottai. Solving linear inverse problems provably via posterior sampling with latent diffusion models. Advances in Neural Information Process- ing Systems, 36, 2023. 2, 5
2023
-
[93]
Adversarial diffusion distillation
Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin Rombach. Adversarial diffusion distillation. In European Conference on Computer Vision, pages 87–103. Springer,
-
[94]
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural In- f...
2022
-
[95]
Generative gaussian splatting: Generating 3d scenes with video diffusion priors
Katja Schwarz, Norman Mueller, and Peter Kontschieder. Generative gaussian splatting: Generating 3d scenes with video diffusion priors. arXiv preprint arXiv:2503.13272,
-
[96]
Mvdream: Multi-view diffusion for 3d generation
Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023. 3
2023 arXiv
-
[97]
Realmdreamer: Text-driven 3d scene genera- tion with inpainting and depth diffusion
Jaidev Shriram, Alex Trevithick, Lingjie Liu, and Ravi Ra- mamoorthi. Realmdreamer: Text-driven 3d scene genera- tion with inpainting and depth diffusion. arXiv preprint arXiv:2404.07199, 2024. 3
2024 arXiv
-
[98]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International confer- ence on machine learning, pages 2256–2265. PMLR, 2015. 2, 3
2015
-
[99]
Pseudoinverse-guided diffusion models for inverse problems
Jiaming Song, Arash Vahdat, Morteza Mardani, and Jan Kautz. Pseudoinverse-guided diffusion models for inverse problems. In International Conference on Learning Repre- sentations, 2023. 2, 5
2023
-
[100]
Loss-guided diffusion models for plug-and-play controllable generation
Jiaming Song, Qinsheng Zhang, Hongxu Yin, Morteza Mar- dani, Ming-Yu Liu, Jan Kautz, Yongxin Chen, and Arash Vahdat. Loss-guided diffusion models for plug-and-play controllable generation. In International Conference on Machine Learning, pages 32483–32498. PMLR, 2023. 2
2023
-
[101]
Score-based generative modeling through stochastic differential equa- tions
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions. arXiv preprint arXiv:2011.13456, 2020. 2, 3
2011 arXiv
-
[102]
Viewset diffusion:(0-) image-conditioned 3d gener- ative models from 2d data
Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Viewset diffusion:(0-) image-conditioned 3d gener- ative models from 2d data. In Proceedings of the IEEE/CVF international conference on computer vision, pages 8863– 8873, 2023. 21
2023
-
[103]
Splatter image: Ultra-fast single-view 3d recon- struction
Stanislaw Szymanowicz, Chrisitian Rupprecht, and Andrea Vedaldi. Splatter image: Ultra-fast single-view 3d recon- struction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 10208– 10217, 2024. 4
2024
-
[104]
Bolt3d: Generating 3d scenes in seconds
Stanislaw Szymanowicz, Jason Y Zhang, Pratul Srinivasan, Ruiqi Gao, Arthur Brussee, Aleksander Holynski, Ricardo Martin-Brualla, Jonathan T Barron, and Philipp Henzler. Bolt3d: Generating 3d scenes in seconds. arXiv preprint arXiv:2503.14445, 2025. 21
2025
-
[105]
Lgm: Large multi-view gaussian model for high-resolution 3d content creation
Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. In European Conference on Computer Vision, pages 1–18. Springer, 2024. 2, 3, 4
2024
-
[106]
Dreamgaussian: Generative gaussian splatting for 12 efficient 3d content creation
Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. Dreamgaussian: Generative gaussian splatting for 12 efficient 3d content creation. In The Twelfth International Conference on Learning Representations, 2024. 2, 3
2024
-
[107]
Lion: Latent point diffu- sion models for 3d shape generation
Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, Karsten Kreis, et al. Lion: Latent point diffu- sion models for 3d shape generation. Advances in Neural Information Processing Systems, 35:10021–10039, 2022. 2
2022
-
[108]
Cg3d: Compositional generation for text-to-3d via gaussian splatting
Alexander Vilesov, Pradyumna Chari, and Achuta Kadambi. Cg3d: Compositional generation for text-to-3d via gaussian splatting. arXiv preprint arXiv:2311.17907, 2023. 2
2023 arXiv
-
[109]
Sv3d: Novel multi-view synthesis and 3d generation from a single image using la- tent video diffusion
Vikram V oleti, Chun-Han Yao, Mark Boss, Adam Letts, David Pankratz, Dmitry Tochilkin, Christian Laforte, Robin Rombach, and Varun Jampani. Sv3d: Novel multi-view synthesis and 3d generation from a single image using la- tent video diffusion. In European Conference on Computer...
2025
-
[110]
Mesh-guided neural implicit field editing
Can Wang, Mingming He, Menglei Chai, Dongdong Chen, and Jing Liao. Mesh-guided neural implicit field editing. arXiv preprint arXiv:2312.02157, 2023. 3
2023 arXiv
-
[111]
Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation
Haochen Wang, Xiaodan Du, Jiahao Li, Raymond A Yeh, and Greg Shakhnarovich. Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12619–12629, 2023. 19
2023
-
[112]
Gaussianeditor: Editing 3d gaussians delicately with text instructions
Junjie Wang, Jiemin Fang, Xiaopeng Zhang, Lingxi Xie, and Qi Tian. Gaussianeditor: Editing 3d gaussians delicately with text instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20902–20911, 2024. 2, 3
2024
-
[113]
View-consistent 3d editing with gaus- sian splatting
Yuxuan Wang, Xuanyu Yi, Zike Wu, Na Zhao, Long Chen, and Hanwang Zhang. View-consistent 3d editing with gaus- sian splatting. In European Conference on Computer Vision, pages 404–420. Springer, 2024. 2, 3, 5
2024
-
[114]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 18
2004
-
[115]
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distilla- tion
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distilla- tion. Advances in Neural Information Processing Systems, 36, 2024. 19
2024
-
[116]
Harmonyview: Harmonizing consis- tency and diversity in one-image-to-3d
Sangmin Woo, Byeongjun Park, Hyojun Go, Jin-Young Kim, and Changick Kim. Harmonyview: Harmonizing consis- tency and diversity in one-image-to-3d. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10574–10584, 2024. 3
2024
-
[117]
Gaussc- trl: multi-view consistent text-driven 3d gaussian splatting editing
Jing Wu, Jia-Wang Bian, Xinghui Li, Guangrun Wang, Ian Reid, Philip Torr, and Victor Adrian Prisacariu. Gaussc- trl: multi-view consistent text-driven 3d gaussian splatting editing. arXiv preprint arXiv:2403.08733, 2024. 2, 3, 5
2024 arXiv
-
[118]
Lo- calized gaussian splatting editing with contextual awareness
Hanyuan Xiao, Yingshu Chen, Huajian Huang, Haolin Xiong, Jing Yang, Pratusha Prasad, and Yajie Zhao. Lo- calized gaussian splatting editing with contextual awareness. arXiv preprint arXiv:2408.00083, 2024. 2, 3
2024 arXiv
-
[119]
DMV3d: Denoising multi- view diffusion using 3d large reconstruction model
Yinghao Xu, Hao Tan, Fujun Luan, Sai Bi, Peng Wang, Ji- ahao Li, Zifan Shi, Kalyan Sunkavalli, Gordon Wetzstein, Zexiang Xu, and Kai Zhang. DMV3d: Denoising multi- view diffusion using 3d large reconstruction model. In The Twelfth International Conference on Learning Represent...
2024
-
[120]
Depth anything v2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xi- aogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything v2. arXiv preprint arXiv:2406.09414, 2024. 4, 8, 17, 20
2024 arXiv
-
[121]
Prometheus: 3d-aware latent diffusion models for feed-forward text-to-3d scene genera- tion
Yuanbo Yang, Jiahao Shao, Xinyang Li, Yujun Shen, An- dreas Geiger, and Yiyi Liao. Prometheus: 3d-aware latent diffusion models for feed-forward text-to-3d scene genera- tion. arXiv preprint arXiv:2412.21117, 2024. 21
2024 arXiv
-
[122]
Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models
Taoran Yi, Jiemin Fang, Junjie Wang, Guanjun Wu, Lingxi Xie, Xiaopeng Zhang, Wenyu Liu, Qi Tian, and Xinggang Wang. Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision...
2024
-
[123]
Diffusion time-step curricu- lum for one image to 3d generation
Xuanyu Yi, Zike Wu, Qingshan Xu, Pan Zhou, Joo-Hwee Lim, and Hanwang Zhang. Diffusion time-step curricu- lum for one image to 3d generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9948–9958, 2024. 3
2024
-
[124]
Edit-diffnerf: Editing 3d neural radiance fields using 2d diffusion model
Lu Yu, Wei Xiang, and Kang Han. Edit-diffnerf: Editing 3d neural radiance fields using 2d diffusion model. arXiv preprint arXiv:2306.09551, 2023. 1, 3
2023 arXiv
-
[125]
Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis.arXiv preprint arXiv:2409.02048, 2024
Wangbo Yu, Jinbo Xing, Li Yuan, Wenbo Hu, Xiaoyu Li, Zhipeng Huang, Xiangjun Gao, Tien-Tsin Wong, Ying Shan, and Yonghong Tian. Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis.arXiv preprint arXiv:2409.02048, 2024. 2
2024 arXiv
-
[126]
Mvimgnet: A large- scale dataset of multi-view images
Xianggang Yu, Mutian Xu, Yidan Zhang, Haolin Liu, Chongjie Ye, Yushuang Wu, Zizheng Yan, Chenming Zhu, Zhangyang Xiong, Tianyou Liang, et al. Mvimgnet: A large- scale dataset of multi-view images. In Proceedings of the IEEE/CVF conference on computer vision and pattern recog- ...
2023
-
[127]
Gaussiancube: Structuring gaussian splatting using opti- mal transport for 3d generative modeling
Bowen Zhang, Yiji Cheng, Jiaolong Yang, Chunyu Wang, Feng Zhao, Yansong Tang, Dong Chen, and Baining Guo. Gaussiancube: Structuring gaussian splatting using opti- mal transport for 3d generative modeling. arXiv preprint arXiv:2403.19655, 2024. 2
2024 arXiv
-
[128]
Text2nerf: Text-driven 3d scene generation with neu- ral radiance fields
Jingbo Zhang, Xiaoyu Li, Ziyu Wan, Can Wang, and Jing Liao. Text2nerf: Text-driven 3d scene generation with neu- ral radiance fields. IEEE Transactions on Visualization & Computer Graphics, 30(12):7749–7762, 2024. 1
2024
-
[129]
Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani
Jason Y . Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani. Cameras as rays: Pose estimation via ray diffusion. In The Twelfth Inter- national Conference on Learning Representations, 2024. 2, 5, 7, 8, 15, 16
2024
-
[130]
Gs-lrm: Large reconstruction model for 3d gaussian splatting
Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, and Zexiang Xu. Gs-lrm: Large reconstruction model for 3d gaussian splatting. In European Conference on Computer Vision, pages 1–19. Springer, 2024. 2
2024
-
[131]
Point cloud part editing: Segmentation, generation, assembly, and selection
Kaiyi Zhang, Yang Chen, Ximing Yang, Weizhong Zhang, and Cheng Jin. Point cloud part editing: Segmentation, generation, assembly, and selection. In Proceedings of 13 the AAAI Conference on Artificial Intelligence, pages 7187– 7195, 2024. 3
2024
-
[132]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 5, 8, 15, 18
2018
-
[133]
Dragtex: Generative point-based texture editing on 3d mesh
Yudi Zhang, Qi Xu, and Lei Zhang. Dragtex: Generative point-based texture editing on 3d mesh. arXiv preprint arXiv:2403.02217, 2024. 3
2024 arXiv
-
[134]
3d shape generation and completion through point-voxel diffusion
Linqi Zhou, Yilun Du, and Jiajun Wu. 3d shape generation and completion through point-voxel diffusion. In Proceed- ings of the IEEE/CVF international conference on computer vision, pages 5826–5835, 2021. 2
2021
-
[135]
Dreamscene360: Uncon- strained text-to-3d scene generation with panoramic gaus- sian splatting
Shijie Zhou, Zhiwen Fan, Dejia Xu, Haoran Chang, Pradyumna Chari, Tejas Bharadwaj, Suya You, Zhangyang Wang, and Achuta Kadambi. Dreamscene360: Uncon- strained text-to-3d scene generation with panoramic gaus- sian splatting. In European Conference on Computer Vision, pages 324...
2025
-
[136]
Tip-editor: An accurate 3d editor fol- lowing both text-prompts and image-prompts
Jingyu Zhuang, Di Kang, Yan-Pei Cao, Guanbin Li, Liang Lin, and Ying Shan. Tip-editor: An accurate 3d editor fol- lowing both text-prompts and image-prompts. ACM Trans- actions on Graphics (TOG), 43(4):1–12, 2024. 2, 3 14 SplatFlow: Multi-View Rectified Flow Model for 3D Gauss...
2024
-
[138]
Keep other details consistent with the original setting ( e.g., background, lighting)
Object Replacement Focus: Change the main object in the caption to a different but plausible one for the scene. Keep other details consistent with the original setting ( e.g., background, lighting)
-
[139]
Avoid improbable replacements that clash with the scene’s context or elements
Natural Integration: Ensure the new object fits logically within the environment described. Avoid improbable replacements that clash with the scene’s context or elements
-
[140]
Clarity and Directness: Use clear and straightforward language to describe the new object in place of the original, reflecting the same style as the given caption
-
[141]
Single Object Focus: Most images will contain a single object, so focus solely on replacing that object without altering other aspects of the scene unless explicitly instructed
-
[142]
If the given caption describe empty scene like empty hallway, add a ’new object’ to the scene
-
[143]
A white tag with a green leaf design and the text \
Just return the text. Example: - Input Caption: "A white tag with a green leaf design and the text \"HEY!\" on it.|Leaf- shaped tag on hanger, black and white checkered background." - Target Caption: "A sleek metallic spoon with a reflective surface on a plaid fabric." 4Implem...
-
[144]
Two durians with spiky skins and a purple tag
Vision-aided GAN loss: Adding vision-aided GAN loss after 200K iterations yielded the best performance across all metrics. Compared to training without vision-aided GAN loss, it significantly improved perceptual quality, as evidenced by better LPIPS and FID scores. Additionall...
-
[453]
Springer, 2025. 1, 3
2025
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.