Pith. sign in

REVIEW 1 major objections 2 minor 5 cited by

A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation

T0 review · 1 major / 2 minor · reviewed 2026-05-18 · grok-4.3

Pith's one-line read 3D Gaussian Splatting applications are organized into segmentation, editing, and generation tasks.

desk verdict This is a standard survey that organizes 3DGS work into segmentation, editing, and generation with decent coverage of methods and benchmarks, but the taxonomy has overlaps and possible gaps. read the letter →

arxiv 2508.09977 v4 submitted 2025-08-13 cs.CV

classification cs.CV
keywords 3DGaussianSplattingSegmentationEditingGenerationNovelviewsynthesisSurveyComputervisionDownstreamapplications
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to map recent uses of 3D Gaussian Splatting beyond basic scene rendering. It groups these uses into segmentation to identify parts of scenes, editing to modify them, and generation to create new content, plus some related functional tasks. For each group the survey collects example methods along with their supervision strategies and learning approaches. It also reviews common datasets, evaluation methods, and benchmark comparisons to show patterns and trends. The structure is intended to help researchers find relevant work and spot shared ideas across the area.

What carries the argument

The three-category taxonomy of segmentation, editing, and generation that structures the review and brings out shared supervision and learning patterns.

What would settle it

Publication of many 3D Gaussian Splatting application papers that cannot be placed in segmentation, editing, or generation would show the categorization misses major parts of the field.

Watch

Extended reading notes

Core claim

The survey establishes that the explicit and compact form of 3D Gaussian Splatting supports a range of tasks needing geometric and semantic understanding, and that these tasks can be grouped into segmentation, editing, and generation as foundational categories, with methods drawing on 2D foundation models and prior NeRF work to reveal common design principles.

Load-bearing premise

The chosen split of applications into segmentation, editing, and generation plus related functions accurately reflects the main structure of the research area.

Editorial extensions

If this is right

  • Common supervision strategies and learning paradigms become visible across task types.
  • Datasets and evaluation protocols enable direct comparisons of methods on public benchmarks.
  • Design principles identified in each category can guide development of new techniques.
  • The maintained repository of papers and code supports tracking further progress.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same taxonomy approach could be used to organize applications of other explicit 3D representations.
  • New application types may appear that require expanding or revising the current categories.
  • Benchmark comparisons could point to performance differences that suggest specific future improvements.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 2 minor

Summary. This survey reviews 3D Gaussian Splatting (3DGS) applications beyond novel view synthesis. It first covers reconstruction preliminaries, problem formulations, 2D foundation models, and related NeRF work, then organizes applications into three foundational tasks—segmentation, editing, and generation—plus additional functional applications. For each category the paper summarizes representative methods, supervision strategies, and learning paradigms, highlights shared principles and trends, and provides datasets, benchmarks, and comparative analyses while maintaining a public GitHub repository of resources.

Significance. A well-executed survey in this rapidly growing area would help researchers navigate the literature on 3DGS downstream tasks. The explicit maintenance of a continually updated repository (https://github.com/heshuting555/Awesome-3DGS-Applications) is a concrete strength that aids reproducibility and community use. If the taxonomy is justified and coverage is representative, the work would usefully synthesize supervision and paradigm trends across the three core tasks.

major comments (1)
  1. [Categorization section] Categorization section (following the preliminaries review): the central claim that the tripartite taxonomy plus functional applications delivers a comprehensive overview rests on the unstated assumption that the chosen framing accurately reflects field structure. The paper should add an explicit discussion of boundary porosity (e.g., editing methods that presuppose segmentation) and state inclusion/exclusion criteria for surveyed works to address possible omissions in areas such as physics-aware simulation or medical volumetric analysis.
minor comments (2)
  1. [Abstract] Abstract: the phrase 'additional functional applications built upon or tightly coupled with these foundational capabilities' is vague; a short parenthetical list of examples would improve clarity.
  2. [Introduction / Resources] The GitHub link is given but no statement is made about how frequently it is updated or what curation process is used; adding one sentence on maintenance policy would strengthen the reproducibility claim.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the positive assessment and recommendation for minor revision. We address the single major comment below and will incorporate the suggested clarifications to improve transparency of the taxonomy.

read point-by-point responses
  1. Referee: [Categorization section] Categorization section (following the preliminaries review): the central claim that the tripartite taxonomy plus functional applications delivers a comprehensive overview rests on the unstated assumption that the chosen framing accurately reflects field structure. The paper should add an explicit discussion of boundary porosity (e.g., editing methods that presuppose segmentation) and state inclusion/exclusion criteria for surveyed works to address possible omissions in areas such as physics-aware simulation or medical volumetric analysis.

    Authors: We agree that an explicit justification of the taxonomy and its boundaries will strengthen the manuscript. In the revised version we will add a short dedicated paragraph (or subsection) right after the preliminaries that (i) states the rationale for the tripartite core (segmentation, editing, generation) plus functional applications, namely that these categories correspond to the dominant research threads observed in the literature at the time of writing; (ii) discusses boundary porosity with concrete examples, such as editing pipelines that first invoke segmentation to obtain semantic Gaussians or generation methods that condition on previously edited or segmented representations; and (iii) articulates inclusion/exclusion criteria: we survey methods that directly extend or apply 3DGS to the listed tasks in general scenes, drawing from peer-reviewed and arXiv papers up to our literature cutoff date. We will note that physics-aware simulation and medical volumetric analysis are emerging but still sparsely represented within the 3DGS literature; they are mentioned briefly under functional applications where relevant and flagged as promising directions for future dedicated surveys rather than being omitted by oversight. These additions preserve the existing structure while making the framing assumptions transparent. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: survey organizes external literature without derivations or self-referential reductions

full rationale

This paper is a literature survey that reviews reconstruction preliminaries, categorizes 3DGS applications into segmentation/editing/generation plus functional tasks, and summarizes methods from cited external works. No original equations, parameter fitting, or derivation chain exists that could reduce to the paper's own inputs by construction. The taxonomy is an organizational framework drawn from the field rather than a fitted or self-defined result, and all referenced supervision strategies and benchmarks originate from independent prior publications. Self-citations, if present, are not load-bearing for any central claim.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Being a survey, the paper does not introduce or rely on new free parameters, axioms, or invented entities; its contributions are organizational summaries of prior literature.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation." pith.science (2026). https://pith.science/paper/2508.09977

@misc{pith2026250809977,
  author       = {Pith},
  title        = {Pith review of: A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2508.09977}},
  note         = {Machine review of arXiv:2508.09977}
}
read the original abstract

In the context of novel view synthesis, 3D Gaussian Splatting (3DGS) has recently emerged as an efficient and competitive counterpart to Neural Radiance Field (NeRF), enabling high-fidelity photorealistic rendering in real time. Beyond novel view synthesis, the explicit and compact nature of 3DGS enables a wide range of downstream applications that require geometric and semantic understanding. This survey provides a comprehensive overview of recent progress in 3DGS applications. It first reviews the reconstruction preliminaries of 3DGS, followed by the problem formulation, 2D foundation models, and related NeRF-based research areas that inform downstream 3DGS applications. We then categorize 3DGS applications into three foundational tasks: segmentation, editing, and generation, alongside additional functional applications built upon or tightly coupled with these foundational capabilities. For each, we summarize representative methods, supervision strategies, and learning paradigms, highlighting shared design principles and emerging trends. Commonly used datasets and evaluation protocols are also summarized, along with comparative analyses of recent methods across public benchmarks. To support ongoing research and development, a continually updated repository of papers, code, and resources is maintained at https://github.com/heshuting555/Awesome-3DGS-Applications.

Figures

Figures reproduced from arXiv: 2508.09977 by the authors.

Figure 1
Figure 1. Overview of the three main 3DGS applications. (a) [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of the structure of this survey. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Examples from 13 commonly used datasets for segmenta [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 3DEditSafe: Defending 3D Editing Pipelines from Unsafe Generation

    cs.GR 2026-05 unverdicted novelty 7.0 of 10

    3DEditSafe adds generation-stage guidance, 3D safety regularization, semantic projection, residue suppression, and mask-aware preservation to reduce unsafe semantic alignment in 3D editing while noting a safety-qualit...

  2. NG-GS: NeRF-Guided 3D Gaussian Splatting Segmentation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    NG-GS uses NeRF guidance and RBF interpolation on 3DGS to produce smoother, higher-quality object segmentation boundaries.

  3. GS4City: Hierarchical Semantic Gaussian Splatting via City-Model Priors

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    GS4City derives geometry-grounded semantic masks from LoD3 CityGML models via raycasting and fuses them with 2D foundation model outputs to supervise identity encodings on Gaussians, improving coarse and fine semantic...

  4. SVR-GS: Spatially Variant Regularization for Probabilistic Masks in 3D Gaussian Splatting

    cs.CV 2025-09 conditional novelty 6.0 of 10

    SVR-GS replaces MaskGS's global mask average with a per-pixel spatial mask regularizer, cutting Gaussian counts by up to 5.63x over 3DGS with about 0.4-0.5 dB average PSNR loss.

  5. Meshless Domain Randomization via Explicit Parameter Perturbation of 3D Gaussian Splatting

    cs.GR 2026-07 conditional novelty 5.0 of 10

    Perturbing spherical-harmonic colors and overlaying procedural 3D noise on 3D Gaussian Splat reconstructions creates randomized, meshless synthetic datasets for domain randomization.

Reference graph

Works this paper leans on

300 extracted references · 300 canonical work pages · cited by 5 Pith papers

  1. [1]

    3d gaussian splatting for real-time radiance field rendering

    B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering.” ACM TOG, 2023

  2. [2]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in ECCV, 2020

  3. [3]

    Scaffold- gs: Structured 3d gaussians for view-adaptive rendering,

    T. Lu, M. Yu, L. Xu, Y . Xiangli, L. Wang, D. Lin, and B. Dai, “Scaffold- gs: Structured 3d gaussians for view-adaptive rendering,” in CVPR, 2024

  4. [4]

    Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images,

    Y . Chen, H. Xu, C. Zheng, B. Zhuang, M. Pollefeys, A. Geiger, T.-J. Cham, and J. Cai, “Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images,” in ECCV, 2024

  5. [5]

    Gs-slam: Dense visual slam with 3d gaussian splatting,

    C. Yan, D. Qu, D. Xu, B. Zhao, Z. Wang, D. Wang, and X. Li, “Gs-slam: Dense visual slam with 3d gaussian splatting,” in CVPR, 2024

  6. [6]

    Gaussianavatars: Photorealistic head avatars with rigged 3d gaussians,

    S. Qian, T. Kirschstein, L. Schoneveld, D. Davoli, S. Giebenhain, and M. Nießner, “Gaussianavatars: Photorealistic head avatars with rigged 3d gaussians,” in CVPR, 2024

  7. [7]

    ReferSplat: Referring segmentation in 3d gaussian splatting,

    S. He, G. Jie, C. Wang, Y . Zhou, S. Hu, G. Li, and H. Ding, “ReferSplat: Referring segmentation in 3d gaussian splatting,” in ICML, 2025

  8. [8]

    Gaussianeditor: Swift and controllable 3d editing with gaussian splatting,

    Y . Chen, Z. Chen, C. Zhang, F. Wang, X. Yang, Y . Wang, Z. Cai, L. Yang, H. Liu, and G. Lin, “Gaussianeditor: Swift and controllable 3d editing with gaussian splatting,” in CVPR, 2024

Show all 300 references
  1. [9]

    Dreamgaussian: Generative gaussian splatting for efficient 3d content creation,

    J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng, “Dreamgaussian: Generative gaussian splatting for efficient 3d content creation,” inICLR, 2024

  2. [10]

    A survey on 3d gaussian splatting,

    G. Chen and W. Wang, “A survey on 3d gaussian splatting,” arXiv preprint arXiv:2401.03890, 2024

  3. [11]

    3d gaussian splatting in robotics: A survey,

    S. Zhu, G. Wang, X. Kong, D. Kong, and H. Wang, “3d gaussian splatting in robotics: A survey,”arXiv preprint arXiv:2410.12262, 2024

  4. [12]

    3dgs. zip: A survey on 3d gaussian splatting compression methods,

    M. T. Bagdasarian, P. Knoll, Y .-H. Li, F. Barthel, A. Hilsmann, P. Eisert, and W. Morgenstern, “3dgs. zip: A survey on 3d gaussian splatting compression methods,” arXiv preprint arXiv:2407.09510, 2024

  5. [13]

    Compression in 3d gaussian splatting: A survey of methods, trends, and future directions,

    M. S. Ali, C. Zhang, M. Cagnazzo, G. Valenzise, E. Tartaglione, and S.-H. Bae, “Compression in 3d gaussian splatting: A survey of methods, trends, and future directions,” arXiv preprint arXiv:2502.19457, 2025

  6. [14]

    Recent advances in 3d gaussian splatting,

    T. Wu, Y .-J. Yuan, L.-X. Zhang, J. Yang, Y .-P. Cao, L.-Q. Yan, and L. Gao, “Recent advances in 3d gaussian splatting,” Computational Visual Media, 2024

  7. [15]

    3d gaussian splatting as new era: A survey,

    B. Fei, J. Xu, R. Zhang, Q. Zhou, W. Yang, and Y . He, “3d gaussian splatting as new era: A survey,” IEEE TVCG, 2024

  8. [16]

    3d gaussian splatting: Survey, technologies, challenges, and opportunities,

    Y . Bao, T. Ding, J. Huo, Y . Liu, Y . Li, W. Li, Y . Gao, and J. Luo, “3d gaussian splatting: Survey, technologies, challenges, and opportunities,” IEEE TCSVT, 2025

  9. [17]

    Rt-gs2: Real-time generalizable semantic segmentation for 3d gaussian representations of radiance fields,

    M.-B. Jurca, R. Royen, I. Giosan, and A. Munteanu, “Rt-gs2: Real-time generalizable semantic segmentation for 3d gaussian representations of radiance fields,” BMVC, 2024

  10. [18]

    Omniseg3d: Omniversal 3d segmentation via hierarchical contrastive learning,

    H. Ying, Y . Yin, J. Zhang, F. Wang, T. Yu, R. Huang, and L. Fang, “Omniseg3d: Omniversal 3d segmentation via hierarchical contrastive learning,” in CVPR, 2024

  11. [19]

    Gaga: Group any gaussians via 3d-aware memory bank,

    W. Lyu, X. Li, A. Kundu, Y .-H. Tsai, and M.-H. Yang, “Gaga: Group any gaussians via 3d-aware memory bank,” arXiv preprint arXiv:2404.07977, 2024

  12. [20]

    Rethinking end-to-end 2d to 3d scene segmentation in gaussian splatting,

    R. Zhu, S. Qiu, Z. Liu, K.-H. Hui, Q. Wu, P.-A. Heng, and C.-W. Fu, “Rethinking end-to-end 2d to 3d scene segmentation in gaussian splatting,” in CVPR, 2025

  13. [21]

    Pointmap association and piecewise-plane constraint for consistent and compact 3d gaussian segmentation field,

    W. Hu, W. Chai, S. Hao, X. Cui, X. Wen, J.-N. Hwang, and G. Wang, “Pointmap association and piecewise-plane constraint for consistent and compact 3d gaussian segmentation field,” arXiv, 2025

  14. [22]

    Gaussian grouping: Segment and edit anything in 3d scenes,

    M. Ye, M. Danelljan, F. Yu, and L. Ke, “Gaussian grouping: Segment and edit anything in 3d scenes,” in ECCV, 2024

  15. [23]

    Sagd: Boundary-enhanced segment anything in 3d gaussian via gaus- sian decomposition,

    X. Hu, Y . Wang, L. Fan, J. Fan, J. Peng, Z. Lei, Q. Li, and Z. Zhang, “Sagd: Boundary-enhanced segment anything in 3d gaussian via gaus- sian decomposition,” arXiv preprint arXiv:2401.17857, 2024

  16. [24]

    Click-gaussian: Interactive segmentation to any 3d gaussians,

    S. Choi, H. Song, J. Kim, T. Kim, and H. Do, “Click-gaussian: Interactive segmentation to any 3d gaussians,” in ECCV, 2024

  17. [25]

    isegman: Interactive segment-and-manipulate 3d gaussians,

    Y . Zhao, W. Xu, R. Zheng, P. Qiao, C. Liu, and J. Chen, “isegman: Interactive segment-and-manipulate 3d gaussians,” in CVPR, 2025

  18. [26]

    Langsplat: 3d language gaussian splatting,

    M. Qin, W. Li, J. Zhou, H. Wang, and H. Pfister, “Langsplat: 3d language gaussian splatting,” in CVPR, 2024

  19. [27]

    Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields,

    S. Zhou, H. Chang, S. Jiang, Z. Fan, Z. Zhu, D. Xu, P. Chari, S. You, Z. Wang, and A. Kadambi, “Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields,” in CVPR, 2024

  20. [28]

    3d vision-language gaussian splatting,

    Q. Peng, B. Planche, Z. Gao, M. Zheng, A. Choudhuri, T. Chen, C. Chen, and Z. Wu, “3d vision-language gaussian splatting,” ICLR, 2025

  21. [29]

    Gaussiancut: Interactive segmentation via graph cut for 3d gaussian splatting,

    U. Jain, A. Mirzaei, and I. Gilitschenski, “Gaussiancut: Interactive segmentation via graph cut for 3d gaussian splatting,” inNeurIPS, 2024

  22. [30]

    Segment any 3d gaussians,

    J. Cen, J. Fang, C. Yang, L. Xie, X. Zhang, W. Shen, and Q. Tian, “Segment any 3d gaussians,” in AAAI, 2025

  23. [31]

    Instancegaussian: Appearance-semantic joint gaussian representation for 3d instance-level perception,

    H. Li, Y . Wu, J. Meng, Q. Gao, Z. Zhang, R. Wang, and J. Zhang, “Instancegaussian: Appearance-semantic joint gaussian representation for 3d instance-level perception,” CVPR, 2024

  24. [32]

    Opengaussian: Towards point-level 3d gaussian-based open vocabulary understanding,

    Y . Wu, J. Meng, H. Li, C. Wu, Y . Shi, X. Cheng, C. Zhao, H. Feng, E. Ding, J. Wang et al. , “Opengaussian: Towards point-level 3d gaussian-based open vocabulary understanding,” in NeurIPS, 2024

  25. [33]

    Tip-editor: An accurate 3d editor following both text-prompts and image-prompts,

    J. Zhuang, D. Kang, Y .-P. Cao, G. Li, L. Lin, and Y . Shan, “Tip-editor: An accurate 3d editor following both text-prompts and image-prompts,” ACM TOG, 2024

  26. [34]

    View- consistent 3d editing with gaussian splatting,

    Y . Wang, X. Yi, Z. Wu, N. Zhao, L. Chen, and H. Zhang, “View- consistent 3d editing with gaussian splatting,” in ECCV, 2024

  27. [35]

    Gaussianeditor: Editing 3d gaussians delicately with text instructions,

    J. Wang, J. Fang, X. Zhang, L. Xie, and Q. Tian, “Gaussianeditor: Editing 3d gaussians delicately with text instructions,” in CVPR, 2024

  28. [36]

    Dreamcatalyst: Fast and high-quality 3d editing via controlling editability and identity preservation,

    J. Kim, S. Lee, J. Shin, J. Choi, and H. Shim, “Dreamcatalyst: Fast and high-quality 3d editing via controlling editability and identity preservation,” in ICLR, 2025

  29. [37]

    Lgm: Large multi-view gaussian model for high-resolution 3d content creation,

    J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu, “Lgm: Large multi-view gaussian model for high-resolution 3d content creation,” in ECCV, 2024

  30. [38]

    Grm: Large gaussian reconstruction model for efficient 3d reconstruction and generation,

    Y . Xu, Z. Shi, W. Yifan, H. Chen, C. Yang, S. Peng, Y . Shen, and G. Wetzstein, “Grm: Large gaussian reconstruction model for efficient 3d reconstruction and generation,” in ECCV, 2024

  31. [39]

    Gs-lrm: Large reconstruction model for 3d gaussian splatting,

    K. Zhang, S. Bi, H. Tan, Y . Xiangli, N. Zhao, K. Sunkavalli, and Z. Xu, “Gs-lrm: Large reconstruction model for 3d gaussian splatting,” in ECCV, 2024

  32. [40]

    Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models,

    T. Yi, J. Fang, J. Wang, G. Wu, L. Xie, X. Zhang, W. Liu, Q. Tian, and X. Wang, “Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models,” in CVPR, 2024

  33. [41]

    Luciddreamer: Towards high-fidelity text-to-3d generation via interval score matching,

    Y . Liang, X. Yang, J. Lin, H. Li, X. Xu, and Y . Chen, “Luciddreamer: Towards high-fidelity text-to-3d generation via interval score matching,” in CVPR, 2024

  34. [42]

    Text2room: Extracting textured 3d meshes from 2d text-to-image models,

    L. Höllein, A. Cao, A. Owens, J. Johnson, and M. Nießner, “Text2room: Extracting textured 3d meshes from 2d text-to-image models,” inICCV, 2023

  35. [43]

    Dreamscene: 3d gaussian-based text-to-3d scene generation via formation pattern sampling,

    H. Li, H. Shi, W. Zhang, W. Wu, Y . Liao, L. Wang, L.-h. Lee, and P. Y . Zhou, “Dreamscene: 3d gaussian-based text-to-3d scene generation via formation pattern sampling,” in ECCV, 2024

  36. [44]

    Dreamscene360: Unconstrained text-to-3d scene generation with panoramic gaussian splatting,

    S. Zhou, Z. Fan, D. Xu, H. Chang, P. Chari, T. Bharadwaj, S. You, Z. Wang, and A. Kadambi, “Dreamscene360: Unconstrained text-to-3d scene generation with panoramic gaussian splatting,” in ECCV, 2024

  37. [45]

    Emerging properties in self-supervised vision transformers,

    M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in ICCV, 2021

  38. [46]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al. , “Dinov2: Learning robust visual features without supervision,” TMLR, 2024

  39. [47]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in ICML, 2021

  40. [48]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Loet al., “Segment anything,” in ICCV, 2023

  41. [49]

    Sam 2: Segment anything in images and videos,

    N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson et al., “Sam 2: Segment anything in images and videos,” in ICLR, 2025. 16

  42. [50]

    MOSEv2: A more challenging dataset for video object segmentation in complex scenes,

    H. Ding, K. Ying, C. Liu, S. He, X. Jiang, Y .-G. Jiang, P. H. Torr, and S. Bai, “MOSEv2: A more challenging dataset for video object segmentation in complex scenes,” arXiv preprint arXiv:2508.05630 , 2025

  43. [51]

    MOSE: A new dataset for video object segmentation in complex scenes,

    H. Ding, C. Liu, S. He, X. Jiang, P. H. Torr, and S. Bai, “MOSE: A new dataset for video object segmentation in complex scenes,” in ICCV, 2023

  44. [52]

    MeViS: A large-scale benchmark for video segmentation with motion expressions,

    H. Ding, C. Liu, S. He, X. Jiang, and C. C. Loy, “MeViS: A large-scale benchmark for video segmentation with motion expressions,” in ICCV, 2023

  45. [53]

    MeViS: A multi-modal dataset for referring motion expression video segmentation,

    H. Ding, C. Liu, S. He, K. Ying, X. Jiang, C. C. Loy, and Y .-G. Jiang, “MeViS: A multi-modal dataset for referring motion expression video segmentation,” IEEE TPAMI, 2025

  46. [54]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” NeurIPS, 2020

  47. [55]

    High- resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High- resolution image synthesis with latent diffusion models,” in CVPR, 2022

  48. [56]

    Scalable diffusion models with transformers,

    W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in ICCV, 2023

  49. [57]

    B. F. Labs, “Flux,” https://github.com/black-forest-labs/flux, 2024

  50. [58]

    In-place scene labelling and understanding with implicit scene representation,

    S. Zhi, T. Laidlow, S. Leutenegger, and A. J. Davison, “In-place scene labelling and understanding with implicit scene representation,” in CVPR, 2021

  51. [59]

    Decomposing nerf for editing via feature field distillation,

    S. Kobayashi, E. Matsumoto, and V . Sitzmann, “Decomposing nerf for editing via feature field distillation,” NeurIPS, 2022

  52. [60]

    Neural feature fu- sion fields: 3d distillation of self-supervised 2d image representations,

    V . Tschernezki, I. Laina, D. Larlus, and A. Vedaldi, “Neural feature fu- sion fields: 3d distillation of self-supervised 2d image representations,” in 3DV, 2022

  53. [61]

    Interactive segmen- tation of radiance fields,

    R. Goel, D. Sirikonda, S. Saini, and P. Narayanan, “Interactive segmen- tation of radiance fields,” in CVPR, 2023

  54. [62]

    Lerf: Language embedded radiance fields,

    J. Kerr, C. M. Kim, K. Goldberg, A. Kanazawa, and M. Tancik, “Lerf: Language embedded radiance fields,” in ICCV, 2023

  55. [63]

    Weakly supervised 3d open-vocabulary segmenta- tion,

    K. Liu, F. Zhan, J. Zhang, M. Xu, Y . Yu, A. El Saddik, C. Theobalt, E. Xing, and S. Lu, “Weakly supervised 3d open-vocabulary segmenta- tion,” NeurIPS, 2023

  56. [64]

    Dm-nerf: 3d scene geometry decom- position and manipulation from 2d images,

    B. Wang, L. Chen, and B. Yang, “Dm-nerf: 3d scene geometry decom- position and manipulation from 2d images,” in ICLR, 2023

  57. [65]

    Panoptic lifting for 3d scene understanding with neural fields,

    Y . Siddiqui, L. Porzi, S. R. Buló, N. Müller, M. Nießner, A. Dai, and P. Kontschieder, “Panoptic lifting for 3d scene understanding with neural fields,” in CVPR, 2023

  58. [66]

    Contrastive lift: 3d object instance segmentation by slow-fast con- trastive fusion,

    Y . Bhalgat, I. Laina, J. F. Henriques, A. Zisserman, and A. Vedaldi, “Contrastive lift: 3d object instance segmentation by slow-fast con- trastive fusion,” in NeurIPS, 2023

  59. [67]

    Spin-nerf: Multiview segmentation and perceptual inpainting with neural radiance fields,

    A. Mirzaei, T. Aumentado-Armstrong, K. G. Derpanis, J. Kelly, M. A. Brubaker, I. Gilitschenski, and A. Levinshtein, “Spin-nerf: Multiview segmentation and perceptual inpainting with neural radiance fields,” in CVPR, 2023

  60. [68]

    Segment anything in 3d with nerfs,

    J. Cen, Z. Zhou, J. Fang, W. Shen, L. Xie, D. Jiang, X. Zhang, Q. Tian et al., “Segment anything in 3d with nerfs,” in NeurIPS, 2023

  61. [69]

    Garfield: Group anything with radiance fields,

    C. M. Kim, M. Wu, J. Kerr, K. Goldberg, M. Tancik, and A. Kanazawa, “Garfield: Group anything with radiance fields,” in CVPR, 2024

  62. [70]

    Editing conditional radiance fields,

    S. Liu, X. Zhang, Z. Zhang, R. Zhang, J.-Y . Zhu, and B. Russell, “Editing conditional radiance fields,” in CVPR, 2021

  63. [71]

    Clip-nerf: Text-and- image driven manipulation of neural radiance fields,

    C. Wang, M. Chai, M. He, D. Chen, and J. Liao, “Clip-nerf: Text-and- image driven manipulation of neural radiance fields,” in CVPR, 2022

  64. [72]

    Learning object-compositional neural radiance field for editable scene rendering,

    B. Yang, Y . Zhang, Y . Xu, Y . Li, H. Zhou, H. Bao, G. Zhang, and Z. Cui, “Learning object-compositional neural radiance field for editable scene rendering,” in ICCV, 2021

  65. [73]

    Laterf: Label and text driven object radiance fields,

    A. Mirzaei, Y . Kant, J. Kelly, and I. Gilitschenski, “Laterf: Label and text driven object radiance fields,” in ECCV, 2022

  66. [74]

    Conerf: Controllable neural radiance fields,

    K. Kania, K. M. Yi, M. Kowalski, T. Trzci ´nski, and A. Tagliasacchi, “Conerf: Controllable neural radiance fields,” in CVPR, 2022

  67. [75]

    Instruct-nerf2nerf: Editing 3d scenes with instructions,

    A. Haque, M. Tancik, A. A. Efros, A. Holynski, and A. Kanazawa, “Instruct-nerf2nerf: Editing 3d scenes with instructions,” inICCV, 2023

  68. [76]

    Graf: Generative radiance fields for 3d-aware image synthesis,

    K. Schwarz, Y . Liao, M. Niemeyer, and A. Geiger, “Graf: Generative radiance fields for 3d-aware image synthesis,” NeurIPS, 2020

  69. [77]

    Gram: Generative radiance manifolds for 3d-aware image generation,

    Y . Deng, J. Yang, J. Xiang, and X. Tong, “Gram: Generative radiance manifolds for 3d-aware image generation,” in CVPR, 2022

  70. [78]

    Dreamfusion: Text- to-3d using 2d diffusion,

    B. Poole, A. Jain, J. T. Barron, and B. Mildenhall, “Dreamfusion: Text- to-3d using 2d diffusion,” ICLR, 2023

  71. [79]

    Magic3d: High-resolution text-to- 3d content creation,

    C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y . Liu, and T.-Y . Lin, “Magic3d: High-resolution text-to- 3d content creation,” in CVPR, 2023

  72. [80]

    Latent-nerf for shape-guided generation of 3d shapes and textures,

    G. Metzer, E. Richardson, O. Patashnik, R. Giryes, and D. Cohen-Or, “Latent-nerf for shape-guided generation of 3d shapes and textures,” in CVPR, 2023

  73. [81]

    Score jaco- bian chaining: Lifting pretrained 2d diffusion models for 3d generation,

    H. Wang, X. Du, J. Li, R. A. Yeh, and G. Shakhnarovich, “Score jaco- bian chaining: Lifting pretrained 2d diffusion models for 3d generation,” in CVPR, 2023

  74. [82]

    Efficient geometry-aware 3d generative adversarial networks,

    E. R. Chan, C. Z. Lin, M. A. Chan, K. Nagano, B. Pan, S. De Mello, O. Gallo, L. J. Guibas, J. Tremblay, S. Khamis et al. , “Efficient geometry-aware 3d generative adversarial networks,” in CVPR, 2022

  75. [83]

    Text2nerf: Text-driven 3d scene generation with neural radiance fields,

    J. Zhang, X. Li, Z. Wan, C. Wang, and J. Liao, “Text2nerf: Text-driven 3d scene generation with neural radiance fields,” IEEE TVCG, 2024

  76. [84]

    Language embedded 3d gaussians for open-vocabulary scene understanding,

    J.-C. Shi, M. Wang, H.-B. Duan, and S.-H. Guan, “Language embedded 3d gaussians for open-vocabulary scene understanding,” inCVPR, 2024

  77. [85]

    Fast and efficient: Mask neural fields for 3d scene segmentation,

    Z. Gao, L. Li, L. Jiao, F. Liu, X. Liu, W. Ma, Y . Guo, and S. Yang, “Fast and efficient: Mask neural fields for 3d scene segmentation,”arXiv preprint arXiv:2407.01220, 2024

  78. [86]

    Semantic gaussians: Open- vocabulary scene understanding with 3d gaussian splatting,

    J. Guo, X. Ma, Y . Fan, H. Liu, and Q. Li, “Semantic gaussians: Open- vocabulary scene understanding with 3d gaussian splatting,” arXiv preprint arXiv:2403.15624, 2024

  79. [87]

    Language-driven semantic segmentation,

    B. Li, K. Q. Weinberger, S. Belongie, V . Koltun, and R. Ranftl, “Language-driven semantic segmentation,” in ICLR, 2022

  80. [88]

    N2f2: Hierarchical scene understanding with nested neural feature fields,

    Y . Bhalgat, I. Laina, J. F. Henriques, A. Zisserman, and A. Vedaldi, “N2f2: Hierarchical scene understanding with nested neural feature fields,” in ECCV, 2024

  81. [89]

    Langsurf: Language-embedded surface gaussians for 3d scene under- standing,

    H. Li, R. Qin, Z. Zou, D. He, B. Li, B. Dai, D. Zhang, and J. Han, “Langsurf: Language-embedded surface gaussians for 3d scene under- standing,” arXiv preprint arXiv:2412.17635, 2024

  82. [90]

    Semantic consistent language gaussian splatting for point-level open-vocabulary querying,

    H. Yin, H. Zhan, Y . Xu, and R. A. Yeh, “Semantic consistent language gaussian splatting for point-level open-vocabulary querying,” arXiv preprint arXiv:2503.21767, 2025

  83. [91]

    Supergseg: Open-vocabulary 3d segmentation with structured super-gaussians,

    S. Liang, S. Wang, K. Li, M. Niemeyer, S. Gasperini, N. Navab, and F. Tombari, “Supergseg: Open-vocabulary 3d segmentation with structured super-gaussians,” arXiv preprint arXiv:2412.10231, 2024

  84. [92]

    Bootstraping clustering of gaussians for view-consistent 3d scene understanding,

    W. Zhang, L. Zhang, P. Hu, L. Ma, Y . Zhuge, and H. Lu, “Bootstraping clustering of gaussians for view-consistent 3d scene understanding,” in AAAI, 2025

  85. [93]

    Gls: Geometry-aware 3d language gaussian splatting,

    J. Qiu, L. Liu, Z. Su, and T. Lin, “Gls: Geometry-aware 3d language gaussian splatting,” arXiv preprint arXiv:2411.18066, 2024

  86. [94]

    Clip-gs: Clip-informed gaussian splatting for real-time and view-consistent 3d semantic understanding,

    G. Liao, J. Li, Z. Bao, X. Ye, J. Wang, Q. Li, and K. Liu, “Clip-gs: Clip-informed gaussian splatting for real-time and view-consistent 3d semantic understanding,” ACM TOMM, 2025

  87. [95]

    Fmgs: Foundation model embedded 3d gaussian splatting for holistic 3d scene understand- ing,

    X. Zuo, P. Samangouei, Y . Zhou, Y . Di, and M. Li, “Fmgs: Foundation model embedded 3d gaussian splatting for holistic 3d scene understand- ing,” IJCV, 2024

  88. [96]

    Goi: Find 3d gaussians of interest with an optimizable open-vocabulary semantic- space hyperplane,

    Y . Qu, S. Dai, X. Li, J. Lin, L. Cao, S. Zhang, and R. Ji, “Goi: Find 3d gaussians of interest with an optimizable open-vocabulary semantic- space hyperplane,” in ACM MM, 2024

  89. [97]

    Fastlgs: Speeding up language embedded gaussians with feature grid mapping,

    Y . Ji, H. Zhu, J. Tang, W. Liu, Z. Zhang, X. Tan, and Y . Xie, “Fastlgs: Speeding up language embedded gaussians with feature grid mapping,” in AAAI, 2025

  90. [98]

    Fmlgs: Fast multilevel language embedded gaussians for part-level interactive agents,

    X. Tan, Y . Ji, H. Zhu, and Y . Xie, “Fmlgs: Fast multilevel language embedded gaussians for part-level interactive agents,” arXiv, 2025

  91. [99]

    Efficient decoupled feature 3d gaussian splatting via hierarchical compression,

    Z. Dai, T. Liu, and Y . Zhang, “Efficient decoupled feature 3d gaussian splatting via hierarchical compression,” in CVPR, 2025

  92. [100]

    Gradiseg: Gradient-guided gaussian segmentation with enhanced 3d boundary precision,

    Z. Li, W. Han, Y . Cai, H. Jiang, B. Bi, S. Gao, H. Zhao, and Z. Wang, “Gradiseg: Gradient-guided gaussian segmentation with enhanced 3d boundary precision,” arXiv preprint arXiv:2412.00392, 2024

  93. [101]

    Cosseggaussians: Compact and swift scene segmenting 3d gaussians,

    B. Dou, T. Zhang, Y . Ma, Z. Wang, and Z. Yuan, “Cosseggaussians: Compact and swift scene segmenting 3d gaussians,” arXiv, 2024

  94. [102]

    Segment then splat: A unified approach for 3d open-vocabulary segmentation based on gaussian splatting,

    Y . Lu, Y . Zhou, Y . Qiao, C. Song, T. Liang, J. Ma, and Y . Yin, “Segment then splat: A unified approach for 3d open-vocabulary segmentation based on gaussian splatting,” arXiv preprint arXiv:2503.22204, 2025

  95. [103]

    Cags: Open-vocabulary 3d scene understanding with context-aware gaussian splatting,

    W. Sun, Y . Zhou, J. Jiao, and Y . Li, “Cags: Open-vocabulary 3d scene understanding with context-aware gaussian splatting,” arXiv, 2025

  96. [104]

    Segment anything in 3d with radiance fields,

    J. Cen, J. Fang, Z. Zhou, C. Yang, L. Xie, X. Zhang, W. Shen, and Q. Tian, “Segment anything in 3d with radiance fields,” IJCV, 2025

  97. [105]

    Contrastive gaussian clustering: Weakly supervised 3d scene segmentation,

    M. C. Silva, M. Dahaghin, M. Toso, and A. Del Bue, “Contrastive gaussian clustering: Weakly supervised 3d scene segmentation,” ICPR, 2024

  98. [106]

    Opensplat3d: Open-vocabulary 3d instance segmentation using gaussian splatting,

    J. Piekenbrinck, C. Schmidt, A. Hermans, N. Vaskevicius, T. Linder, and B. Leibe, “Opensplat3d: Open-vocabulary 3d instance segmentation using gaussian splatting,” in CVPRW, 2025

  99. [107]

    econsg: Efficient and multi-view consistent open-vocabulary 3d semantic gaussians,

    C. Zhang and G. H. Lee, “econsg: Efficient and multi-view consistent open-vocabulary 3d semantic gaussians,” in ICLR, 2025

  100. [108]

    Ccl-lgs: Contrastive codebook learning for 3d language gaussian splatting,

    L. Tian, X. Li, L. Ma, H. Huang, Z. Zheng, H. Yin, T. Li, H. Lu, and X. Jia, “Ccl-lgs: Contrastive codebook learning for 3d language gaussian splatting,” in ICCV, 2025

  101. [109]

    V ote- splat: Hough voting gaussian splatting for 3d scene understanding,

    M. Jiang, S. Jia, J. Gu, X. Lu, G. Zhu, A. Dong, and L. Zhang, “V ote- splat: Hough voting gaussian splatting for 3d scene understanding,” in ICCV, 2025. 17

  102. [110]

    Identity-aware language gaussian splatting for open-vocabulary 3d semantic segmentation,

    S. Jang and W. Kim, “Identity-aware language gaussian splatting for open-vocabulary 3d semantic segmentation,” in ICCV, 2025

  103. [111]

    Cob-gs: Clear object boundaries in 3dgs segmentation based on boundary-adaptive gaussian splitting,

    J. Zhang, J. Jiang, Y . Chen, K. Jiang, and X. Liu, “Cob-gs: Clear object boundaries in 3dgs segmentation based on boundary-adaptive gaussian splitting,” in CVPR, 2025

  104. [112]

    Panogs: Gaussian-based panoptic segmentation for 3d open vocabulary scene understanding,

    H. Zhai, H. Li, Z. Li, X. Pan, Y . He, and G. Zhang, “Panogs: Gaussian-based panoptic segmentation for 3d open vocabulary scene understanding,” in CVPR, 2025

  105. [113]

    Flashsplat: 2d to 3d gaussian splatting segmentation solved optimally,

    Q. Shen, X. Yang, and X. Wang, “Flashsplat: 2d to 3d gaussian splatting segmentation solved optimally,” in ECCV, 2024

  106. [114]

    Training-free hierarchical scene understanding for gaussian splatting with superpoint graphs,

    S. Dai, Y . Qu, Z. Li, X. Li, S. Zhang, and L. Cao, “Training-free hierarchical scene understanding for gaussian splatting with superpoint graphs,” arXiv preprint arXiv:2504.13153, 2025

  107. [115]

    Lifting by gaus- sians: A simple, fast and flexible method for 3d instance segmentation,

    R. Chacko, N. Häni, E. Khaliullin, L. Sun, and D. Lee, “Lifting by gaus- sians: A simple, fast and flexible method for 3d instance segmentation,” in WACV, 2025

  108. [116]

    Ludvig: Learning-free uplifting of 2d visual features to gaussian splatting scenes,

    J. Marrie, R. Ménégaux, M. Arbel, D. Larlus, and J. Mairal, “Ludvig: Learning-free uplifting of 2d visual features to gaussian splatting scenes,” in ICCV, 2025

  109. [117]

    Slgaussian: Fast language gaussian splatting in sparse views,

    K. Chen, B. Dai, M. Qin, D. Zhang, P. Li, Y . Zou, and H. Wang, “Slgaussian: Fast language gaussian splatting in sparse views,” in ACM MM, 2025

  110. [118]

    Gsemsplat: General- izable semantic 3d gaussian splatting from uncalibrated image pairs,

    X. Wang, C. Lan, H. Zhu, Z. Chen, and Y . Lu, “Gsemsplat: General- izable semantic 3d gaussian splatting from uncalibrated image pairs,” arXiv preprint arXiv:2412.16932, 2024

  111. [119]

    Dr. splat: Directly referring 3d gaussian splatting via direct language embedding registration,

    K. Jun-Seong, G. Kim, K. Yu-Ji, Y .-C. F. Wang, J. Choe, and T.-H. Oh, “Dr. splat: Directly referring 3d gaussian splatting via direct language embedding registration,” in CVPR, 2025

  112. [120]

    Semanticsplat: Feed- forward 3d scene understanding with language-aware gaussian fields,

    Q. Li, J. Sun, L. An, Z. Su, H. Zhang, and Y . Liu, “Semanticsplat: Feed- forward 3d scene understanding with language-aware gaussian fields,” arXiv preprint arXiv:2506.09565, 2025

  113. [121]

    Langscene-x: Reconstruct generalizable 3d language-embedded scenes with trimap video diffusion,

    F. Liu, H. Li, J. Chi, H. Wang, M. Yang, F. Wang, and Y . Duan, “Langscene-x: Reconstruct generalizable 3d language-embedded scenes with trimap video diffusion,” in ICCV, 2025

  114. [122]

    Spatialsplat: Efficient semantic 3d from sparse unposed images,

    Y . Sheng, J. Deng, X. Zhang, Y . Zhang, B. Hua, Y . Zhang, and J. Ji, “Spatialsplat: Efficient semantic 3d from sparse unposed images,”arXiv preprint arXiv:2505.23044, 2025

  115. [123]

    Large spatial model: End-to-end unposed images to semantic 3d,

    Z. Fan, J. Zhang, W. Cong, P. Wang, R. Li, K. Wen, S. Zhou, A. Kadambi, Z. Wang, D. Xu et al. , “Large spatial model: End-to-end unposed images to semantic 3d,” in NeurIPS, 2024

  116. [124]

    Dust3r: Geometric 3d vision made easy,

    S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud, “Dust3r: Geometric 3d vision made easy,” in CVPR, 2024

  117. [125]

    Point transformer,

    H. Zhao, L. Jiang, J. Jia, P. H. Torr, and V . Koltun, “Point transformer,” in ICCV, 2021

  118. [126]

    Cogvideox: Text-to-video diffusion models with an expert transformer,

    Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y . Yang, W. Hong, X. Zhang, G. Fenget al., “Cogvideox: Text-to-video diffusion models with an expert transformer,” in ICLR, 2025

  119. [127]

    Scenesplat: Gaussian splatting-based scene understanding with vision-language pretraining,

    Y . Li, Q. Ma, R. Yang, H. Li, M. Ma, B. Ren, N. Popovic, N. Sebe, E. Konukoglu, T. Gevers et al. , “Scenesplat: Gaussian splatting-based scene understanding with vision-language pretraining,” in ICCV, 2025

  120. [128]

    Instructpix2pix: Learning to follow image editing instructions,

    T. Brooks, A. Holynski, and A. A. Efros, “Instructpix2pix: Learning to follow image editing instructions,” in CVPR, 2023

  121. [129]

    Prompt-softbox- prompt: A free-text embedding control for image editing,

    Y . Yang, Y . Wang, T. Zhang, J. Wang, and S. He, “Prompt-softbox- prompt: A free-text embedding control for image editing,” inACM MM, 2025

  122. [130]

    Editsplat: Multi-view fusion and attention-guided optimization for view-consistent 3d scene editing with 3d gaussian splatting,

    D. I. Lee, H. Park, J. Seo, E. Park, H. Park, H. D. Baek, S. Shin, S. Kim, and S. Kim, “Editsplat: Multi-view fusion and attention-guided optimization for view-consistent 3d scene editing with 3d gaussian splatting,” in CVPR, 2025

  123. [131]

    Gsedit: Efficient text-guided editing of 3d objects via gaussian splatting,

    F. Palandra, A. Sanchietti, D. Baieri, and E. Rodolà, “Gsedit: Efficient text-guided editing of 3d objects via gaussian splatting,” arXiv preprint arXiv:2403.05154, 2024

  124. [132]

    Dacapo: Score distillation as stacked bridge for fast and high-quality 3d editing,

    Y . Huang, B. Liao, Y . Hu, H. Lin, L. Wu, S. Li, C. Tan, Z. Liu, Y . Liu, Z. Zang et al., “Dacapo: Score distillation as stacked bridge for fast and high-quality 3d editing,” in CVPR, 2025

  125. [133]

    Gseditpro: 3d gaussian splatting editing with attention-based progressive localization,

    Y . Sun, R. Tian, X. Han, X. Liu, Y . Zhang, and K. Xu, “Gseditpro: 3d gaussian splatting editing with attention-based progressive localization,” in Computer Graphics F orum, 2024

  126. [134]

    Dreambooth: Fine tuning text-to-image diffusion models for subject- driven generation,

    N. Ruiz, Y . Li, V . Jampani, Y . Pritch, M. Rubinstein, and K. Aberman, “Dreambooth: Fine tuning text-to-image diffusion models for subject- driven generation,” in CVPR, 2023

  127. [135]

    Localized gaussian splatting editing with contextual awareness,

    H. Xiao, Y . Chen, H. Huang, H. Xiong, J. Yang, P. Prasad, and Y . Zhao, “Localized gaussian splatting editing with contextual awareness,” in WACV, 2025

  128. [136]

    Gaussctrl: Multi-view consistent text-driven 3d gaussian splatting editing,

    J. Wu, J.-W. Bian, X. Li, G. Wang, I. Reid, P. Torr, and V . A. Prisacariu, “Gaussctrl: Multi-view consistent text-driven 3d gaussian splatting editing,” in ECCV, 2024

  129. [137]

    Trame: Trajectory-anchored multi-view editing for text-guided 3d gaussian splatting manipulation,

    C. Luo, D. Di, X. Yang, Y . Ma, Z. Xue, C. Wei, and Y . Liu, “Trame: Trajectory-anchored multi-view editing for text-guided 3d gaussian splatting manipulation,” IEEE TMM, 2025

  130. [138]

    Diffusion-based attention warping for consis- tent 3d scene editing,

    E. Gomel and L. Wolf, “Diffusion-based attention warping for consis- tent 3d scene editing,” arXiv preprint arXiv:2412.07984, 2024

  131. [139]

    Dge: Direct gaussian 3d editing by consistent multi-view editing,

    M. Chen, I. Laina, and A. Vedaldi, “Dge: Direct gaussian 3d editing by consistent multi-view editing,” in ECCV, 2024

  132. [140]

    Splatflow: Multi-view rectified flow model for 3d gaussian splatting synthesis,

    H. Go, B. Park, J. Jang, J.-Y . Kim, S. Kwon, and C. Kim, “Splatflow: Multi-view rectified flow model for 3d gaussian splatting synthesis,” in CVPR, 2025

  133. [141]

    Intergsedit: Interactive 3d gaussian splatting editing with 3d geometry-consistent attention prior,

    M. Wen, S. Wu, K. Wang, and D. Liang, “Intergsedit: Interactive 3d gaussian splatting editing with 3d geometry-consistent attention prior,” in ICCV, 2025

  134. [142]

    Progdf: Progressive gaussian differential field for controllable and flexible 3d editing,

    Y . Zhao, W. Xu, Y . Wu, W. Huang, Z. Sun, and W. Yang, “Progdf: Progressive gaussian differential field for controllable and flexible 3d editing,” arXiv preprint arXiv:2412.08152, 2024

  135. [143]

    3dsceneeditor: Controllable 3d scene editing with gaussian splatting,

    Z. Yan, L. Li, Y . Shao, S. Chen, W. Kai, J.-N. Hwang, H. Zhao, and F. Remondino, “3dsceneeditor: Controllable 3d scene editing with gaussian splatting,” arXiv preprint arXiv:2412.01583, 2024

  136. [144]

    3ditscene: Editing any scene via language-guided disentan- gled gaussian splatting,

    Q. Zhang, Y . Xu, C. Wang, H.-Y . Lee, G. Wetzstein, B. Zhou, and C. Yang, “3ditscene: Editing any scene via language-guided disentan- gled gaussian splatting,” in ICLR, 2025

  137. [145]

    Learning 3d geometry and feature consistent gaussian splatting for object removal,

    Y . Wang, Q. Wu, G. Zhang, and D. Xu, “Learning 3d geometry and feature consistent gaussian splatting for object removal,” in ECCV, 2024

  138. [146]

    D-miso: Editing dynamic 3d scenes using multi-gaussians soup,

    J. Waczynska, P. Borycki, J. Kaleta, S. Tadeja, and P. Spurek, “D-miso: Editing dynamic 3d scenes using multi-gaussians soup,”NeurIPS, 2024

  139. [147]

    Texture- gs: Disentangling the geometry and texture for 3d gaussian splatting editing,

    T.-X. Xu, W. Hu, Y .-K. Lai, Y . Shan, and S.-H. Zhang, “Texture- gs: Disentangling the geometry and texture for 3d gaussian splatting editing,” in ECCV, 2024

  140. [148]

    Ctrl-d: Controllable dynamic 3d scene editing with personalized 2d diffusion,

    K. He, C.-H. Wu, and I. Gilitschenski, “Ctrl-d: Controllable dynamic 3d scene editing with personalized 2d diffusion,” in CVPR, 2025

  141. [149]

    Lora: Low-rank adaptation of large language models,

    E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in ICLR, 2022

  142. [150]

    Gs-vton: Controllable 3d virtual try-on with gaussian splatting,

    Y . Cao, M. Hadi, L. Pan, and Z. Liu, “Gs-vton: Controllable 3d virtual try-on with gaussian splatting,” arXiv preprint arXiv:2410.05259, 2024

  143. [151]

    Perse: Personalized 3d generative avatars from a single portrait,

    H. Cha, I. Lee, and H. Joo, “Perse: Personalized 3d generative avatars from a single portrait,” arXiv preprint arXiv:2412.21206, 2024

  144. [152]

    Tiger: Text-instructed 3d gaussian retrieval and coherent editing,

    T. Xu, P. Chen, Jiamin an d Chen, Y . Zhang, J. Yu, and W. Yang, “Tiger: Text-instructed 3d gaussian retrieval and coherent editing,” arXiv preprint arXiv:2405.14455, 2024

  145. [153]

    Point’n move: Inter- active scene object manipulation on gaussian splatting radiance fields,

    J. Huang, H. Yu, J. Zhang, and H. Nait-Charif, “Point’n move: Inter- active scene object manipulation on gaussian splatting radiance fields,” IET Image Processing, 2024

  146. [154]

    Neural surface priors for editable gaussian splatting,

    J. Szymkowiak, W. Jakubowska, D. Malarz, W. Smolak-Dy ˙zewska, M. Zi˛ eba, P. Musialski, W. Pałubicki, and P. Spurek, “Neural surface priors for editable gaussian splatting,” arXiv, 2024

  147. [155]

    Gaussianvton: 3d human virtual try-on via multi-stage gaussian splatting editing with image prompting,

    H. Chen, Y . Huang, H. Huang, X. Ge, and D. Shao, “Gaussianvton: 3d human virtual try-on via multi-stage gaussian splatting editing with image prompting,” arXiv preprint arXiv:2405.07472, 2024

  148. [156]

    Sgsst: Scaling gaussian splatting styletransfer,

    B. Galerne, J. Wang, L. Raad, and J.-M. Morel, “Sgsst: Scaling gaussian splatting styletransfer,” in CVPR, 2025

  149. [157]

    Regs: Reference-based controllable scene stylization with gaussian splatting,

    Y . Mei, J. Xu, and V . Patel, “Regs: Reference-based controllable scene stylization with gaussian splatting,” NeurIPS, 2024

  150. [158]

    Wast-3d: Wasserstein-2 distance for scene-to-scene stylization on 3d gaussians,

    D. Kotovenko, O. Grebenkova, N. Sarafianos, A. Paliwal, P. Ma, O. Poursaeed, S. Mohan, Y . Fan, Y . Li, R. Ranjan et al. , “Wast-3d: Wasserstein-2 distance for scene-to-scene stylization on 3d gaussians,” in ECCV, 2024

  151. [159]

    Multi-stylegs: Stylized gaussian splatting with multiple styles,

    Y . Lin, J. Lei, and K. Jia, “Multi-stylegs: Stylized gaussian splatting with multiple styles,” in AAAI, 2025

  152. [160]

    Instantstyle- gaussian: Efficient art style transfer with 3d gaussian splatting,

    X.-Y . Yu, J.-X. Yu, L.-B. Zhou, Y . Wei, and L.-L. Ou, “Instantstyle- gaussian: Efficient art style transfer with 3d gaussian splatting,” arXiv preprint arXiv:2408.04249, 2024

  153. [161]

    Artnvg: Content-style separated artistic neighboring-view gaussian stylization,

    Z. Gu, Z. Zhang, M. Li, Z. Ji, R. Chen, Z. Hu, and G. Ye, “Artnvg: Content-style separated artistic neighboring-view gaussian stylization,” in ICMR, 2025

  154. [162]

    Morpheus: Text-driven 3d gaussian splat shape and color stylization,

    J. Wynn, Z. Qureshi, J. Powierza, J. Watson, and M. Sayed, “Morpheus: Text-driven 3d gaussian splat shape and color stylization,” in CVPR, 2025

  155. [163]

    Fantasystyle: Controllable stylized distillation for 3d gaussian splatting,

    Y . Yang, Y . Wang, C. Wang, H. Wang, and S. He, “Fantasystyle: Controllable stylized distillation for 3d gaussian splatting,”arXiv, 2025

  156. [164]

    Gaussian splatting in style,

    A. Saroha, M. Gladkova, C. Curreli, D. Muhle, T. Yenamandra, and D. Cremers, “Gaussian splatting in style,” arXiv:2403.08498, 2024

  157. [165]

    Stylesplat: 3d object style transfer with gaussian splatting,

    S. Jain, A. Kuthiala, P. S. Sethi, and P. Saxena, “Stylesplat: 3d object style transfer with gaussian splatting,” arXiv:2407.09473, 2024. 18

  158. [166]

    Stylegaussian: Instant 3d style transfer with gaussian splatting,

    K. Liu, F. Zhan, M. Xu, C. Theobalt, L. Shao, and S. Lu, “Stylegaussian: Instant 3d style transfer with gaussian splatting,” in SIGGRAPH Asia , 2024

  159. [167]

    Semanticsplatstylization: Semantic scene stylization based on 3d gaussian splatting and class- based style transfer,

    S. N. Sinha, H. Graf, and M. Weinmann, “Semanticsplatstylization: Semantic scene stylization based on 3d gaussian splatting and class- based style transfer,” in GCH, 2024

  160. [168]

    Language-driven physics- based scene synthesis and editing via feature splatting,

    R.-Z. Qiu, G. Yang, W. Zeng, and X. Wang, “Language-driven physics- based scene synthesis and editing via feature splatting,” inECCV, 2024

  161. [169]

    Mvdrag3d: Drag-based creative 3d editing via multi-view generation-reconstruction priors,

    H. Chen, Y . Lan, Y . Chen, Y . Zhou, and X. Pan, “Mvdrag3d: Drag-based creative 3d editing via multi-view generation-reconstruction priors,” arXiv preprint arXiv:2410.16272, 2024

  162. [170]

    Drag your gaussian: Effective drag-based editing with score distillation for 3d gaussian splatting,

    Y . Qu, D. Chen, X. Li, X. Li, S. Zhang, L. Cao, and R. Ji, “Drag your gaussian: Effective drag-based editing with score distillation for 3d gaussian splatting,” arXiv preprint arXiv:2501.18672, 2025

  163. [171]

    3dego: 3d editing on the go!

    U. Khalid, H. Iqbal, A. Farooq, J. Hua, and C. Chen, “3dego: 3d editing on the go!” in ECCV, 2024

  164. [172]

    Portrait video editing empowered by multimodal generative priors,

    X. Gao, H. Xiao, C. Zhong, S. Hu, Y . Guo, and J. Zhang, “Portrait video editing empowered by multimodal generative priors,” in SIGGRAPH Asia, 2024

  165. [173]

    Infusion: Inpainting 3d gaussians via learning depth completion from diffusion prior,

    Z. Liu, H. Ouyang, Q. Wang, K. L. Cheng, J. Xiao, K. Zhu, N. Xue, Y . Liu, Y . Shen, and Y . Cao, “Infusion: Inpainting 3d gaussians via learning depth completion from diffusion prior,” arXiv preprint arXiv:2404.11613, 2024

  166. [174]

    Reffusion: Reference adapted diffusion models for 3d scene inpainting,

    A. Mirzaei, R. De Lutio, S. W. Kim, D. Acuna, J. Kelly, S. Fidler, I. Gilitschenski, and Z. Gojcic, “Reffusion: Reference adapted diffusion models for 3d scene inpainting,” arXiv preprint arXiv:2404.10765 , 2024

  167. [175]

    Text-to-3d gaussian splatting with physics- grounded motion generation,

    W. Wang and Y . Fu, “Text-to-3d gaussian splatting with physics- grounded motion generation,” arXiv preprint arXiv:2412.05560, 2024

  168. [176]

    Hash3d: Training-free acceleration for 3d generation,

    X. Yang, S. Liu, and X. Wang, “Hash3d: Training-free acceleration for 3d generation,” in CVPR, 2025

  169. [177]

    Layoutdreamer: Physics-guided layout for text-to-3d compositional scene generation,

    Y . Zhou, Z. He, Q. Li, and C. Wang, “Layoutdreamer: Physics-guided layout for text-to-3d compositional scene generation,” arXiv preprint arXiv:2502.01949, 2025

  170. [178]

    Richdreamer: A generalizable normal-depth diffusion model for detail richness in text-to-3d,

    L. Qiu, G. Chen, X. Gu, Q. Zuo, M. Xu, Y . Wu, W. Yuan, Z. Dong, L. Bo, and X. Han, “Richdreamer: A generalizable normal-depth diffusion model for detail richness in text-to-3d,” in CVPR, 2024

  171. [179]

    Text-to-3d using gaussian splatting,

    Z. Chen, F. Wang, Y . Wang, and H. Liu, “Text-to-3d using gaussian splatting,” in CVPR, 2024

  172. [180]

    Gaussiandreamerpro: Text to ma- nipulable 3d gaussians with highly enhanced quality,

    T. Yi, J. Fang, Z. Zhou, J. Wang, G. Wu, L. Xie, X. Zhang, W. Liu, X. Wang, and Q. Tian, “Gaussiandreamerpro: Text to ma- nipulable 3d gaussians with highly enhanced quality,” arXiv preprint arXiv:2406.18462, 2024

  173. [181]

    Compgs: Unleashing 2d compositionality for compositional text-to-3d via dynamically optimizing 3d gaussians,

    C. Ge, C. Xu, Y . Ji, C. Peng, M. Tomizuka, P. Luo, M. Ding, V . Jampani, and W. Zhan, “Compgs: Unleashing 2d compositionality for compositional text-to-3d via dynamically optimizing 3d gaussians,” arXiv preprint arXiv:2410.20723, 2024

  174. [182]

    Cg3d: Compositional generation for text-to-3d via gaussian splatting,

    A. Vilesov, P. Chari, and A. Kadambi, “Cg3d: Compositional generation for text-to-3d via gaussian splatting,” arXiv preprint arXiv:2311.17907, 2023

  175. [183]

    Hyper- 3dg: Text-to-3d gaussian generation via hypergraph,

    D. Di, J. Yang, C. Luo, Z. Xue, W. Chen, X. Yang, and Y . Gao, “Hyper- 3dg: Text-to-3d gaussian generation via hypergraph,” IJCV, 2025

  176. [184]

    Apply hierarchical-chain-of-generation to complex attributes text-to-3d generation,

    Y . Qin, Z. Xu, and Y . Liu, “Apply hierarchical-chain-of-generation to complex attributes text-to-3d generation,” in CVPR, 2025

  177. [185]

    Scalinggaussian: Enhancing 3d content creation with generative gaussian splatting,

    S. Chen, J. Zhou, Z. Jiang, T. Zhang, Z. Wu, J.-N. Hwang, and L. Li, “Scalinggaussian: Enhancing 3d content creation with generative gaussian splatting,” arXiv preprint arXiv:2407.19035, 2024

  178. [186]

    Physics3d: Learning physical properties of 3d gaussians via video diffusion,

    F. Liu, H. Wang, S. Yao, S. Zhang, J. Zhou, and Y . Duan, “Physics3d: Learning physical properties of 3d gaussians via video diffusion,” arXiv preprint arXiv:2406.04338, 2024

  179. [187]

    Stabledreamer: tam- ing noisy score distillation sampling for text-to-3d,

    P. Guo, H. Hao, A. Caccavale, Z. Ren, E. Zhang, Q. Shan, A. Sankar, A. G. Schwing, A. Colburn, and F. Ma, “Stabledreamer: tam- ing noisy score distillation sampling for text-to-3d,” arXiv preprint arXiv:2312.02189, 2023

  180. [188]

    Connecting consistency distillation to score distillation for text-to-3d generation,

    Z. Li, M. Hu, Q. Zheng, and X. Jiang, “Connecting consistency distillation to score distillation for text-to-3d generation,” in ECCV, 2024

  181. [189]

    Dreammapping: High-fidelity text-to-3d generation via variational dis- tribution mapping,

    Z. Cai, D. Wang, Y . Liang, Z. Shao, Y .-C. Chen, X. Zhan, and Z. Wang, “Dreammapping: High-fidelity text-to-3d generation via variational dis- tribution mapping,” arXiv preprint arXiv:2409.05099, 2024

  182. [190]

    Humangaussian: Text-driven 3d human generation with gaussian splat- ting,

    X. Liu, X. Zhan, J. Tang, Y . Shan, G. Zeng, D. Lin, X. Liu, and Z. Liu, “Humangaussian: Text-driven 3d human generation with gaussian splat- ting,” in CVPR, 2024

  183. [191]

    Dreamer xl: Towards high-resolution text-to-3d generation via trajec- tory score matching,

    X. Miao, H. Duan, V . Ojha, J. Song, T. Shah, Y . Long, and R. Ranjan, “Dreamer xl: Towards high-resolution text-to-3d generation via trajec- tory score matching,” arXiv preprint arXiv:2405.11252, 2024

  184. [192]

    Denoising diffusion implicit models,

    J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” arXiv preprint arXiv:2010.02502, 2020

  185. [193]

    Gaussianmotion: End-to-end learning of animatable gaussian avatars with pose guidance from text,

    G. Shim, S. Lee, and J. Choo, “Gaussianmotion: End-to-end learning of animatable gaussian avatars with pose guidance from text,” arXiv preprint arXiv:2502.11642, 2025

  186. [194]

    Enhancing single image to 3d generation using gaussian splatting and hybrid diffusion priors,

    H. Basak, H. Tabatabaee, S. Gayaka, M.-F. Li, X. Yang, C.-H. Kuo, A. Sen, M. Sun, and Z. Yin, “Enhancing single image to 3d generation using gaussian splatting and hybrid diffusion priors,” arXiv preprint arXiv:2410.09467, 2024

  187. [195]

    Geco: Generative image- to-3d within a second,

    C. Wang, J. Gu, X. Long, Y . Liu, and L. Liu, “Geco: Generative image- to-3d within a second,” arXiv preprint arXiv:2405.20327, 2024

  188. [196]

    Prolific- dreamer: High-fidelity and diverse text-to-3d generation with variational score distillation,

    Z. Wang, C. Lu, Y . Wang, F. Bao, C. Li, H. Su, and J. Zhu, “Prolific- dreamer: High-fidelity and diverse text-to-3d generation with variational score distillation,” NeurIPS, 2023

  189. [197]

    Dreamphysics: Learning physical properties of dynamic 3d gaussians with video diffusion priors,

    T. Huang, H. Zhang, Y . Zeng, Z. Zhang, H. Li, W. Zuo, and R. W. Lau, “Dreamphysics: Learning physical properties of dynamic 3d gaussians with video diffusion priors,” arXiv preprint arXiv:2406.01476, 2024

  190. [198]

    Mvgaussian: High- fidelity text-to-3d content generation with multi-view guidance and surface densification,

    P. Pham, A. N. Mathur, O. Sharma, and A. Bera, “Mvgaussian: High- fidelity text-to-3d content generation with multi-view guidance and surface densification,” arXiv preprint arXiv:2409.06620, 2024

  191. [199]

    Gradeadreamer: Enhanced text-to-3d generation using gaussian splatting and multi-view diffusion,

    T. Ukarapol and K. Pruvost, “Gradeadreamer: Enhanced text-to-3d generation using gaussian splatting and multi-view diffusion,” arXiv preprint arXiv:2406.09850, 2024

  192. [200]

    GALA3d: Towards text-to-3d complex scene generation via layout-guided generative gaussian splatting,

    X. Zhou, X. Ran, Y . Xiong, J. He, Z. Lin, Y . Wang, D. Sun, and M.- H. Yang, “GALA3d: Towards text-to-3d complex scene generation via layout-guided generative gaussian splatting,” in ICML, 2024

  193. [201]

    MVDream: Multi- view diffusion for 3d generation,

    Y . Shi, P. Wang, J. Ye, L. Mai, K. Li, and X. Yang, “MVDream: Multi- view diffusion for 3d generation,” in ICLR, 2024

  194. [202]

    Controllable text-to-3d generation via surface-aligned gaussian splatting,

    Z. Li, Y . Chen, L. Zhao, and P. Liu, “Controllable text-to-3d generation via surface-aligned gaussian splatting,” arXiv preprint arXiv:2403.09981, 2024

  195. [203]

    Vfusion3d: Learning scalable 3d generative models from video diffusion models,

    J. Han, F. Kokkinos, and P. Torr, “Vfusion3d: Learning scalable 3d generative models from video diffusion models,” in ECCV, 2024

  196. [204]

    Gvgen: Text-to-3d generation with xxvolu- metric representation,

    X. He, J. Chen, S. Peng, D. Huang, Y . Li, X. Huang, C. Yuan, W. Ouyang, and T. He, “Gvgen: Text-to-3d generation with xxvolu- metric representation,” in ECCV, 2024

  197. [205]

    Taming feed-forward reconstruction models as latent encoders for 3d generative models,

    S. Wizadwongsa, J. Zhou, E. Li, and J. J. Park, “Taming feed-forward reconstruction models as latent encoders for 3d generative models,” arXiv preprint arXiv:2501.00651, 2024

  198. [206]

    Gaussiananything: Interactive point cloud flow matching for 3d generation,

    L. Yushi, S. Zhou, Z. Lyu, F. Hong, S. Yang, B. Dai, X. Pan, and C. C. Loy, “Gaussiananything: Interactive point cloud flow matching for 3d generation,” in ICLR, 2025

  199. [207]

    Atlas gaussians diffusion for 3d generation,

    H. Yang, Y . Dong, H. Jiang, D. Xu, G. Pavlakos, and Q. Huang, “Atlas gaussians diffusion for 3d generation,” in ICLR, 2025

  200. [208]

    Turbo3d: Ultra-fast text-to-3d generation,

    H. Hu, T. Yin, F. Luan, Y . Hu, H. Tan, Z. Xu, S. Bi, S. Tulsiani, and K. Zhang, “Turbo3d: Ultra-fast text-to-3d generation,” in CVPR, 2025

  201. [209]

    AGG: Amortized generative 3d gaussians for single image to 3d,

    D. Xu, Y . Yuan, M. Mardani, S. Liu, J. Song, Z. Wang, and A. Vahdat, “AGG: Amortized generative 3d gaussians for single image to 3d,” TMLR, 2024

  202. [210]

    Humansplat: Generalizable single-image human gaussian splatting with structure priors,

    P. Pan, Z. Su, C. Lin, Z. Fan, Y . Zhang, Z. Li, T. Shen, Y . Mu, and Y . Liu, “Humansplat: Generalizable single-image human gaussian splatting with structure priors,” NeurIPS, 2024

  203. [211]

    Triplane meets gaussian splatting: Fast and generalizable single-view 3d reconstruction with transformers,

    Z.-X. Zou, Z. Yu, Y .-C. Guo, Y . Li, D. Liang, Y .-P. Cao, and S.- H. Zhang, “Triplane meets gaussian splatting: Fast and generalizable single-view 3d reconstruction with transformers,” in CVPR, 2024

  204. [212]

    Flex3d: Feed- forward 3d generation with flexible reconstruction model and input view curation,

    J. Han, J. Wang, A. Vedaldi, P. Torr, and F. Kokkinos, “Flex3d: Feed- forward 3d generation with flexible reconstruction model and input view curation,” arXiv preprint arXiv:2410.00890, 2024

  205. [213]

    Geogs3d: Single-view 3d reconstruction via geometric-aware diffusion model and gaussian splatting,

    Q. Feng, Z. Xing, Z. Wu, and Y .-G. Jiang, “Geogs3d: Single-view 3d reconstruction via geometric-aware diffusion model and gaussian splatting,” arXiv preprint arXiv:2403.10242, 2024

  206. [214]

    Brightdreamer: Generic 3d gaussian gen- erative framework for fast text-to-3d synthesis,

    L. Jiang and L. Wang, “Brightdreamer: Generic 3d gaussian gen- erative framework for fast text-to-3d synthesis,” arXiv preprint arXiv:2403.11273, 2024

  207. [215]

    Large point-to-gaussian model for image-to-3d generation,

    L. Lu, H. Gao, T. Dai, Y . Zha, Z. Hou, J. Wu, and S.-T. Xia, “Large point-to-gaussian model for image-to-3d generation,” in ACM MM , 2024

  208. [216]

    Unigs: Modeling unitary 3d gaussians for novel view synthesis from sparse-view images,

    J. Wu, K. Liu, Y . Shi, X. Jiang, Y . Yao, and L. Zhang, “Unigs: Modeling unitary 3d gaussians for novel view synthesis from sparse-view images,” in ICCV, 2025

  209. [217]

    Nov- elgs: Consistent novel-view denoising via large gaussian reconstruction model,

    J. Liu, J. Xu, W. Cheng, Y . Gao, X. Wang, Y . Shan, and Y . Tang, “Nov- elgs: Consistent novel-view denoising via large gaussian reconstruction model,” arXiv preprint arXiv:2411.16779, 2024

  210. [218]

    Ouroboros3d: Image-to-3d generation via 3d-aware recursive diffusion,

    H. Wen, Z. Huang, Y . Wang, X. Chen, and L. Sheng, “Ouroboros3d: Image-to-3d generation via 3d-aware recursive diffusion,” in CVPR, 2025. 19

  211. [219]

    Cycle3d: High-quality and consistent image-to-3d generation via generation-reconstruction cycle,

    Z. Tang, J. Zhang, X. Cheng, W. Yu, C. Feng, Y . Pang, B. Lin, and L. Yuan, “Cycle3d: High-quality and consistent image-to-3d generation via generation-reconstruction cycle,” in AAAI, 2025

  212. [220]

    Baking gaussian splatting into diffusion denoiser for fast and scalable single-stage image-to-3d generation,

    Y . Cai, H. Zhang, K. Zhang, Y . Liang, M. Ren, F. Luan, Q. Liu, S. Y . Kim, J. Zhang, Z. Zhang et al., “Baking gaussian splatting into diffusion denoiser for fast and scalable single-stage image-to-3d generation,” in ICCV, 2025

  213. [221]

    Hi3d: Pursuing high-resolution image-to-3d generation with video diffusion models,

    H. Yang, Y . Chen, Y . Pan, T. Yao, Z. Chen, C.-W. Ngo, and T. Mei, “Hi3d: Pursuing high-resolution image-to-3d generation with video diffusion models,” in ACM MM, 2024

  214. [222]

    Text-to-3d with classifier score distillation,

    X. Yu, Y .-C. Guo, Y . Li, D. Liang, S.-H. Zhang, and X. QI, “Text-to-3d with classifier score distillation,” in ICLR, 2024

  215. [223]

    Fastscene: Text-driven fast 3d indoor scene generation via panoramic gaussian splatting,

    Y . Ma, D. Zhan, and Z. Jin, “Fastscene: Text-driven fast 3d indoor scene generation via panoramic gaussian splatting,” in IJCAI, 2024

  216. [224]

    Taming video diffusion prior with scene-grounding guidance for 3d gaussian splatting from sparse inputs,

    Y . Zhong, Z. Li, D. Z. Chen, L. Hong, and D. Xu, “Taming video diffusion prior with scene-grounding guidance for 3d gaussian splatting from sparse inputs,” in CVPR, 2025

  217. [225]

    Wonder- world: Interactive 3d scene generation from a single image,

    H.-X. Yu, H. Duan, C. Herrmann, W. T. Freeman, and J. Wu, “Wonder- world: Interactive 3d scene generation from a single image,” in CVPR, 2025

  218. [226]

    Scene4u: Hierarchical layered 3d scene reconstruction from single panoramic image for your immerse exploration,

    Z. Huang, J. He, J. Ye, L. Jiang, W. Li, Y . Chen, and T. Han, “Scene4u: Hierarchical layered 3d scene reconstruction from single panoramic image for your immerse exploration,” in CVPR, 2025

  219. [227]

    Text2immersion: Generative immersive scene with 3d gaussians,

    H. Ouyang, K. Heal, S. Lombardi, and T. Sun, “Text2immersion: Generative immersive scene with 3d gaussians,” arXiv, 2023

  220. [228]

    Realmdreamer: Text-driven 3d scene generation with inpainting and depth diffusion,

    J. Shriram, A. Trevithick, L. Liu, and R. Ramamoorthi, “Realmdreamer: Text-driven 3d scene generation with inpainting and depth diffusion,” 3DV, 2025

  221. [229]

    Wonderjourney: Going from anywhere to everywhere,

    H.-X. Yu, H. Duan, J. Hur, K. Sargent, M. Rubinstein, W. T. Freeman, F. Cole, D. Sun, N. Snavely, J. Wu et al., “Wonderjourney: Going from anywhere to everywhere,” in CVPR, 2024

  222. [230]

    Holodreamer: Holistic 3d panoramic world generation from text descriptions,

    H. Zhou, X. Cheng, W. Yu, Y . Tian, and L. Yuan, “Holodreamer: Holistic 3d panoramic world generation from text descriptions,” arXiv preprint arXiv:2407.15187, 2024

  223. [231]

    Luciddreamer: Domain-free generation of 3d gaussian splatting scenes,

    J. Chung, S. Lee, H. Nam, J. Lee, and K. M. Lee, “Luciddreamer: Domain-free generation of 3d gaussian splatting scenes,” arXiv preprint arXiv:2311.13384, 2023

  224. [232]

    Vistadream: Sampling multiview consistent images for single-view scene reconstruc- tion,

    H. Wang, Y . Liu, Z. Liu, W. Wang, Z. Dong, and B. Yang, “Vistadream: Sampling multiview consistent images for single-view scene reconstruc- tion,” in ICCV, 2025

  225. [233]

    Multi-view geometry-aware diffusion transformer for indoor novel view synthesis,

    X. Kang, Z. Xiang, Z. Zhang, and K. Khoshelham, “Multi-view geometry-aware diffusion transformer for indoor novel view synthesis,” in ICLR, 2025

  226. [234]

    Textsplat: Text-guided semantic fusion for generalizable gaussian splatting,

    Z. Wu, H. Xu, G. Xu, P. Nie, Z. Yan, J. Zheng, L. Qu, M. Li, and L. Nie, “Textsplat: Text-guided semantic fusion for generalizable gaussian splatting,” arXiv preprint arXiv:2504.09588, 2025

  227. [235]

    pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction,

    D. Charatan, S. L. Li, A. Tagliasacchi, and V . Sitzmann, “pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction,” in CVPR, 2024

  228. [236]

    Splatter-360: Generalizable 360 gaussian splatting for wide- baseline panoramic images,

    Z. Chen, C. Wu, Z. Shen, C. Zhao, W. Ye, H. Feng, E. Ding, and S.-H. Zhang, “Splatter-360: Generalizable 360 gaussian splatting for wide- baseline panoramic images,” in CVPR, 2025

  229. [237]

    Mvsplat360: Feed-forward 360 scene synthesis from sparse views,

    Y . Chen, C. Zheng, H. Xu, B. Zhuang, A. Vedaldi, T.-J. Cham, and J. Cai, “Mvsplat360: Feed-forward 360 scene synthesis from sparse views,” in NeurIPS, 2024

  230. [238]

    Generative gaussian splatting for unbounded 3d city generation,

    H. Xie, Z. Chen, F. Hong, and Z. Liu, “Generative gaussian splatting for unbounded 3d city generation,” in CVPR, 2025

  231. [239]

    Selfsplat: Pose-free and 3d prior-free generalizable 3d gaussian splat- ting,

    G. Kang, J. Yoo, J. Park, S. Nam, H. Im, S. Shin, S. Kim, and E. Park, “Selfsplat: Pose-free and 3d prior-free generalizable 3d gaussian splat- ting,” in CVPR, 2025

  232. [240]

    Catsplat: Context-aware transformer with spatial guidance for generalizable 3d gaussian splatting from a single- view image,

    W. Roh, H. Jung, J. W. Kim, S. Lee, I. Yoo, A. Lugmayr, S. Chi, K. Ramani, and S. Kim, “Catsplat: Context-aware transformer with spatial guidance for generalizable 3d gaussian splatting from a single- view image,” in ICCV, 2025

  233. [241]

    Om- nisplat: Taming feed-forward 3d gaussian splatting for omnidirectional images with editable capabilities,

    S. Lee, J. Chung, K. Kim, J. Huh, G. Lee, M. Lee, and K. M. Lee, “Om- nisplat: Taming feed-forward 3d gaussian splatting for omnidirectional images with editable capabilities,” in CVPR, 2025

  234. [242]

    Prometheus: 3d-aware latent diffusion models for feed-forward text-to-3d scene generation,

    Y . Yang, J. Shao, X. Li, Y . Shen, A. Geiger, and Y . Liao, “Prometheus: 3d-aware latent diffusion models for feed-forward text-to-3d scene generation,” in CVPR, 2025

  235. [243]

    Videorf- splat: Direct scene-level text-to-3d gaussian splatting generation with flexible pose and multi-view joint modeling,

    H. Go, B. Park, H. Nam, B.-H. Kim, H. Chung, and C. Kim, “Videorf- splat: Direct scene-level text-to-3d gaussian splatting generation with flexible pose and multi-view joint modeling,” in ICCV, 2025

  236. [244]

    Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis,

    W. Yu, J. Xing, L. Yuan, W. Hu, X. Li, Z. Huang, X. Gao, T.-T. Wong, Y . Shan, and Y . Tian, “Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis,” arXiv preprint arXiv:2409.02048 , 2024

  237. [245]

    Wonderland: Navigating 3d scenes from a single image,

    H. Liang, J. Cao, V . Goel, G. Qian, S. Korolev, D. Terzopoulos, K. N. Plataniotis, S. Tulyakov, and J. Ren, “Wonderland: Navigating 3d scenes from a single image,” in CVPR, 2025

  238. [246]

    Generative gaussian splatting: Generating 3d scenes with video diffusion priors,

    K. Schwarz, N. Mueller, and P. Kontschieder, “Generative gaussian splatting: Generating 3d scenes with video diffusion priors,” in ICCV, 2025

  239. [247]

    Videoscene: Distilling video diffusion model to generate 3d scenes in one step,

    H. Wang, F. Liu, J. Chi, and Y . Duan, “Videoscene: Distilling video diffusion model to generate 3d scenes in one step,” in CVPR, 2025

  240. [248]

    Scene splatter: Momentum 3d scene generation from single image with video diffusion model,

    S. Zhang, J. Li, X. Fei, H. Liu, and Y . Duan, “Scene splatter: Momentum 3d scene generation from single image with video diffusion model,” in CVPR, 2025

  241. [249]

    Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar model- ing,

    Z. Li, Z. Zheng, L. Wang, and Y . Liu, “Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar model- ing,” in CVPR, 2024

  242. [250]

    Hifi4g: High-fidelity human performance rendering via compact gaussian splatting,

    Y . Jiang, Z. Shen, P. Wang, Z. Su, Y . Hong, Y . Zhang, J. Yu, and L. Xu, “Hifi4g: High-fidelity human performance rendering via compact gaussian splatting,” in CVPR, 2024

  243. [251]

    Ggavatar: Reconstructing garment-separated 3d gaussian splatting avatars from monocular video,

    J. Chen, “Ggavatar: Reconstructing garment-separated 3d gaussian splatting avatars from monocular video,” in ACM MM, 2024

  244. [252]

    Taoavatar: Real-time lifelike full-body talking avatars for augmented reality via 3d gaussian splatting,

    J. Chen, J. Hu, G. Wang, Z. Jiang, T. Zhou, Z. Chen, and C. Lv, “Taoavatar: Real-time lifelike full-body talking avatars for augmented reality via 3d gaussian splatting,” in CVPR, 2025

  245. [253]

    Hugs: Human gaussian splats,

    M. Kocabas, J.-H. R. Chang, J. Gabriel, O. Tuzel, and A. Ranjan, “Hugs: Human gaussian splats,” in CVPR, 2024

  246. [254]

    Gaussian shell maps for efficient 3d human generation,

    R. Abdal, W. Yifan, Z. Shi, Y . Xu, R. Po, Z. Kuang, Q. Chen, D.-Y . Yeung, and G. Wetzstein, “Gaussian shell maps for efficient 3d human generation,” in CVPR, 2024

  247. [255]

    Smpl: A skinned multi-person linear model,

    M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, “Smpl: A skinned multi-person linear model,” ACM TOG, 2023

  248. [256]

    Expressive body capture: 3d hands, face, and body from a single image,

    G. Pavlakos, V . Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black, “Expressive body capture: 3d hands, face, and body from a single image,” in CVPR, 2019

  249. [257]

    Drivable 3d gaussian avatars,

    W. Zielonka, T. Bagautdinov, S. Saito, M. Zollhöfer, J. Thies, and J. Romero, “Drivable 3d gaussian avatars,” in 3DV, 2025

  250. [258]

    Gauhuman: Articulated gaussian splatting from monocular human videos,

    S. Hu, T. Hu, and Z. Liu, “Gauhuman: Articulated gaussian splatting from monocular human videos,” in CVPR, 2024

  251. [259]

    Gpavatar: High-fidelity head avatars by learning efficient gaussian projections,

    W.-Q. Feng, D. Han, Z.-K. Zhou, S. Li, X. Liu, P. Wan, D. Zhang, and M. Wang, “Gpavatar: High-fidelity head avatars by learning efficient gaussian projections,” in CVPR, 2025

  252. [260]

    Hravatar: High-quality and relightable gaussian head avatar,

    D. Zhang, Y . Liu, L. Lin, Y . Zhu, K. Chen, M. Qin, Y . Li, and H. Wang, “Hravatar: High-quality and relightable gaussian head avatar,” inCVPR, 2025

  253. [261]

    Strandhead: Text to hair-disentangled 3d head avatars using human-centric priors,

    X. Sun, Z. Cai, Y . Tai, J. Yang, and Z. Zhang, “Strandhead: Text to hair-disentangled 3d head avatars using human-centric priors,” inICCV, 2025

  254. [262]

    Learning a model of facial shape and expression from 4d scans

    T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero, “Learning a model of facial shape and expression from 4d scans.” ACM TOG, 2017

  255. [263]

    Gaussian head avatar: Ultra high-fidelity head avatar via dynamic gaussians,

    Y . Xu, B. Chen, Z. Li, H. Zhang, L. Wang, Z. Zheng, and Y . Liu, “Gaussian head avatar: Ultra high-fidelity head avatar via dynamic gaussians,” in CVPR, 2024

  256. [264]

    Dynagslam: Real-time gaussian-splatting slam for online rendering, tracking, motion predictions of moving objects in dynamic scenes,

    R. B. Li, M. Shaghaghi, K. Suzuki, X. Liu, V . Moparthi, B. Du, W. Curtis, M. Renschler, K. M. B. Lee, N. Atanasovet al., “Dynagslam: Real-time gaussian-splatting slam for online rendering, tracking, motion predictions of moving objects in dynamic scenes,” in ICCV, 2025

  257. [265]

    Outdoor monocular slam with global scale-consistent 3d gaussian pointmaps,

    C. Cheng, S. Yu, Z. Wang, Y . Zhou, and H. Wang, “Outdoor monocular slam with global scale-consistent 3d gaussian pointmaps,” in ICCV, 2025

  258. [266]

    Segs-slam: Structure-enhanced 3d gaussian splatting slam with appearance embedding,

    Y . F. Tianci Wen, Zhiang Liu, “Segs-slam: Structure-enhanced 3d gaussian splatting slam with appearance embedding,” in ICCV, 2025

  259. [267]

    Wildgs-slam: Monocular gaussian splatting slam in dynamic environ- ments,

    J. Zheng, Z. Zhu, V . Bieri, M. Pollefeys, S. Peng, and I. Armeni, “Wildgs-slam: Monocular gaussian splatting slam in dynamic environ- ments,” in CVPR, 2025

  260. [268]

    Splatam: Splat track & map 3d gaussians for dense rgb-d slam,

    N. Keetha, J. Karhade, K. M. Jatavallabhula, G. Yang, S. Scherer, D. Ramanan, and J. Luiten, “Splatam: Splat track & map 3d gaussians for dense rgb-d slam,” in CVPR, 2024

  261. [269]

    Gaussian splatting slam,

    H. Matsuki, R. Murai, P. H. Kelly, and A. J. Davison, “Gaussian splatting slam,” in CVPR, 2024

  262. [270]

    Photo-slam: Real-time simultaneous localization and photorealistic mapping for monocular stereo and rgb-d cameras,

    H. Huang, L. Li, H. Cheng, and S.-K. Yeung, “Photo-slam: Real-time simultaneous localization and photorealistic mapping for monocular stereo and rgb-d cameras,” in CVPR, 2024

  263. [271]

    Sgs- slam: Semantic gaussian splatting for neural dense slam,

    M. Li, S. Liu, H. Zhou, G. Zhu, N. Cheng, T. Deng, and H. Wang, “Sgs- slam: Semantic gaussian splatting for neural dense slam,” in ECCV, 2024

  264. [272]

    Opengs-slam: Open-set dense semantic slam with 3d gaussian splatting for object- level scene understanding,

    D. Yang, Y . Gao, X. Wang, Y . Yue, Y . Yang, and M. Fu, “Opengs-slam: Open-set dense semantic slam with 3d gaussian splatting for object- level scene understanding,” in ICRA, 2025. 20

  265. [273]

    Gs3lam: Gaussian semantic splatting slam,

    L. Li, L. Zhang, Z. Wang, and Y . Shen, “Gs3lam: Gaussian semantic splatting slam,” in ACM MM, 2024

  266. [274]

    Language- embedded gaussian splats (legs): Incrementally building room-scale representations with a mobile robot,

    J. Yu, K. Hari, K. Srinivas, K. El-Refai, A. Rashid, C. M. Kim, J. Kerr, R. Cheng, M. Z. Irshad, A. Balakrishna et al. , “Language- embedded gaussian splats (legs): Incrementally building room-scale representations with a mobile robot,” in IROS, 2024

  267. [275]

    A neural representation framework with llm-driven spatial reasoning for open-vocabulary 3d visual grounding,

    Z. Liu, S. Zheng, S. Chen, C. Zhao, L. Liang, X. Xue, and Y . Fu, “A neural representation framework with llm-driven spatial reasoning for open-vocabulary 3d visual grounding,” in ACM MM, 2025

  268. [276]

    Gaussian-det: Learning closed-surface gaussians for 3d object detection,

    H. Yan, Y . Zheng, and Y . Duan, “Gaussian-det: Learning closed-surface gaussians for 3d object detection,” in ICLR, 2024

  269. [277]

    3dgs-det: Empower 3d gaussian splatting with boundary guidance and box-focused sampling for 3d object detection,

    Y . Cao, Y . Jv, and D. Xu, “3dgs-det: Empower 3d gaussian splatting with boundary guidance and box-focused sampling for 3d object detection,” arXiv preprint arXiv:2410.01647, 2024

  270. [278]

    Matt-gs: Masked attention- based 3dgs for robot perception and object detection,

    J. W. Lee, H. Lim, S. Yang, and J. B. Choi, “Matt-gs: Masked attention- based 3dgs for robot perception and object detection,” in IROS, 2025

  271. [279]

    Nex: Real-time view synthesis with neural basis expansion,

    S. Wizadwongsa, P. Phongthawee, J. Yenphraphai, and S. Suwa- janakorn, “Nex: Real-time view synthesis with neural basis expansion,” in CVPR, 2021

  272. [280]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes,

    A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in CVPR, 2017

  273. [281]

    Mip-nerf 360: Unbounded anti-aliased neural radiance fields,

    J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman, “Mip-nerf 360: Unbounded anti-aliased neural radiance fields,” in CVPR, 2022

  274. [282]

    Objaverse: A universe of annotated 3d objects,

    M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi, “Objaverse: A universe of annotated 3d objects,” in CVPR, 2023

  275. [283]

    Google scanned objects: A high-quality dataset of 3d scanned household items,

    L. Downs, A. Francis, N. Koenig, B. Kinman, R. Hickman, K. Rey- mann, T. B. McHugh, and V . Vanhoucke, “Google scanned objects: A high-quality dataset of 3d scanned household items,” in ICRA, 2022

  276. [284]

    Shapenet: An information- rich 3d model repository,

    A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Suet al., “Shapenet: An information- rich 3d model repository,” arXiv preprint arXiv:1512.03012, 2015

  277. [285]

    The replica dataset: A digital replica of indoor spaces,

    J. Straub, T. Whelan, L. Ma, Y . Chen, E. Wijmans, S. Green, J. J. Engel, R. Mur-Artal, C. Ren, S. Verma et al. , “The replica dataset: A digital replica of indoor spaces,” arXiv preprint arXiv:1906.05797, 2019

  278. [286]

    Scenesplat++: A large dataset and comprehensive benchmark for language gaussian splatting,

    M. Ma, Q. Ma, Y . Li, J. Cheng, R. Yang, B. Ren, N. Popovic, M. Wei, N. Sebe, L. Van Gool et al. , “Scenesplat++: A large dataset and comprehensive benchmark for language gaussian splatting,” arXiv preprint arXiv:2506.08710, 2025

  279. [287]

    Neural volumetric object selection,

    Z. Ren, A. Agarwala, B. Russell, A. G. Schwing, and O. Wang, “Neural volumetric object selection,” in CVPR, 2022

  280. [288]

    Multimodal referring segmentation: A survey,

    H. Ding, S. Tang, S. He, C. Liu, Z. Wu, and Y .-G. Jiang, “Multimodal referring segmentation: A survey,” arXiv preprint arXiv:2508.00265 , 2025

  281. [289]

    GRES: Generalized referring expression segmentation,

    C. Liu, H. Ding, and X. Jiang, “GRES: Generalized referring expression segmentation,” in CVPR, 2023

  282. [290]

    Aligning and prompting everything all at once for universal visual perception,

    Y . Shen, C. Fu, P. Chen, M. Zhang, K. Li, X. Sun, Y . Wu, S. Lin, and R. Ji, “Aligning and prompting everything all at once for universal visual perception,” in CVPR, 2024

  283. [291]

    Objectgs: Object-aware scene reconstruction and scene understanding via gaussian splatting,

    R. Zhu, M. Yu, L. Xu, L. Jiang, Y . Li, T. Zhang, J. Pang, and B. Dai, “Objectgs: Object-aware scene reconstruction and scene understanding via gaussian splatting,” in ICCV, 2025

  284. [292]

    Large scale multi-view stereopsis evaluation,

    R. Jensen, A. Dahl, G. V ogiatzis, E. Tola, and H. Aanæs, “Large scale multi-view stereopsis evaluation,” in CVPR, 2014

  285. [293]

    Tanks and temples: Benchmarking large-scale scene reconstruction,

    A. Knapitsch, J. Park, Q.-Y . Zhou, and V . Koltun, “Tanks and temples: Benchmarking large-scale scene reconstruction,” ACM TOG, 2017

  286. [294]

    Matchable image retrieval by learning from surface reconstruction,

    T. Shen, Z. Luo, L. Zhou, R. Zhang, S. Zhu, T. Fang, and L. Quan, “Matchable image retrieval by learning from surface reconstruction,” in ACCV, 2018

  287. [295]

    Local light field fusion: Practical view synthesis with prescriptive sampling guidelines,

    B. Mildenhall, P. P. Srinivasan, R. Ortiz-Cayon, N. K. Kalantari, R. Ramamoorthi, R. Ng, and A. Kar, “Local light field fusion: Practical view synthesis with prescriptive sampling guidelines,”ACM TOG, 2019

  288. [296]

    Blendedmvs: A large-scale dataset for generalized multi-view stereo networks,

    Y . Yao, Z. Luo, S. Li, J. Zhang, Y . Ren, L. Zhou, T. Fang, and L. Quan, “Blendedmvs: A large-scale dataset for generalized multi-view stereo networks,” in CVPR, 2020

  289. [297]

    Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction,

    J. Reizenstein, R. Shapovalov, P. Henzler, L. Sbordone, P. Labatut, and D. Novotny, “Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction,” in ICCV, 2021

  290. [298]

    Nerfstudio: A modular framework for neural radiance field development,

    M. Tancik, E. Weber, E. Ng, R. Li, B. Yi, T. Wang, A. Kristoffersen, J. Austin, K. Salahi, A. Ahuja et al., “Nerfstudio: A modular framework for neural radiance field development,” in ACM SIGGRAPH, 2023

  291. [299]

    Scannet++: A high- fidelity dataset of 3d indoor scenes,

    C. Yeshwanth, Y .-C. Liu, M. Nießner, and A. Dai, “Scannet++: A high- fidelity dataset of 3d indoor scenes,” in ICCV, 2023

  292. [300]

    Aurafusion360: Augmented unseen region alignment for reference-based 360deg un- bounded scene inpainting,

    C.-H. Wu, Y .-J. Chen, Y .-H. Chen, J.-Y . Lee, B.-H. Ke, C.-W. T. Mu, Y .-C. Huang, C.-Y . Lin, M.-H. Chen, Y .-Y . Linet al., “Aurafusion360: Augmented unseen region alignment for reference-based 360deg un- bounded scene inpainting,” in CVPR, 2025

Pith tools

Reviewed May 18, 2026 · model on record in the stance chip above.