Pith. sign in

REVIEW 4 major objections 8 minor 123 references

AI-powered Contextual 3D Environment Generation: A Systematic Review

T0 review · 4 major / 8 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This systematic review of 136 works argues that diffusion models are the best current architecture for large-scale, unbounded 3D scene generation, but their computational cost and slow training remain the main obstacle to practical…

desk verdict A useful but methodologically sloppy systematic review: the classification tables are valuable, but the PRISMA numbers don't add up and the source policy contradicts itself, so the corpus is not yet convincingly representative. read the letter →

arxiv 2506.05449 v1 pith:AFG23QOJ submitted 2025-06-05 cs.GR cs.CVcs.LG

classification cs.GRcs.CVcs.LG
keywords generativeAI3Dscenegenerationdiffusionmodelstext-to-3Dproceduralcontentvirtualenvironmentsrealismassessmentevaluationmetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish a reliable map of how generative AI is currently used to create 3D environments, based on a structured screening of 5,340 records that left 136 fully read works. Its central finding is a comparative one: diffusion models are the best current fit for large-scale, unbounded scene generation because they capture fine detail and realism, but their computational demands and slow training keep them from real-time use, while GANs and VAEs have complementary weaknesses. The review also argues that cross-attention and latent space alignment are the techniques that make text-to-3D generation effective, and that data diversity plus combined geometric and perceptual metrics are the main levers for future progress. A reader should care because these conclusions pin down where the bottleneck is for automating the creation of 3D worlds for gaming, virtual reality, and cinema.

What carries the argument

The carrying structure is the review's screening and classification pipeline. Four database queries define the search space; eligibility criteria then filter records by publication type, date, full-text access, and an automatic keyword-tagging requirement that a record carry at least one of Generative, GAN, Graphics, Model, AI, Mesh, 3D, Autoencoder, or Attention. The 136 surviving works are read in full and annotated with a fixed questionnaire covering target industry, problem solved, architecture, multi-modal techniques, input, output, experimental method, metrics, datasets, and limitations. The resulting tables mapping architectures, inputs, outputs, metrics, and industries are the evidence base for the comparative claims.

What would settle it

Re-run the four database queries in the same three databases used in the review for the same 2021-2024 window, drop only the automatic keyword-tagging criterion, and screen the unique records by title and abstract; if the additional relevant records materially change the reported shares of diffusion, GAN, and VAE papers, or the input and metric tables, then the review's central characterization is not reliable.

Watch

Extended reading notes

Core claim

The paper's central claim is that the field of AI-driven 3D scene generation has a clear architectural pecking order. Diffusion models, with their ability to capture fine details and realism, are better suited for generating large-scale, unbounded scenes, but their high computational requirements and slower training times remain significant hurdles. GANs produce sharp outputs but struggle to maintain diversity in large or complex scenes, making them less ideal for expansive environments, and VAEs are computationally efficient but limited in detail and scalability. The review further claims that multi-modal integration techniques such as cross-attention and latent space alignment are what enable text-guided 3D generation, and that the quality and diversity of training data, together with combined geometric and perceptual evaluation metrics, are critical for scalable and reliable output. These conclusions come from a structured systematic review that screened 5,340 records, read 136 in full, and classified them by architecture, input, output, dataset, metrics, and application domain.

Load-bearing premise

The load-bearing premise is that the automatic keyword-tagging filter, which keeps only records tagged with at least one of Generative, GAN, Graphics, Model, AI, Mesh, 3D, Autoencoder, or Attention, selects the relevant literature without systematically skewing the corpus; if it drops relevant work, the review's distributional conclusions are not representative.

Editorial extensions

If this is right

  • Future 3D generation research should concentrate on cutting the training and inference cost of diffusion models rather than replacing the architecture, since the review identifies them as the best fit for unbounded scenes.
  • Text-to-3D systems should build on cross-attention and latent space alignment techniques, which the review singles out as the effective bridges between text prompts and 3D outputs.
  • The scarcity of diverse, high-quality datasets becomes a first-order obstacle: improving dataset breadth and quality should directly improve generalization and realism.
  • Evaluation practice should combine geometric metrics like IoU and Chamfer Distance with perceptual metrics like FID and LPIPS, since no single metric captures realism.
  • For large-environment applications, practitioners should expect GANs to risk mode collapse and VAEs to cap out in detail, and plan around those constraints.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's claims, the exclusion of preprints and non-peer-reviewed venues likely undercounts the fastest-moving text-to-3D work, so the reported diffusion dominance may be understated.
  • Beyond the paper's claims, the keyword-tagging filter could systematically favor geometry-heavy papers and underrepresent layout-focused scene generation, a bias the paper itself notes in its exclusion section.
  • Beyond the paper's claims, applying the same screening to 2025 onward records could test whether Gaussian splatting, which appears only once in the architecture table, becomes a serious competitor to diffusion as its cost profile improves.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 8 minor

Summary. This manuscript conducts a PRISMA-based systematic review of AI-driven 3D environment generation. The authors screened 5,340 records from IEEE, ACM, and Scopus, retaining 136 papers, and classify them according to tasks, architectures, input/output types, multi-modal techniques, metrics, and application domains. The review's main findings are that diffusion models are better suited for large-scale, unbounded scene generation than GANs or VAEs (though computationally expensive), that cross-attention and latent-space alignment are effective for text-to-3D integration, and that data quality/diversity plus multi-metric evaluation are crucial for robust 3D scene generation. The paper provides numerous classification tables (Tables III-X) and a thematic analysis of the selected literature.

Significance. If the screening process were reliable, this review would offer a useful map of a fast-moving field, and its taxonomy of architectures, inputs, outputs, and metrics could help researchers position new work. The classification tables are extensive, and the discussion of evaluation metrics in Section III-G is a practical contribution. However, the review's descriptive statistics and the central comparative conclusion about diffusion models rest on the representativeness of the 136-paper corpus, which is not established due to screening inconsistencies and a partially unvalidated automatic filter. The manuscript also does not include reproducible code or machine-checked derivations; its value is as a literature synthesis. The methodological issues described below are load-bearing because they directly affect the empirical claims about the distribution of architectures and the conclusions drawn from that distribution.

major comments (4)
  1. [Section II.C] The PRISMA screening counts are internally inconsistent. The text states that 5,340 records were identified and 2,984 duplicates removed, leaving 2,355 records, but 5,340 - 2,984 = 2,356. It then states that 1,811 records were excluded at screening, leaving 544, and that 408 additional records were excluded after reading titles/abstracts, leaving 136. This is arithmetically consistent for the 2,355 starting point (1,811 + 408 + 136 = 2,355), but Section II.C.2 reports 556 manual exclusions and 1,663 automatic exclusions, which sum to 2,219, not 1,811. The relationship between these two sets of numbers is never explained, and the flow in Figure 1 is not legible in the manuscript. These inconsistencies make the screening process non-reproducible and undermine the claim of PRISMA compliance.
  2. [Section II.A and Section II.C, criterion 8] The methodology excludes arXiv and preprints, but the manuscript itself relies on preprints. Section II.A explicitly excludes arXiv because of preprint heterogeneity, and eligibility criterion 8 excludes any 'work-in-progress, a pre-print, or any other type of document different from a research article or survey.' Yet the introduction cites arXiv preprints [4]-[6] (Text2Room, RealmDreamer, 3D-LLM) and the included work [8] (Edify3D) is an arXiv preprint admitted as one of the 136 reviewed records. Since a substantial share of 2021-2024 text-to-3D and unbounded-scene-generation research appears on arXiv, this contradiction may systematically bias the corpus toward ACM/IEEE/Scopus-indexed publications. The authors should either justify the inclusion of [8] under the stated criteria or revise the eligibility rules, and they should quantify the effect of this choice on the architectural and input distributions in Tables IV-VII.
  3. [Section II.C.1] The automatic keyword-tagging filter is a load-bearing inclusion criterion whose behavior is undocumented. Records must be automatically tagged with at least one of 'Generative, GAN, Graphics, Model, AI, Mesh, 3D, Autoencoder, Attention' in Zotero. The paper does not report how many records were removed by this filter, nor does it validate that the filter does not silently drop relevant records whose metadata uses different vocabulary (e.g., 'neural radiance field', 'NeRF', 'radiance field', 'neural rendering'). Because the filter operates on metadata, records with sparse or nonstandard tags could be excluded before manual screening, biasing the corpus. At minimum, the authors should report the number of records excluded by this criterion and discuss its sensitivity with respect to the final corpus composition.
  4. [Section III-D] The abstract and Section III-D present the comparative claim that 'Diffusion models, with their ability to capture fine details and realism, are better suited for generating large-scale, unbounded scenes' than GANs or VAEs. This is stated as a settled conclusion of the review, but the evidence supplied is narrative and piecemeal: the text cites [48] for GAN diversity limits, [113] for VAE limitations, and [112] for diffusion's computational cost, without a systematic comparison of architectures across scene scales within the reviewed corpus. Given the unresolved representativeness issues in the screening process (previous comments), this conclusion overreaches the evidence presented. The authors should either qualify the claim as a synthesis of the surveyed authors' reported limitations or provide a structured comparison that controls for the uneven representation of architectures and application domains in Tables IV and V.
minor comments (8)
  1. [Section II.C.1] The sentence 'eligibility criteriaRecords must be automatically tagged with the desired keywords for the topic' is missing punctuation; it should read 'eligibility criteria: records must be automatically tagged...'
  2. [Section III-B and Table IV] 'V AE' should be 'VAE' (also in Table IV).
  3. [Section III-G] The metric names 'Frechet Point Distance' and 'Frechet Inception Distance' should use the accented form 'Fréchet' for consistency with standard terminology.
  4. [Section III-H] The phrase 'enabling prompt-based bibliographies [9]' is unclear; it likely refers to generating design references or variations from prompts, and should be rephrased.
  5. [Section II.A] 'arXiv.org' is written with an inconsistent space ('arXiv.org' and 'arXiv.org'); please standardize the spelling.
  6. [References] The reference list includes preprints [4]-[6] and [8] despite the stated exclusion of preprints; please reconcile the reference policy with the eligibility criteria.
  7. [Global] The manuscript includes the line 'This work has been submitted to the IEEE for possible publication' and a reference to 'this thesis' in Section III-B; these should be removed or clarified before submission to a journal.
  8. [Figure 1] The PRISMA flowchart should display numbers that match the corrected screening counts; the current text implies 2,356 unique records, while the reported duplicates removal leads to 2,355.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the review's claims are aggregations of the 136 cited papers, with no fitted inputs, equations, or author self-citations that reduce the conclusions to their inputs.

full rationale

This paper is a PRISMA systematic review rather than a derivation or modeling paper. Its central comparative claim (Section III-D) that diffusion models are better suited for large-scale, unbounded scenes is presented as a synthesis of the reviewed literature, supported by citations such as [112]; it is not derived from any equation, fit, or parameter estimated in the paper. The architecture, output, input, and metric tables (IV-VIII) are classifications built from the included papers' own descriptions, which is standard review practice and does not constitute definitional circularity. The reference list contains no citations to the authors' own prior work, so there is no self-citation chain. The weaknesses identified by a skeptical reader — the exclusion of arXiv (Section II-A), the undocumented Zotero tag filter (Section II-C.1), the inconsistent PRISMA arithmetic (5340 - 2984 = 2356, not 2355), and the inclusion of the arXiv preprint Edify3D [8] despite criterion 8 — are threats to corpus representativeness and internal consistency, but they do not make any conclusion equivalent to its input by construction. A biased or incomplete corpus can undermine the review's conclusions, but that is a validity concern, not a circularity concern. Therefore the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The review's conclusions rest on assumptions about corpus completeness and classification reliability: the three chosen databases plus one website record represent the relevant literature; the automatic keyword tagging filter identifies relevant records without bias; and the manual reading correctly classifies each paper by task, architecture, input, output, and metric. These are domain assumptions of the review, not independently verified.

assumptions (3)
  • domain assumption The three databases (IEEE, ACM, Scopus) plus one additional website record cover the relevant peer-reviewed literature on AI-driven 3D scene generation from 2021 onward.
    Section II-A defines the data sources; the review excludes Google Scholar and arXiv, so its comprehensiveness depends on this coverage being sufficient.
  • ad hoc to paper Zotero's automatic keyword tagging with the required terms (Generative, GAN, Graphics, Model, AI, Mesh, 3D, Autoencoder, Attention) identifies relevant records without systematically biasing the corpus.
    Section II-C 'Challenges when Excluding Records' describes this filter; it is an ad hoc inclusion criterion chosen by the authors and applied through a specific tool.
  • domain assumption Manual title/abstract and full-text reading yields correct classification of each record into the task, architecture, output, input, metric, and industry tables.
    Section II-D 'Annotation' describes a questionnaire but no inter-rater reliability or independent validation; the tables and conclusions rely on the accuracy of these classifications.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI-powered Contextual 3D Environment Generation: A Systematic Review." pith.science (2026). https://pith.science/paper/AFG23QOJ

@misc{pith2026250605449,
  author       = {Pith},
  title        = {Pith review of: AI-powered Contextual 3D Environment Generation: A Systematic Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AFG23QOJ}},
  note         = {Machine review of arXiv:2506.05449}
}
read the original abstract

The generation of high-quality 3D environments is crucial for industries such as gaming, virtual reality, and cinema, yet remains resource-intensive due to the reliance on manual processes. This study performs a systematic review of existing generative AI techniques for 3D scene generation, analyzing their characteristics, strengths, limitations, and potential for improvement. By examining state-of-the-art approaches, it presents key challenges such as scene authenticity and the influence of textual inputs. Special attention is given to how AI can blend different stylistic domains while maintaining coherence, the impact of training data on output quality, and the limitations of current models. In addition, this review surveys existing evaluation metrics for assessing realism and explores how industry professionals incorporate AI into their workflows. The findings of this study aim to provide a comprehensive understanding of the current landscape and serve as a foundation for future research on AI-driven 3D content generation. Key findings include that advanced generative architectures enable high-quality 3D content creation at a high computational cost, effective multi-modal integration techniques like cross-attention and latent space alignment facilitate text-to-3D tasks, and the quality and diversity of training data combined with comprehensive evaluation metrics are critical to achieving scalable, robust 3D scene generation.

Figures

Figures reproduced from arXiv: 2506.05449 by the authors.

Figure 1
Figure 1. PRISMA Flowchart exploratory work, which in this case resulted in the review of one additional record [8]. B. Search Queries The search queries are constructed iteratively, with a strong emphasis on refining and expanding the scope of results to comprehensively capture the applications of Gen AI in 3D synthesis. However, records on this topic for application domains such as gaming, cinema, and robotics were also tar… view at source ↗
Figure 2
Figure 2. Distribution of Records before Screening [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Distribution of Records After Screening even though those were not the main focus of research for such articles. One solution that greatly improved the quality of the included records for review was the addition of eligibility criteria Records must be automatically tagged with the desired keywords for the topic. Using the reference manager’s automatic record tagging feature, a few keywords were selected as crucial a… view at source ↗
Figures from the paper (7 more)
Figure 5
Figure 5. Figure 5: BlockFusion generates new blocks given 2D Conditional Layout. [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 6
Figure 6. Figure 6: Albedo Aware 3D Content Generation by HyperDreamer [ [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 7
Figure 7. Figure 7: NeRF-IS [10]: Encoding Semantic Information in Diffusion Process. SceneHGN: Hierarchical Graph Networks for 3D Indoor Scene Generation With Fine-Grained Geometry [22] utilizes multi-stage training strategies to produce stable and detailed indoor scenes, pushing the bou…
Figure 8
Figure 8. Figure 8: Usage of Cross-Attention in Spice-E [89] Architecture. Diffusion models incorporating innovative approaches to noise and optimization mechanisms have also emerged as pivotal tools. Blue noise for diffusion models [32] explores novel noise signals, such as blue Noise, w…
Figure 10
Figure 10. Figure 10: 3D Scene Generation Pipeline from CLAY [ [PITH_FULL_IMAGE:figures/full_fig_p007_10.png]
Figure 9
Figure 9. Figure 9: SPAGHETTI [12] architecture enhances Control over Shape Parts. D. Limitations in lifelike 3D scene generation The limitations of current AI models in the generation of lifelike 3D scenes stem from several critical factors, one of them being the immense processing power…
Figure 11
Figure 11. Figure 11: Shaping Autoencoding Pipeline in 3DShape2VecSet [ [PITH_FULL_IMAGE:figures/full_fig_p008_11.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

123 extracted references · 31 canonical work pages

  1. [8]

    NVIDIA et al.Edify 3D: Scalable High-Quality 3D Asset Generation. 2024. arXiv: 2411.07135[cs.CV]. URL: https://arxiv.org/abs/2411.07135

  2. [4]

    Yining Hong et al.3D-LLM: Injecting the 3D World into Large Language Models. 2023. arXiv: 2307.12981 [cs.CV].URL: https://arxiv.org/abs/2307.12981

  3. [6]

    Jaidev Shriram et al.RealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth Diffu- sion. 2024. arXiv: 2404.07199[cs.CV].URL: https: //arxiv.org/abs/2404.07199

  4. [48]

    SinGRAF: Learning a 3D Gen- erative Radiance Field for a Single Scene

    Minjung Son et al. “SinGRAF: Learning a 3D Gen- erative Radiance Field for a Single Scene”. In:2023 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR). ISSN: 2575-7075. June 2023, pp. 8507–8517.DOI: 10 . 1109 / CVPR52729 . 2023.00822

  5. [113]

    VDAM: V AE based domain adap- tation for cloud property retrieval from multi-satellite data

    Xin Huang et al. “VDAM: V AE based domain adap- tation for cloud property retrieval from multi-satellite data”. In:Proceedings of the 30th International Con- ference on Advances in Geographic Information Sys- tems. SIGSPATIAL ’22. event-place: Seattle, Washing- ton. New York, NY , USA: Association for Computing Machinery, 2022.ISBN: 978-1-4503-9529-8.DOI:...

  6. [112]

    ThemeStation: Generating Theme-Aware 3D Assets from Few Exemplars

    Zhenwei Wang et al. “ThemeStation: Generating Theme-Aware 3D Assets from Few Exemplars”. In: ACM SIGGRAPH 2024 Conference Papers. SIG- GRAPH ’24. event-place: Denver, CO, USA. New York, NY , USA: Association for Computing Machin- ery, 2024.ISBN: 9798400705250.DOI: 10 . 1145 / 3641519 . 3657471.URL: https : / / doi . org / 10 . 1145 / 3641519.3657471

  7. [1]

    A survey on procedural modelling for virtual worlds

    R. M. Smelik et al. “A survey on procedural modelling for virtual worlds”. In:Computer Graphics Forum33.6 (2014), pp. 31–50

  8. [2]

    Deep learning- based 3D reconstruction: a survey

    Taha Samavati and Mohsen Soryani. “Deep learning- based 3D reconstruction: a survey”. In:Artificial Intel- ligence Review56 (Jan. 2023).DOI: 10.1007/s10462- 023-10399-2

Show all 123 references
  1. [3]

    Procedural content generation via machine learning (PCGML)

    A. Summerville et al. “Procedural content generation via machine learning (PCGML)”. In:IEEE Transac- tions on Games10.3 (2018), pp. 257–270

  2. [5]

    Lukas H ¨ollein et al.Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models. 2023. arXiv: 2303.11989[cs.CV].URL: https://arxiv.org/ abs/2303.11989

  3. [7]

    The PRISMA 2020 statement: an updated guideline for reporting systematic reviews

    Matthew J Page et al. “The PRISMA 2020 statement: an updated guideline for reporting systematic reviews”. In:BMJ372 (2021).DOI: 10.1136/bmj.n71. eprint: https://www.bmj.com/content/372/bmj.n71.full.pdf. URL: https://www.bmj.com/content/372/bmj.n71

  4. [10]

    NeRF-IS: Explicit Neural Radi- ance Fields in Semantic Space

    Jiansong Sha et al. “NeRF-IS: Explicit Neural Radi- ance Fields in Semantic Space”. In:Proceedings of the 5th ACM International Conference on Multimedia in Asia. MMAsia ’23. event-place: Tainan, Taiwan. New York, NY , USA: Association for Computing Machinery, 2024.ISBN: 979840...

  5. [11]

    V AIDE: Virtual AI Designer for Web3D Exhibition Layout Creation

    Bixiao Zhao et al. “V AIDE: Virtual AI Designer for Web3D Exhibition Layout Creation”. In:2024 IEEE International Conference on Web Services (ICWS). ISSN: 2836-3868. July 2024, pp. 1314–1320.DOI: 10. 1109/ICWS62655.2024.00158

  6. [12]

    SPAGHETTI: editing implicit shapes through part aware generation

    Amir Hertz et al. “SPAGHETTI: editing implicit shapes through part aware generation”. In:ACM Trans. Graph.41.4 (July 2022). Place: New York, NY , USA Publisher: Association for Computing Ma- chinery.ISSN: 0730-0301.DOI: 10 . 1145 / 3528223 . 3530084.URL: https : / / doi . org ...

  7. [13]

    Generative Terrain Authoring with Mid-air Hand Sketching in Virtual Reality

    Yushen Hu et al. “Generative Terrain Authoring with Mid-air Hand Sketching in Virtual Reality”. In:Pro- ceedings of the 30th ACM Symposium on Virtual Real- ity Software and Technology. VRST ’24. event-place: Trier, Germany. New York, NY , USA: Association for Computing Machine...

  8. [14]

    SP-GAN: sphere-guided 3D shape generation and manipulation

    Ruihui Li et al. “SP-GAN: sphere-guided 3D shape generation and manipulation”. In:ACM Trans. Graph. 40.4 (July 2021). Place: New York, NY , USA Pub- lisher: Association for Computing Machinery.ISSN: 0730-0301.DOI: 10 . 1145 / 3450626 . 3459766.URL: https://doi.org/10.1145/3450...

  9. [16]

    Creating and Experiencin 3D Im- mersion Using Generative 2D Diffusion: An Integrated Framework

    Ziming He et al. “Creating and Experiencin 3D Im- mersion Using Generative 2D Diffusion: An Integrated Framework”. In:2024 IEEE International Confer- ence on Multimedia and Expo Workshops (ICMEW). ISSN: 2995-1429. July 2024, pp. 1–6.DOI: 10.1109/ ICMEW63481.2024.10645466

  10. [17]

    LLMR: Real-time Prompting of Interactive Worlds using Large Language Models

    Fernanda De La Torre et al. “LLMR: Real-time Prompting of Interactive Worlds using Large Language Models”. In:Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. CHI ’24. event-place: Honolulu, HI, USA. New York, NY , USA: Association for Computing Ma...

  11. [18]

    NeuralField-LDM: Scene Generation with Hierarchical Latent Diffusion Mod- els

    Seung Wook Kim et al. “NeuralField-LDM: Scene Generation with Hierarchical Latent Diffusion Mod- els”. In:2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). ISSN: 2575-

  12. [22]

    SceneHGN: Hierarchical Graph Net- works for 3D Indoor Scene Generation With Fine- Grained Geometry

    Lin Gao et al. “SceneHGN: Hierarchical Graph Net- works for 3D Indoor Scene Generation With Fine- Grained Geometry”. In:IEEE Transactions on Pattern Analysis and Machine Intelligence45.7 (July 2023), pp. 8902–8919.ISSN: 1939-3539.DOI: 10 . 1109 / TPAMI.2023.3237577

  13. [24]

    iControl3D: An Interactive System for Controllable 3D Scene Generation

    Xingyi Li et al. “iControl3D: An Interactive System for Controllable 3D Scene Generation”. In:Proceed- ings of the 32nd ACM International Conference on Multimedia. MM ’24. event-place: Melbourne VIC, Australia. New York, NY , USA: Association for Com- puting Machinery, 2024, p...

  14. [25]

    MemoVis: A GenAI-Powered Tool for Creating Companion Reference Images for 3D Design Feedback

    Chen Chen et al. “MemoVis: A GenAI-Powered Tool for Creating Companion Reference Images for 3D Design Feedback”. In:ACM Trans. Comput.-Hum. In- teract.(Sept. 2024). Place: New York, NY , USA Pub- lisher: Association for Computing Machinery.ISSN: 1073-0516.DOI: 10.1145/3694681....

  15. [26]

    ProteusNeRF: Fast Lightweight NeRF Editing using 3D-Aware Image Context

    Binglun Wang, Niladri Shekhar Dutt, and Niloy J. Mitra. “ProteusNeRF: Fast Lightweight NeRF Editing using 3D-Aware Image Context”. English. In:Pro- ceedings of the ACM on Computer Graphics and Interactive Techniques7.1 (2024). Publisher: Associ- ation for Computing Machinery T...

  16. [28]

    A Multi-Stage Advanced Deep Learning Graphics Pipeline

    Mark Wesley Harris and Sudhanshu Kumar Semwal. “A Multi-Stage Advanced Deep Learning Graphics Pipeline”. In:SIGGRAPH Asia 2021 Technical Com- munications. SA ’21. event-place: Tokyo, Japan. New York, NY , USA: Association for Computing Machin- ery, 2021.ISBN: 978-1-4503-9073-6...

  17. [30]

    A Survey on Generative Adversarial Networks: Variants, Appli- cations, and Training

    Abdul Jabbar, Xi Li, and Bourahla Omar. “A Survey on Generative Adversarial Networks: Variants, Appli- cations, and Training”. In:ACM Comput. Surv.54.8 (Oct. 2021). Place: New York, NY , USA Publisher: Association for Computing Machinery.ISSN: 0360- 0300.DOI: 10.1145/3463475.U...

  18. [33]

    A Neural Space-Time Represen- tation for Text-to-Image Personalization

    Yuval Alaluf et al. “A Neural Space-Time Represen- tation for Text-to-Image Personalization”. In:ACM Trans. Graph.42.6 (Dec. 2023). Place: New York, NY , USA Publisher: Association for Computing Machin- ery.ISSN: 0730-0301.DOI: 10.1145/3618322.URL: https://doi.org/10.1145/3618322

  19. [34]

    Artin- ter: AI-powered Boundary Objects for Commission- ing Visual Arts

    John Joon Young Chung and Eytan Adar. “Artin- ter: AI-powered Boundary Objects for Commission- ing Visual Arts”. In:Proceedings of the 2023 ACM Designing Interactive Systems Conference. DIS ’23. event-place: Pittsburgh, PA, USA. New York, NY , USA: Association for Computing Ma...

  20. [35]

    Autoencoder-Based Collaborative Attention GAN for Multi-Modal Image Synthesis

    Bing Cao et al. “Autoencoder-Based Collaborative Attention GAN for Multi-Modal Image Synthesis”. In: IEEE Transactions on Multimedia26 (2024), pp. 995– 13 1010.ISSN: 1941-0077.DOI: 10 . 1109 / TMM . 2023 . 3274990

  21. [36]

    FIND: Fine-tuning Initial Noise Distribution with Policy Optimization for Diffusion Models

    Changgu Chen et al. “FIND: Fine-tuning Initial Noise Distribution with Policy Optimization for Diffusion Models”. In:Proceedings of the 32nd ACM Inter- national Conference on Multimedia. MM ’24. event- place: Melbourne VIC, Australia. New York, NY , USA: Association for Comput...

  22. [37]

    ConceptLab: Creative Concept Generation using VLM-Guided Diffusion Prior Con- straints

    Elad Richardson et al. “ConceptLab: Creative Concept Generation using VLM-Guided Diffusion Prior Con- straints”. In:ACM Trans. Graph.43.3 (June 2024). Place: New York, NY , USA Publisher: Association for Computing Machinery.ISSN: 0730-0301.DOI: 10. 1145/3659578.URL: https://do...

  23. [38]

    GeoLatent: A Geometric Approach to Latent Space Design for Deformable Shape Gen- erators

    Haitao Yang et al. “GeoLatent: A Geometric Approach to Latent Space Design for Deformable Shape Gen- erators”. In:ACM Trans. Graph.42.6 (Dec. 2023). Place: New York, NY , USA Publisher: Association for Computing Machinery.ISSN: 0730-0301.DOI: 10. 1145/3618371.URL: https://doi....

  24. [39]

    Sat2Scene: 3D Urban Scene Gen- eration from Satellite Images with Diffusion

    Zuoyue Li et al. “Sat2Scene: 3D Urban Scene Gen- eration from Satellite Images with Diffusion”. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). ISSN: 2575-7075. June 2024, pp. 7141–7150.DOI: 10 . 1109 / CVPR52733 . 2024.00682

  25. [40]

    Synthesizing Indoor Scene Layouts in Complicated Architecture Using Dynamic Convolution Networks

    Hao Jiang et al. “Synthesizing Indoor Scene Layouts in Complicated Architecture Using Dynamic Convolution Networks”. In:Proc. ACM Comput. Graph. Interact. Tech.4.1 (Apr. 2021). Place: New York, NY , USA Publisher: Association for Computing Machinery.DOI: 10 . 1145 / 3451267.UR...

  26. [41]

    DG3D: Generating High Quality 3D Textured Shapes by Learning to Discriminate Multi- Modal Diffusion-Renderings

    Qi Zuo et al. “DG3D: Generating High Quality 3D Textured Shapes by Learning to Discriminate Multi- Modal Diffusion-Renderings”. In:2023 IEEE/CVF In- ternational Conference on Computer Vision (ICCV). ISSN: 2380-7504. Oct. 2023, pp. 14529–14538.DOI: 10.1109/ICCV51070.2023.01340

  27. [42]

    3DP3: 3D Scene Perception via Probabilistic Programming

    Nishad Gothoskar et al. “3DP3: 3D Scene Perception via Probabilistic Programming”. English. In:Advances in Neural Information Processing Systems. Ed. by Ranzato M et al. V ol. 12. ISSN: 10495258 Type: Con- ference paper. Neural information processing systems foundation, 2021, ...

  28. [43]

    Sparse Query Dense: Enhanc- ing 3D Object Detection with Pseudo Points

    Yujian Mo et al. “Sparse Query Dense: Enhanc- ing 3D Object Detection with Pseudo Points”. In: Proceedings of the 32nd ACM International Confer- ence on Multimedia. MM ’24. event-place: Melbourne VIC, Australia. New York, NY , USA: Association for Computing Machinery, 2024, pp...

  29. [46]

    rlty2rlty: Transitioning Between Realities with Gen- erative AI

    Matt Gottsacker, Gerd Bruder, and Gregory F. Welch. “rlty2rlty: Transitioning Between Realities with Gen- erative AI”. In:2024 IEEE Conference on Virtual Re- ality and 3D User Interfaces Abstracts and Workshops (VRW). Mar. 2024, pp. 1160–1161.DOI: 10 . 1109 / VRW62533.2024.00374

  30. [49]

    RIP-NeRF: Learning Rotation- Invariant Point-based Neural Radiance Field for Fine- grained Editing and Compositing

    Yuze Wang et al. “RIP-NeRF: Learning Rotation- Invariant Point-based Neural Radiance Field for Fine- grained Editing and Compositing”. In:Proceedings of the 2023 ACM International Conference on Mul- timedia Retrieval. ICMR ’23. event-place: Thessa- loniki, Greece. New York, NY...

  31. [50]

    In- teractive Latent Variable Evolution for the Generation of Minecraft Structures

    Timothy Merino, M. Charity, and Julian Togelius. “In- teractive Latent Variable Evolution for the Generation of Minecraft Structures”. English. In:ACM Interna- tional Conference Proceeding Series. Ed. by Lopes P et al. Type: Conference paper. Association for Comput- ing Machin...

  32. [51]

    World-GAN: a Generative Model for Minecraft Worlds

    Maren Awiszus, Frederik Schubert, and Bodo Rosen- hahn. “World-GAN: a Generative Model for Minecraft Worlds”. In:2021 IEEE Conference on Games (CoG). ISSN: 2325-4289. Aug. 2021, pp. 1–8.DOI: 10.1109/ CoG52621.2021.9619133

  33. [53]

    CLIP-Mesh: Generat- ing textured meshes from text using pretrained image- text models

    Nasir Mohammad Khalid et al. “CLIP-Mesh: Generat- ing textured meshes from text using pretrained image- text models”. In:SIGGRAPH Asia 2022 Conference Papers. SA ’22. event-place: Daegu, Republic of Ko- rea. New York, NY , USA: Association for Computing Machinery, 2022.ISBN: 9...

  34. [54]

    VRCopilot: Authoring 3D Layouts with Generative AI Models in VR

    Lei Zhang et al. “VRCopilot: Authoring 3D Layouts with Generative AI Models in VR”. In:Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. UIST ’24. event-place: Pitts- burgh, PA, USA. New York, NY , USA: Association for Computing Machinery,...

  35. [55]

    WorldSmith: Iterative and Expressive Prompting for World Building with a Generative AI

    Hai Dang et al. “WorldSmith: Iterative and Expressive Prompting for World Building with a Generative AI”. In:Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. UIST ’23. event-place: San Francisco, CA, USA. New York, NY , USA: Association f...

  36. [56]

    Generating and Integrating Dif- fusion Model-Based Panoramic Views for Virtual In- terview Platform

    Jongwook Si et al. “Generating and Integrating Dif- fusion Model-Based Panoramic Views for Virtual In- terview Platform”. In:2024 IEEE International Con- ference on Artificial Intelligence in Engineering and Technology (IICAIET). Aug. 2024, pp. 343–348.DOI: 10.1109/IICAIET6235...

  37. [58]

    BlockFusion: Expandable 3D Scene Generation using Latent Tri-plane Extrapola- tion

    Zhennan Wu et al. “BlockFusion: Expandable 3D Scene Generation using Latent Tri-plane Extrapola- tion”. In:ACM Trans. Graph.43.4 (July 2024). Place: New York, NY , USA Publisher: Association for Com- puting Machinery.ISSN: 0730-0301.DOI: 10 . 1145 / 3658188.URL: https://doi.or...

  38. [59]

    L-MAGIC: Language Model As- sisted Generation of Images with Coherence

    Zhipeng Cai et al. “L-MAGIC: Language Model As- sisted Generation of Images with Coherence”. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). ISSN: 2575-7075. June 2024, pp. 7049–7058.DOI: 10 . 1109 / CVPR52733 . 2024.00673

  39. [60]

    iNVS: Repurposing Diffusion In- painters for Novel View Synthesis

    Yash Kant et al. “iNVS: Repurposing Diffusion In- painters for Novel View Synthesis”. In:SIGGRAPH Asia 2023 Conference Papers. SA ’23. event-place: Sydney, NSW, Australia. New York, NY , USA: As- sociation for Computing Machinery, 2023.ISBN: 9798400703157.DOI: 10 . 1145 / 3610...

  40. [61]

    WorldGen: A Large Scale Generative Simulator

    Chahat Deep Singh et al. “WorldGen: A Large Scale Generative Simulator”. In:2023 IEEE International Conference on Robotics and Automation (ICRA). May 2023, pp. 9147–9154.DOI: 10.1109/ICRA48891.2023. 10160861

  41. [62]

    ShapeCoder: Discovering Abstractions for Visual Programs from Unstructured Primitives

    R. Kenny Jones et al. “ShapeCoder: Discovering Abstractions for Visual Programs from Unstructured Primitives”. In:ACM Trans. Graph.42.4 (July 2023). Place: New York, NY , USA Publisher: Association for Computing Machinery.ISSN: 0730-0301.DOI: 10. 1145/3592416.URL: https://doi....

  42. [63]

    Assessing the Utility of GAN-Generated 3D Virtual Desert Terrain: A User- Centric Evaluation of Immersion and Realism

    Rahul K. Rai et al. “Assessing the Utility of GAN-Generated 3D Virtual Desert Terrain: A User- Centric Evaluation of Immersion and Realism”. En- glish. In:Smart Innovation, Systems and Technolo- gies382 (2024). Ed. by Nakamatsu K, Patnaik S, and Kountchev R. ISBN: 978-98199901...

  43. [65]

    LART: Neural Correspondence Learning with Latent Regularization Transformer for 3D Motion Transfer

    Haoyu Chen et al. “LART: Neural Correspondence Learning with Latent Regularization Transformer for 3D Motion Transfer”. English. In:Advances in Neu- ral Information Processing Systems. Ed. by Oh A et al. V ol. 36. ISSN: 10495258 Type: Conference paper. Neural information proce...

  44. [66]

    DreamUp3D: Object-Centric Genera- tive Models for Single-View 3D Scene Understanding and Real-to-Sim Transfer

    Yizhe Wu et al. “DreamUp3D: Object-Centric Genera- tive Models for Single-View 3D Scene Understanding and Real-to-Sim Transfer”. In:IEEE Robotics and Automation Letters9.4 (Apr. 2024), pp. 3291–3298. ISSN: 2377-3766.DOI: 10.1109/LRA.2024.3362678

  45. [67]

    CLAY: A Controllable Large- scale Generative Model for Creating High-quality 3D Assets

    Longwen Zhang et al. “CLAY: A Controllable Large- scale Generative Model for Creating High-quality 3D Assets”. In:ACM Trans. Graph.43.4 (July 2024). Place: New York, NY , USA Publisher: Association for Computing Machinery.ISSN: 0730-0301.DOI: 10. 1145/3658146.URL: https://doi....

  46. [68]

    MultiCAD: Contrastive Represen- tation Learning for Multi-modal 3D Computer-Aided Design Models

    Weijian Ma et al. “MultiCAD: Contrastive Represen- tation Learning for Multi-modal 3D Computer-Aided Design Models”. In:Proceedings of the 32nd ACM International Conference on Information and Knowl- edge Management. CIKM ’23. event-place: Birming- ham, United Kingdom. New York...

  47. [70]

    360° Reconstruction From a Single Image Using Space Carved Outpainting

    Nuri Ryu et al. “360° Reconstruction From a Single Image Using Space Carved Outpainting”. In:SIG- GRAPH Asia 2023 Conference Papers. SA ’23. event- place: Sydney, NSW, Australia. New York, NY , USA: Association for Computing Machinery, 2023.ISBN: 9798400703157.DOI: 10 . 1145 /...

  48. [71]

    Cross-modal 3D Shape Gener- ation and Manipulation

    Zezhou Cheng et al. “Cross-modal 3D Shape Gener- ation and Manipulation”. English. In:Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)13663 LNCS (2022). Ed. by Avidan S et al. ISBN: 978-3031200...

  49. [72]

    ChartPointFlow for Topology-Aware 3D Point Cloud Generation

    Takumi Kimura, Takashi Matsubara, and Kuniaki Ue- hara. “ChartPointFlow for Topology-Aware 3D Point Cloud Generation”. In:Proceedings of the 29th ACM International Conference on Multimedia. MM ’21. event-place: Virtual Event, China. New York, NY , USA: Association for Computin...

  50. [74]

    Elevating Perception: Unified Recognition Framework and Vision-Language Pre- Training Using Three-Dimensional Image Reconstruc- tion

    ZhiQiang Wang et al. “Elevating Perception: Unified Recognition Framework and Vision-Language Pre- Training Using Three-Dimensional Image Reconstruc- tion”. In:2023 2nd International Conference on Arti- ficial Intelligence, Human-Computer Interaction and Robotics (AIHCIR). Dec...

  51. [75]

    Self-Supervised Learning on 3D Point Clouds by Learning Discrete Generative Models

    Benjamin Eckart et al. “Self-Supervised Learning on 3D Point Clouds by Learning Discrete Generative Models”. In:2021 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR). ISSN: 2575-7075. June 2021, pp. 8244–8253.DOI: 10.1109/ CVPR46437.2021.00815

  52. [76]

    EASI-Tex: Edge-Aware Mesh Texturing from Single Image

    Sai Raj Kishore Perla et al. “EASI-Tex: Edge-Aware Mesh Texturing from Single Image”. In:ACM Trans. Graph.43.4 (July 2024). Place: New York, NY , USA Publisher: Association for Computing Machin- ery.ISSN: 0730-0301.DOI: 10.1145/3658222.URL: https://doi.org/10.1145/3658222

  53. [77]

    F-3DGS: Factorized Coordinates and Representations for 3D Gaussian Splatting

    Xiangyu Sun et al. “F-3DGS: Factorized Coordinates and Representations for 3D Gaussian Splatting”. In: Proceedings of the 32nd ACM International Confer- ence on Multimedia. MM ’24. event-place: Melbourne VIC, Australia. New York, NY , USA: Association for Computing Machinery, ...

  54. [78]

    3DShape2VecSet: A 3D Shape Rep- resentation for Neural Fields and Generative Diffusion Models

    Biao Zhang et al. “3DShape2VecSet: A 3D Shape Rep- resentation for Neural Fields and Generative Diffusion Models”. In:ACM Trans. Graph.42.4 (July 2023). Place: New York, NY , USA Publisher: Association for Computing Machinery.ISSN: 0730-0301.DOI: 10. 1145/3592442.URL: https://...

  55. [79]

    HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image

    Tong Wu et al. “HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image”. In:SIGGRAPH Asia 2023 Conference Papers. SA ’23. event-place: Sydney, NSW, Australia. New York, NY , USA: Association for Computing Machinery, 2023.ISBN: 9798400703157.DOI: 10...

  56. [80]

    An End-to-End Conditional Generative Adversarial Network Based on Depth Map for 3D Craniofacial Reconstruction

    Niankai Zhang et al. “An End-to-End Conditional Generative Adversarial Network Based on Depth Map for 3D Craniofacial Reconstruction”. In:Proceedings of the 30th ACM International Conference on Multi- media. MM ’22. event-place: Lisboa, Portugal. New York, NY , USA: Associatio...

  57. [81]

    Object-centric Learning with Capsule Networks: A Survey

    Fabio De Sousa Ribeiro et al. “Object-centric Learning with Capsule Networks: A Survey”. In:ACM Com- 16 put. Surv.56.11 (July 2024). Place: New York, NY , USA Publisher: Association for Computing Machin- ery.ISSN: 0360-0300.DOI: 10.1145/3674500.URL: https://doi.org/10.1145/3674500

  58. [82]

    Learning to Gen- erate 3D Shapes from a Single Example

    Rundi Wu and Changxi Zheng. “Learning to Gen- erate 3D Shapes from a Single Example”. In:ACM Trans. Graph.41.6 (Nov. 2022). Place: New York, NY , USA Publisher: Association for Computing Ma- chinery.ISSN: 0730-0301.DOI: 10 . 1145 / 3550454 . 3555480.URL: https : / / doi . org ...

  59. [83]

    Text-Free Controllable 3-D Point Cloud Generation

    Haihong Xiao et al. “Text-Free Controllable 3-D Point Cloud Generation”. English. In:IEEE Transactions on Instrumentation and Measurement73 (2024). Pub- lisher: Institute of Electrical and Electronics Engineers Inc. Type: Article, pp. 1–12.ISSN: 00189456.DOI: 10. 1109/TIM.2024...

  60. [84]

    A Study on V oxel Shape Generation and Reconstruction with VQ-V AE- 2

    Kenta Nakada and Hideaki Kimata. “A Study on V oxel Shape Generation and Reconstruction with VQ-V AE- 2”. In:Proceedings of the 2023 7th International Con- ference on Graphics and Signal Processing. ICGSP ’23. event-place: Fujisawa, Japan. New York, NY , USA: Association for C...

  61. [85]

    Embracing the Dark Knowledge: Domain Generalization Using Regularized Knowl- edge Distillation

    Yufei Wang et al. “Embracing the Dark Knowledge: Domain Generalization Using Regularized Knowl- edge Distillation”. In:Proceedings of the 29th ACM International Conference on Multimedia. MM ’21. event-place: Virtual Event, China. New York, NY , USA: Association for Computing M...

  62. [90]

    Text-to-3D Generative AI on Mobile Devices: Measurements and Optimizations

    Xuechen Zhang et al. “Text-to-3D Generative AI on Mobile Devices: Measurements and Optimizations”. In:Proceedings of the 2023 Workshop on Emerg- ing Multimedia Systems. EMS ’23. event-place: New York, NY , USA. New York, NY , USA: Association for Computing Machinery, 2023, pp....

  63. [91]

    RelScene: A Benchmark and base- line for Spatial Relations in text-driven 3D Scene Generation

    Zhaoda Ye et al. “RelScene: A Benchmark and base- line for Spatial Relations in text-driven 3D Scene Generation”. In:Proceedings of the 32nd ACM Inter- national Conference on Multimedia. MM ’24. event- place: Melbourne VIC, Australia. New York, NY , USA: Association for Comput...

  64. [92]

    CAD Translator: An Effective Drive for Text to 3D Parametric Computer-Aided Design Generative Modeling

    Xueyang Li et al. “CAD Translator: An Effective Drive for Text to 3D Parametric Computer-Aided Design Generative Modeling”. In:Proceedings of the 32nd ACM International Conference on Multimedia. MM ’24. event-place: Melbourne VIC, Australia. New York, NY , USA: Association for...

  65. [93]

    ImageBind3D: Image as Binding Step for Controllable 3D Generation

    Zhenqiang Li et al. “ImageBind3D: Image as Binding Step for Controllable 3D Generation”. In:Proceed- ings of the 32nd ACM International Conference on Multimedia. MM ’24. event-place: Melbourne VIC, Australia. New York, NY , USA: Association for Computing Machinery, 2024, pp. 3...

  66. [94]

    DreamLCM: Towards High Quality Text-to-3D Generation via Latent Consistency Model

    Yiming Zhong et al. “DreamLCM: Towards High Quality Text-to-3D Generation via Latent Consistency Model”. In:Proceedings of the 32nd ACM Interna- tional Conference on Multimedia. MM ’24. event- place: Melbourne VIC, Australia. New York, NY , USA: Association for Computing Machi...

  67. [95]

    Dream Mesh: A Speech-to-3D Model Generative Pipeline in Mixed Reality

    Suibi Che-Chuan Weng, Yan-Ming Chiou, and Ellen Yi-Luen Do. “Dream Mesh: A Speech-to-3D Model Generative Pipeline in Mixed Reality”. In:2024 IEEE International Conference on Artificial Intelligence and 17 eXtended and Virtual Reality (AIxVR). ISSN: 2771-

  68. [96]

    PlacidDreamer: Advancing Har- mony in Text-to-3D Generation

    Shuo Huang et al. “PlacidDreamer: Advancing Har- mony in Text-to-3D Generation”. In:Proceedings of the 32nd ACM International Conference on Multime- dia. MM ’24. event-place: Melbourne VIC, Australia. New York, NY , USA: Association for Computing Ma- chinery, 2024, pp. 6880–68...

  69. [97]

    Text2VRScene: Exploring the Framework of Automated Text-driven Generation Sys- tem for VR Experience

    Zhizhuo Yin et al. “Text2VRScene: Exploring the Framework of Automated Text-driven Generation Sys- tem for VR Experience”. English. In:Proceedings - 2024 IEEE Conference on Virtual Reality and 3D User Interfaces, VR 2024. Type: Conference paper. Institute of Electrical and Ele...

  70. [98]

    Prompt Engineering for X3D Object Creation with LLMs

    Nicholas Polys, Ayat Mohammed, and Ben Sandbrook. “Prompt Engineering for X3D Object Creation with LLMs”. In:Proceedings of the 29th International ACM Conference on 3D Web Technology. Web3D ’24. event- place: Guimar ˜aes, Portugal. New York, NY , USA: Association for Computing...

  71. [99]

    Sketch3D: Style- Consistent Guidance for Sketch-to-3D Generation

    Wangguandong Zheng et al. “Sketch3D: Style- Consistent Guidance for Sketch-to-3D Generation”. In: Proceedings of the 32nd ACM International Confer- ence on Multimedia. MM ’24. event-place: Melbourne VIC, Australia. New York, NY , USA: Association for Computing Machinery, 2024,...

  72. [100]

    SpaceBlender: Creating Context- Rich Collaborative Spaces Through Generative 3D Scene Blending

    Nels Numan et al. “SpaceBlender: Creating Context- Rich Collaborative Spaces Through Generative 3D Scene Blending”. In:Proceedings of the 37th Annual ACM Symposium on User Interface Software and Tech- nology. UIST ’24. event-place: Pittsburgh, PA, USA. New York, NY , USA: Asso...

  73. [101]

    Knowl- edge Generation Pipeline using LLM for Building 3D Object Knowledge Base

    SooHyung Lee, HyeRin Lee, and KiSuk Lee. “Knowl- edge Generation Pipeline using LLM for Building 3D Object Knowledge Base”. In:2023 14th International Conference on Information and Communication Tech- nology Convergence (ICTC). ISSN: 2162-1241. Oct. 2023, pp. 1303–1305.DOI: 10...

  74. [102]

    ControlStyle: Text-Driven Styl- ized Image Generation Using Diffusion Priors

    Jingwen Chen et al. “ControlStyle: Text-Driven Styl- ized Image Generation Using Diffusion Priors”. In: Proceedings of the 31st ACM International Confer- ence on Multimedia. MM ’23. event-place: Ottawa ON, Canada. New York, NY , USA: Association for Computing Machinery, 2023, ...

  75. [103]

    ControlMat: A Controlled Generative Approach to Material Capture

    Giuseppe Vecchio et al. “ControlMat: A Controlled Generative Approach to Material Capture”. In:ACM Trans. Graph.43.5 (Sept. 2024). Place: New York, NY , USA Publisher: Association for Computing Machin- ery.ISSN: 0730-0301.DOI: 10.1145/3688830.URL: https://doi.org/10.1145/3688830

  76. [104]

    Diffusion Texture Painting

    Anita Hu et al. “Diffusion Texture Painting”. In:ACM SIGGRAPH 2024 Conference Papers. SIGGRAPH ’24. event-place: Denver, CO, USA. New York, NY , USA: Association for Computing Machinery, 2024. ISBN: 9798400705250.DOI: 10 . 1145 / 3641519 . 3657458.URL: https : / / doi . org / ...

  77. [105]

    MatFormer: a generative model for procedural materials

    Paul Guerrero et al. “MatFormer: a generative model for procedural materials”. In:ACM Trans. Graph.41.4 (July 2022). Place: New York, NY , USA Publisher: Association for Computing Machinery.ISSN: 0730- 0301.DOI: 10 .1145 /3528223 . 3530173.URL: https : //doi.org/10.1145/352822...

  78. [106]

    MaPa: Text-driven Photoreal- istic Material Painting for 3D Shapes

    Shangzhan Zhang et al. “MaPa: Text-driven Photoreal- istic Material Painting for 3D Shapes”. In:ACM SIG- GRAPH 2024 Conference Papers. SIGGRAPH ’24. event-place: Denver, CO, USA. New York, NY , USA: Association for Computing Machinery, 2024.ISBN: 9798400705250.DOI: 10 . 1145 /...

  79. [107]

    PhotoMat: A Material Generator Learned from Single Flash Photos

    Xilong Zhou et al. “PhotoMat: A Material Generator Learned from Single Flash Photos”. In:ACM SIG- GRAPH 2023 Conference Proceedings. SIGGRAPH ’23. event-place: Los Angeles, CA, USA. New York, NY , USA: Association for Computing Machinery, 2023.ISBN: 9798400701597.DOI: 10.1145/...

  80. [108]

    Large Language and Text-to-3D Models for Engineer- ing Design Optimization

    Thiago Rios, Stefan Menzel, and Bernhard Sendhoff. “Large Language and Text-to-3D Models for Engineer- ing Design Optimization”. In:2023 IEEE Symposium Series on Computational Intelligence (SSCI). ISSN: 2472-8322. Dec. 2023, pp. 1704–1711.DOI: 10.1109/ SSCI52147.2023.10371898

  81. [109]

    FloorGAN: Generative Net- work for Automated Floor Layout Generation

    Abhinav Upadhyay et al. “FloorGAN: Generative Net- work for Automated Floor Layout Generation”. In: Proceedings of the 6th Joint International Confer- ence on Data Science & Management of Data (10th ACM IKDD CODS and 28th COMAD). CODS- COMAD ’23. event-place: Mumbai, India...

  82. [110]

    Automated video editing based on learned styles using LSTM-GAN

    Hsin-I Huang, Chi-Sheng Shih, and Zi-Lin Yang. “Automated video editing based on learned styles using LSTM-GAN”. In:Proceedings of the 37th ACM/SIGAPP Symposium on Applied Computing. 18 SAC ’22. event-place: Virtual Event. New York, NY , USA: Association for Computing Machiner...

  83. [111]

    TexPainter: Generative Mesh Texturing with Multi-view Consistency

    Hongkun Zhang et al. “TexPainter: Generative Mesh Texturing with Multi-view Consistency”. In:ACM SIGGRAPH 2024 Conference Papers. SIGGRAPH ’24. event-place: Denver, CO, USA. New York, NY , USA: Association for Computing Machinery, 2024. ISBN: 9798400705250.DOI: 10 . 1145 / 364...

  84. [114]

    Cascade Variational Auto-Encoder for Hierarchical Disentanglement

    Fudong Lin et al. “Cascade Variational Auto-Encoder for Hierarchical Disentanglement”. In:Proceedings of the 31st ACM International Conference on In- formation & Knowledge Management. CIKM ’22. event-place: Atlanta, GA, USA. New York, NY , USA: Association for Computing Ma...

  85. [115]

    Automated Testing of Graphics Units by Deep-Learning Detection of Vi- sual Anomalies

    Lev Faivishevsky et al. “Automated Testing of Graphics Units by Deep-Learning Detection of Vi- sual Anomalies”. In:Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. KDD ’21. event-place: Virtual Event, Singapore. New York, NY , USA: Associ...

  86. [116]

    RealFill: Reference-Driven Gen- eration for Authentic Image Completion

    Luming Tang et al. “RealFill: Reference-Driven Gen- eration for Authentic Image Completion”. In:ACM Trans. Graph.43.4 (July 2024). Place: New York, NY , USA Publisher: Association for Computing Machin- ery.ISSN: 0730-0301.DOI: 10.1145/3658237.URL: https://doi.org/10.1145/3658237

  87. [117]

    A Survey on Deep Generative 3D-aware Image Synthesis

    Weihao Xia and Jing-Hao Xue. “A Survey on Deep Generative 3D-aware Image Synthesis”. In:ACM Comput. Surv.56.4 (Nov. 2023). Place: New York, NY , USA Publisher: Association for Computing Machin- ery.ISSN: 0360-0300.DOI: 10.1145/3626193.URL: https://doi.org/10.1145/3626193

  88. [118]

    Image Generation Model Applying PCA on Latent Space

    Myung Keun Song, Asim Niaz, and Kwang Nam Choi. “Image Generation Model Applying PCA on Latent Space”. In:Proceedings of the 2023 2nd Asia Conference on Algorithms, Computing and Machine Learning. CACML ’23. event-place: Shanghai, China. New York, NY , USA: Association for Com...

  89. [119]

    Broomrocket: Open Source Text-to-3D Algorithm for 3D Object Placement

    Sanja Bonic, Janos Bonic, and Stefan Schmid. “Broomrocket: Open Source Text-to-3D Algorithm for 3D Object Placement”. In:ACM Games2.3 (Aug. 2024). Place: New York, NY , USA Publisher: Associa- tion for Computing Machinery.DOI: 10.1145/3648233. URL: https://doi.org/10.1145/3648233

  90. [120]

    Image based and Point Cloud based Methods for 3D View Reconstruction in Real- time Environment

    Arya Agrawal et al. “Image based and Point Cloud based Methods for 3D View Reconstruction in Real- time Environment”. English. In:2024 IEEE 14th An- nual Computing and Communication Workshop and Conference, CCWC 2024. Ed. by Paul R and Kundu A. Type: Conference paper. Institut...

  91. [121]

    GenQuery: Supporting Expressive Visual Search with Generative Models

    Kihoon Son et al. “GenQuery: Supporting Expressive Visual Search with Generative Models”. In:Proceed- ings of the 2024 CHI Conference on Human Factors in Computing Systems. CHI ’24. event-place: Honolulu, HI, USA. New York, NY , USA: Association for Com- puting Machinery, 2024...

  92. [122]

    Prompt- Paint: Steering Text-to-Image Generation Through Paint Medium-like Interactions

    John Joon Young Chung and Eytan Adar. “Prompt- Paint: Steering Text-to-Image Generation Through Paint Medium-like Interactions”. In:Proceedings of the 36th Annual ACM Symposium on User Inter- face Software and Technology. UIST ’23. event-place: San Francisco, CA, USA. New York...

  93. [123]

    Controllable Data Generation by Deep Learning: A Review

    Shiyu Wang et al. “Controllable Data Generation by Deep Learning: A Review”. In:ACM Comput. Surv. 56.9 (Apr. 2024). Place: New York, NY , USA Pub- lisher: Association for Computing Machinery.ISSN: 0360-0300.DOI: 10.1145/3648609.URL: https://doi. org/10.1145/3648609

  94. [124]

    Generative AI: A Re- view on Models and Applications

    Kuldeep Singh Kaswan et al. “Generative AI: A Re- view on Models and Applications”. In:2023 Inter- national Conference on Communication, Security and Artificial Intelligence (ICCSAI). Nov. 2023, pp. 699– 704.DOI: 10.1109/ICCSAI59793.2023.10421601. 19

  95. [125]

    Diffusion Models: A Comprehensive Survey of Methods and Applications

    Ling Yang et al. “Diffusion Models: A Comprehensive Survey of Methods and Applications”. In:ACM Com- put. Surv.56.4 (Nov. 2023). Place: New York, NY , USA Publisher: Association for Computing Machin- ery.ISSN: 0360-0300.DOI: 10.1145/3626235.URL: https://doi.org/10.1145/3626235

  96. [126]

    Gen- erative Adversarial Networks in Computer Vision: A Survey and Taxonomy

    Zhengwei Wang, Qi She, and Tom ´as E. Ward. “Gen- erative Adversarial Networks in Computer Vision: A Survey and Taxonomy”. In:ACM Comput. Surv.54.2 (Feb. 2021). Place: New York, NY , USA Publisher: Association for Computing Machinery.ISSN: 0360- 0300.DOI: 10.1145/3439723.URL: ...

  97. [127]

    Understanding and Creat- ing Art with AI: Review and Outlook

    Eva Cetinic and James She. “Understanding and Creat- ing Art with AI: Review and Outlook”. In:ACM Trans. Multimedia Comput. Commun. Appl.18.2 (Feb. 2022). Place: New York, NY , USA Publisher: Association for Computing Machinery.ISSN: 1551-6857.DOI: 10. 1145/3475799.URL: https:...

  98. [128]

    Explainable Convolutional Neural Networks: A Taxonomy, Re- view, and Future Directions

    Rami Ibrahim and M. Omair Shafiq. “Explainable Convolutional Neural Networks: A Taxonomy, Re- view, and Future Directions”. In:ACM Comput. Surv. 55.10 (Feb. 2023). Place: New York, NY , USA Pub- lisher: Association for Computing Machinery.ISSN: 0360-0300.DOI: 10.1145/3563691.U...

  99. [129]

    The Infinite Index: Informa- tion Retrieval on Generative Text-To-Image Models

    Niklas Deckers et al. “The Infinite Index: Informa- tion Retrieval on Generative Text-To-Image Models”. In:Proceedings of the 2023 Conference on Human Information Interaction and Retrieval. CHIIR ’23. event-place: Austin, TX, USA. New York, NY , USA: Association for Computing ...

  100. [130]

    Creativity and Machine Learning: A Survey

    Giorgio Franceschelli and Mirco Musolesi. “Creativity and Machine Learning: A Survey”. In:ACM Com- put. Surv.56.11 (June 2024). Place: New York, NY , USA Publisher: Association for Computing Machin- ery.ISSN: 0360-0300.DOI: 10.1145/3664595.URL: https://doi.org/10.1145/3664595

  101. [131]

    Death of the Design Researcher? Creating Knowledge Resources for De- signers Using Generative AI

    Willem Van Der Maden et al. “Death of the Design Researcher? Creating Knowledge Resources for De- signers Using Generative AI”. In:Companion Pub- lication of the 2024 ACM Designing Interactive Sys- tems Conference. DIS ’24 Companion. event-place: IT University of Copenhagen, D...

  102. [132]

    Design Ideation with AI - Sketching, Thinking and Talking with Gen- erative Machine Learning Models

    Jakob Tholander and Martin Jonsson. “Design Ideation with AI - Sketching, Thinking and Talking with Gen- erative Machine Learning Models”. In:Proceedings of the 2023 ACM Designing Interactive Systems Con- ference. DIS ’23. event-place: Pittsburgh, PA, USA. New York, NY , USA: ...

  103. [133]

    The Impact of Sketch-guided vs. Prompt-guided 3D Generative AIs on the Design Exploration Process

    Seung Won Lee et al. “The Impact of Sketch-guided vs. Prompt-guided 3D Generative AIs on the Design Exploration Process”. In:Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. CHI ’24. event-place: Honolulu, HI, USA. New York, NY , USA: Association f...

  104. [134]

    ”We Are Visual Thinkers, Not Ver- bal Thinkers!

    Hyerim Park et al. “”We Are Visual Thinkers, Not Ver- bal Thinkers!”: A Thematic Analysis of How Profes- sional Designers Use Generative AI Image Generation Tools”. In:Proceedings of the 13th Nordic Confer- ence on Human-Computer Interaction. NordiCHI ’24. event-place: Uppsala...

  105. [135]

    Evolving Roles and Workflows of Creative Practitioners in the Age of Generative AI

    Srishti Palani and Gonzalo Ramos. “Evolving Roles and Workflows of Creative Practitioners in the Age of Generative AI”. In:Proceedings of the 16th Con- ference on Creativity & Cognition. C&C ’24. event-place: Chicago, IL, USA. New York, NY , USA: Association for Comput...

  106. [136]

    Design Guide- lines for Prompt Engineering Text-to-Image Genera- tive Models

    Vivian Liu and Lydia B Chilton. “Design Guide- lines for Prompt Engineering Text-to-Image Genera- tive Models”. In:Proceedings of the 2022 CHI Con- ference on Human Factors in Computing Systems. CHI ’22. event-place: New Orleans, LA, USA. New York, NY , USA: Association for Co...

  107. [137]

    The Eyes, the Hands and the Brain: What can Text-to-Image Models Offer for Game Design and Visual Creativity?

    Hongwei Zhou et al. “The Eyes, the Hands and the Brain: What can Text-to-Image Models Offer for Game Design and Visual Creativity?” In:Proceedings of the 19th International Conference on the Founda- tions of Digital Games. FDG ’24. event-place: Worces- ter, MA, USA. New York, ...

  108. [138]

    Predictive AI for the 3D-IC Design Pro- cess Reduces the Iteration

    Julian Sun. “Predictive AI for the 3D-IC Design Pro- cess Reduces the Iteration”. In:2024 International VLSI Symposium on Technology, Systems and Applica- tions (VLSI TSA). Apr. 2024, pp. 1–5.DOI: 10.1109/ VLSITSA60681.2024.10546344

  109. [139]

    Generative AI in the Wild: Prospects, Challenges, and Strategies

    Yuan Sun et al. “Generative AI in the Wild: Prospects, Challenges, and Strategies”. In:Proceedings of the 2024 CHI Conference on Human Factors in Com- puting Systems. CHI ’24. event-place: Honolulu, HI, USA. New York, NY , USA: Association for Comput- ing Machinery, 2024.ISBN:...

  110. [140]

    ”I’m a Solo Developer but AI is My New Ill-Informed Co- Worker

    Ruchi Panchanadikar and Guo Freeman. “”I’m a Solo Developer but AI is My New Ill-Informed Co- Worker”: Envisioning and Designing Generative AI to Support Indie Game Development”. In:Proc. ACM Hum.-Comput. Interact.8.CHI PLAY (Oct. 2024). Place: New York, NY , USA Publisher: As...

  111. [141]

    Empowering the Meta- verse with Generative AI: Survey and Future Direc- tions

    Hua Xuan Qin and Pan Hui. “Empowering the Meta- verse with Generative AI: Survey and Future Direc- tions”. In:2023 IEEE 43rd International Conference on Distributed Computing Systems Workshops (ICD- CSW). ISSN: 2332-5666. July 2023, pp. 85–90.DOI: 10.1109/ICDCSW60045.2023.00022

  112. [142]

    LEARNING CONTINUOUS ENVI- RONMENT FIELDS VIA IMPLICIT FUNCTIONS

    Xueting Li et al. “LEARNING CONTINUOUS ENVI- RONMENT FIELDS VIA IMPLICIT FUNCTIONS”. English. In:ICLR 2022 - 10th International Con- ference on Learning Representations. Type: Con- ference paper. International Conference on Learn- ing Representations, ICLR, 2022.URL: https : /...

  113. [143]

    A Comprehensive Survey on Generative AI for Metaverse: Enabling Immer- sive Experience

    Vinay Chamola et al. “A Comprehensive Survey on Generative AI for Metaverse: Enabling Immer- sive Experience”. English. In:Cognitive Computa- tion(2024). Publisher: Springer Type: Review.ISSN: 18669956.DOI: 10 . 1007 / s12559 - 024 - 10342 - 9. URL: https : / / www . scopus . ...

  114. [7075]

    8496–8506.DOI: 10

    June 2023, pp. 8496–8506.DOI: 10 . 1109 / CVPR52729.2023.00821

  115. [7453]

    2024, pp

    Jan. 2024, pp. 345–349.DOI: 10 . 1109 / AIxVR59861.2024.00059

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.