Pith. sign in

REVIEW 3 major objections 3 minor 62 references

Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation

T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Interactive visual prompts replace blind trial-and-error in text-to-3D design.

desk verdict A promising system idea with an internal quality-claim contradiction and unreadable validation sections in this artifact; worth refereeing on the strength of the concept, but the claims as stated are unverifiable. read the letter →

arxiv 2508.00428 v1 pith:S3LWA6CP submitted 2025-08-01 cs.GR cs.HC

classification cs.GRcs.HC
keywords text-to-3Dgenerationvisualpromptengineeringmulti-modallanguagemodelsanalyticsrecommendationmulti-viewconsistencycreativitysupportinteractivesystems
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Text-to-3D generation is fast but unpredictable: users type a description, wait through a black-box process, and often get a defective model that needs many blind retries. The paper claims that this trial-and-error loop can be replaced by a guided visual workflow. It introduces Sel3DCraft, an interactive system that proposes candidate models through both retrieval and generation, scores them from multiple rendered views using multimodal large language models, and visualizes scores and defects so users can click on a problem and get a prompt edit. In a user study, designers finished models in 118.83 seconds instead of 402.17 seconds, a 70.5% reduction, with 66.2% fewer prompt iterations and a rated-quality jump from 2.46 to 4.58 out of five. If the results hold, the system turns text-to-3D from a lottery into a designer-controlled process.

What carries the argument

The central mechanism is multi-view hybrid scoring, which renders each candidate model from multiple viewpoints and asks a multimodal large language model to produce high-level quality metrics for each view and aspect. These scores serve two purposes: they cluster similar candidates in a satellite-chart view for quick exploration, and they feed a per-view heatmap that localizes defects such as texture tears or geometry glitches. The second essential piece is the prompt-driven treemap wordle, an image-text matching visualization that links semantic keywords to the multi-view images, so a user can click a suggested word to refine the prompt. Together, the scoring loop and the wordle convert an unstructured search over prompts into a visible cycle of evaluate-defect-recommend-refine.

What would settle it

Re-run the user study with the MLLM scoring replaced by random scores while keeping everything else identical; if time-to-completion and quality ratings stay about the same, the scoring is not the cause of the reported gains. Alternatively, compare MLLM per-view scores with independent human expert ratings on the same candidate set; a correlation near zero would falsify the claim that the scoring matches human-expert consistency.

Watch

Extended reading notes

Core claim

The central claim is that the whole interaction loop — dual-branch candidate synthesis, MLLM-based multi-view scoring, and prompt-driven visual analytics — produces the measured gains, not any single generation model in isolation. The paper argues that rendering each candidate from several views and feeding those images to a multimodal large language model yields high-level quality metrics, such as structural integrity and cross-view coherence, that track what human experts would say. Those scores drive a satellite-chart clustering view and a per-view heatmap that localize defects, while a treemap wordle maps visual defects back to prompt keywords. Closing the loop lets a designer refine a prompt by clicking on words rather than rewriting entire descriptions. The reported outcome is faster model creation, fewer iterations, and higher rated quality than existing text-to-3D systems.

Load-bearing premise

The assumption that the MLLM's multi-view quality scores agree with what human experts consider good 3D models; if the scores are out of line, the prompt recommendations can lead users to models that score well but are not actually better.

Editorial extensions

If this is right

  • If the interaction loop works as reported, professional 3D design tools can move from blind prompt retries to a guided evaluate-and-refine cycle, cutting the per-model cost dramatically.
  • The MLLM-based multi-view scoring could be reused outside the system as a reference-free evaluation metric for text-to-3D output quality, giving researchers a proxy that tracks human judgment.
  • Mixing retrieval from existing 3D assets with freshly generated candidates is a cheap way to broaden exploration without extra generation cost, a design pattern other interfaces could adopt.
  • The defect-to-keyword mapping embodied in the wordle could be applied to text-to-image and other generative model interfaces, making visual analytics a standard part of prompt engineering.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the reported gains are for the full system, the paper does not isolate how much each component contributes; a plausible open question is whether MLLM scoring alone would already provide most of the time savings.
  • The user study numbers come from a specific set of designers and prompts; repeating the study with novice users or with a different base generator would clarify how much of the benefit is interaction design versus the underlying generation speed.
  • One testable extension is to detach the scoring module from the interface and use it to rank prompts offline, which would let the system suggest entire prompt variants rather than single keywords.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper introduces Sel3DCraft, an interactive visual prompt engineering system for text-to-3D (T23D) generation. The system combines a dual retrieval/generation branch, MLLM-based multi-view hybrid scoring, and a visual analytics suite (multi-view satellite charts, hybrid-level scoring heatmaps, and a prompt-driven treemap wordle) to guide users from unstructured prompt trial-and-error to structured exploration. The abstract claims that user studies show a 70.5% reduction in model creation time, a 66.2% reduction in prompt iterations, and significantly higher model quality ratings (4.58 vs 2.46, Fig. 8), concluding that Sel3DCraft surpasses other T23D systems in supporting designer creativity.

Significance. If the reported results are reproducible and the MLLM scoring genuinely aligns with human expert judgment, the system would be a meaningful contribution to visual analytics and creativity support for 3D content creation. The proposed cross-modal visual representation and the focus on multi-view consistency are timely. However, the provided manuscript text is largely unreadable due to encoding corruption, and the internal claims about quality are inconsistent. The significance therefore hinges entirely on material that cannot currently be inspected.

major comments (3)
  1. [Abstract / Contributions] The abstract states that Sel3DCraft yields 'significantly higher model quality ratings (4.58 vs 2.46, Fig. 8)', while the third listed contribution states that the system 'achieves faster model creation while maintaining comparable quality.' These statements describe different outcomes for the same comparison. The manuscript must specify which baseline each figure refers to, report the full statistical details (sample size, test, effect size, confidence intervals), and reconcile the two phrasings; otherwise the central quality advantage is ambiguous.
  2. [Sections 6–7] The validation of the MLLM hybrid scoring (Section 6) and the user study (Section 7) are unreadable in the provided version of the manuscript because the text is corrupted (repeated glyph placeholders replace the actual prose). These sections are load-bearing for every quantitative claim in the abstract, including the 70.5% time reduction, the 66.2% iteration reduction, and the 4.58 vs 2.46 quality ratings. A readable, complete version of these sections with the full experimental protocol, baseline definitions, participant demographics, and raw measurements must be provided before the central claims can be assessed.
  3. [Sec. 4.2 / Abstract] The multi-view hybrid scoring approach relies on the assumption that MLLM-based high-level metrics correlate with human expert judgment of 3D model quality. Because the validation section is unreadable, there is no evidence in the accessible text that this assumption holds. The paper should provide explicit correlation or agreement measures between MLLM scores and human ratings, and should ablate the contribution of the scoring loop versus the retrieval branch versus the generation branch to the reported time and iteration reductions.
minor comments (3)
  1. [Throughout] Large portions of the text, including in-text citations and reference titles, are corrupted with placeholder glyphs, making it impossible to read the full methodology and related work discussion. The authors should ensure a correctly encoded PDF is submitted.
  2. [References] Several reference entries have corrupted titles and venue names (e.g., references [1], [4], [5], [10]). These should be cleaned up so that readers can locate the cited works.
  3. [Fig. 8] The abstract cites Fig. 8 for the quality ratings, but the figure is not viewable in the provided artifact. Please ensure all figures are embedded correctly and referenced consistently.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the system's central claims are anchored to external user-study measurements, not to its own scoring definitions.

full rationale

No circularity found. The paper's central claim is an empirical user-study comparison of an interactive text-to-3D system against baseline T23D systems; the headline quantities (70.5% time reduction, 66.2% iteration reduction, and the 4.58 vs 2.46 quality ratings) are presented as human-measured outcomes, not as outputs of the system's own MLLM scoring function. The multi-view hybrid MLLM scoring in Sec. 4.2 is an internal component whose alignment with human judgment is itself something the paper says it validates in Sec. 6, and the user study in Sec. 7 is an independent external benchmark for the final claims. There is no fitted parameter that is renamed as a prediction, no equation in which an output reduces by construction to an input, and no load-bearing self-citation chain that forces the conclusion. The abstract's 'significantly higher model quality ratings' versus the contribution's 'comparable quality' is an internal inconsistency in reporting, but inconsistency is not circularity. The artifact's corruption makes Secs. 6-7 unreadable, so the empirical support cannot be verified here; that is a verifiability limitation, not evidence that the derivation is circular. Therefore the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities are identified in the available text (abstract, introduction, references). The central claims rest on domain assumptions about multi-view MLLM evaluation, human-expert alignment, and user-study representativeness.

assumptions (3)
  • domain assumption Multi-view images are an adequate representation of 3D model quality for scoring and refinement.
    The system uses multi-view images as the 3D visual medium for MLLM-based scoring (abstract, Sec. 4.2), assuming that quality, geometry, texture, and view consistency can be judged from rendered views.
  • domain assumption MLLM high-level metrics agree with human-expert judgments of 3D model quality.
    The abstract claims the scoring approach assesses 3D models with human-expert consistency, an unproven premise that underpins the usefulness of the prompt refinement loop.
  • domain assumption The user study participants and tasks represent professional text-to-3D usage.
    The reported time and quality improvements assume the study sample and experimental tasks generalize to real workflows; the method and evaluation sections are not readable to confirm sampling or task design.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation." pith.science (2026). https://pith.science/paper/S3LWA6CP

@misc{pith2026250800428,
  author       = {Pith},
  title        = {Pith review of: Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S3LWA6CP}},
  note         = {Machine review of arXiv:2508.00428}
}
read the original abstract

Text-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While visual prompt engineering has advanced in text-to-image domains, its application to 3D generation presents unique challenges requiring multi-view consistency evaluation and spatial understanding. We present Sel3DCraft, a visual prompt engineering system for T23D that transforms unstructured exploration into a guided visual process. Our approach introduces three key innovations: a dual-branch structure combining retrieval and generation for diverse candidate exploration; a multi-view hybrid scoring approach that leverages MLLMs with innovative high-level metrics to assess 3D models with human-expert consistency; and a prompt-driven visual analytics suite that enables intuitive defect identification and refinement. Extensive testing and user studies demonstrate that Sel3DCraft surpasses other T23D systems in supporting creativity for designers.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

62 extracted references · 36 canonical work pages

  1. [60]

    T. Wu, G. Yang, Z. Li, K. Zhang, Z. Liu, L. J. Guibas, D. Lin, and G. Wetzstein. Gpt-4v(ision) is a human-aligned evaluator for text-to-3d generation. In �������� ���������� �� �������� ������ ��� ������� ������������ ���� ����� �������� ��� ���� ���� ������ ����, pp. 22227– 22238. IEEE, 2024. doi: 10.1109/CVPR52733.2024.02098 3, 5

  2. [1]

    Abdelreheem, A

    A. Abdelreheem, A. Eldesokey, M. Ovsjanikov, and P. Wonka. Zero- shot 3d shape correspondence. In J. Kim, M. C. Lin, and B. Bickel, eds., �������� ���� ���� ���������� ������� �� ����� ������� ���� ���������� �������� ������ ����, pp. 59:1–59:11. ACM, 2023. doi: 10. 1145/3610548.3618228 9

  3. [3]

    Averkiou, V

    M. Averkiou, V . G. Kim, Y . Zheng, and N. J. Mitra. Shapesynth: Param- eterizing model collections for coupled shape exploration and synthesis. ������� ������ �����, 33(2):125–134, 2014. doi: 10.1111/CGF.12310 3

  4. [6]

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert- V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Am...

  5. [7]

    D. Chen, R. Chen, S. Zhang, Y . Wang, Y . Liu, H. Zhou, Q. Zhang, Y . Wan, P. Zhou, and L. Sun. Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark. In ���������� ������������� ���������� �� ������� ��������� ���� ����� ������� �������� ���� ������ ����. OpenReview.net, 2024. 3, 5

  6. [8]

    Y . Chen, T. Wang, T. Wu, X. Pan, K. Jia, and Z. Liu. Comboverse: Compositional 3d assets creation using spatially-aware diffusion guidance. In A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, eds., �������� ������ � ���� ���� � ���� �������� ����������� ������ ������ ��������� ���������� �� ����� ������������ ���� ����, vol. 150...

  7. [9]

    Cherry and C

    E. Cherry and C. Latulipe. Quantifying the creativity support of digital tools through the creativity support index. ��� ������ ������� ���� ���������, 21(4):21:1–21:25, 2014. doi: 10.1145/2617588 8

  8. [10]

    J. J. Y . Chung and E. Adar. Promptpaint: Steering text-to-image generation through paint medium-like interactions. pp. 6:1–6:17, 2023. doi: 10.1145/ 3586183.3606777 2

Show all 62 references
  1. [11]

    J. J. Y . Chung, W. Kim, K. M. Yoo, H. Lee, E. Adar, and M. Chang. Talebrush: Sketching stories with generative pretrained language models. In S. D. J. Barbosa, C. Lampe, C. Appert, D. A. Shamma, S. M. Drucker, J. R. Williamson, and K. Yatani, eds.,��� ���� ��� ���������� �� �...

  2. [12]

    Deitke, R

    M. Deitke, R. Liu, M. Wallingford, H. Ngo, O. Michel, A. Kusupati, A. Fan, C. Laforte, V . V oleti, S. Y . Gadre, E. VanderBilt, A. Kemb- havi, C. V ondrick, G. Gkioxari, K. Ehsani, L. Schmidt, and A. Farhadi. Objaverse-xl: A universe of 10m+ 3d objects. In A. Oh, T. Naumann, ...

  3. [13]

    Y . Feng, X. Wang, K. Wong, S. Wang, Y . Lu, M. Zhu, B. Wang, and W. Chen. Promptmagician: Interactive prompt engineering for text-to- image creation. ���� ������ ���� ������� ������, 30(1):295–305, 2024. doi: 10.1109/TVCG.2023.3327168 1, 2, 7

  4. [14]

    T. Gao, A. Fisch, and D. Chen. Making pre-trained language models better few-shot learners. In C. Zong, F. Xia, W. Li, and R. Navigli, eds., ���� �������� �� ��� ���� ������ ������� �� ��� ����������� ��� ������������� ����������� ��� ��� ���� ������������� ����� ���������� ��...

  5. [15]

    H. Han, R. Yang, H. Liao, J. Xing, Z. Xu, X. Yu, J. Zha, X. Li, and W. Li. REPARO: compositional 3d assets generation with differentiable 3d layout alignment. ���� , abs/2405.18525, 2024. doi: 10.48550/ARXIV.2405. 18525 9

  6. [16]

    X. He, J. Chen, S. Peng, D. Huang, Y . Li, X. Huang, C. Yuan, W. Ouyang, and T. He. GVGEN: text-to-3d generation with volumetric representation. In A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, eds., �������� ������ � ���� ���� � ���� �������� ����...

  7. [17]

    Hessel, A

    J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y . Choi. Clipscore: A reference-free evaluation metric for image captioning. In M. Moens, X. Huang, L. Specia, and S. W. Yih, eds.,����������� �� ��� ���� ���� ������� �� ��������� ������� �� ������� �������� ����������� ����...

  8. [18]

    Y . Hong, K. Zhang, J. Gu, S. Bi, Y . Zhou, D. Liu, F. Liu, K. Sunkavalli, T. Bui, and H. Tan. LRM: large reconstruction model for single image to 3d. In ��� ������� ������������� ���������� �� �������� ���������������� ���� ����� ������� �������� ��� ����� ����. OpenReview.ne...

  9. [19]

    Hou and L

    X. Hou and L. Zhang. Saliency detection: A spectral residual approach. In ���� ���� �������� ������� ���������� �� �������� ������ ��� ������� ����������� ����� ������ ����� ���� ����� ������������ ���������� ��� . IEEE Computer Society, 2007. doi: 10.1109/CVPR.2007.383267 4

  10. [20]

    Kam-Kwai, X

    W. Kam-Kwai, X. Wang, Y . Wang, J. He, R. Zhang, and H. Qu. Anchorage: Visual analysis of satisfaction in customer service videos via anchor events. ���� ������ ���� ������� ������, 30(7):4008–4022, 2024. doi: 10.1109/ TVCG.2023.3245609 2

  11. [22]

    J. Li, D. Li, C. Xiong, and S. C. H. Hoi. BLIP: bootstrapping language- image pre-training for unified vision-language understanding and gener- ation. In K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvári, G. Niu, and S. Sabato, eds., ������������� ���������� �� ������� ��������...

  12. [23]

    J. Li, H. Tan, K. Zhang, Z. Xu, F. Luan, Y . Xu, Y . Hong, K. Sunkavalli, G. Shakhnarovich, and S. Bi. Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model. In ��� ������� ������������� ���������� �� �������� ���������������� ���� ����� �������...

  13. [24]

    Y . Li, J. Wang, P. Aboagye, C.-C. M. Yeh, Y . Zheng, L. Wang, W. Zhang, and K.-L. Ma. Visual analytics for efficient image exploration and user- guided image captioning. ���� ������������ �� ������������� ��� ���� ����� ��������, 30(6):2875–2887, 13 pages, apr 2024. doi: 10.1...

  14. [25]

    P. P. Liang, Y . Lyu, G. Chhablani, N. Jain, Z. Deng, X. Wang, L. Morency, and R. Salakhutdinov. Multiviz: Towards visualizing and understand- ing multimodal models. In ��� �������� ������������� ���������� �� �������� ���������������� ���� ����� ������� ������� ��� ���� ����....

  15. [26]

    M. Liu, R. Shi, K. Kuang, Y . Zhu, X. Li, S. Han, H. Cai, F. Porikli, and H. Su. Openshape: Scaling up 3d shape representation towards open-world understanding. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, eds., �������� �� ������ ����������� �������...

  16. [27]

    M. Liu, C. Xu, H. Jin, L. Chen, M. V . T., Z. Xu, and H. Su. One-2-3-45: 10 �� ������ �� ���� ������������ �� ������������� ��� �������� ��������� Any single image to 3d mesh in 45 seconds without per-shape optimization. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt,...

  17. [28]

    R. Liu, R. Wu, B. V . Hoorick, P. Tokmakov, S. Zakharov, and C. V ondrick. Zero-1-to-3: Zero-shot one image to 3d object. In�������� ������������� ���������� �� �������� ������� ���� ����� ������ ������� ������� ���� ����, pp. 9264–9275. IEEE, 2023. doi: 10.1109/ICCV51070.2023.00853 1

  18. [29]

    Liu and L

    V . Liu and L. B. Chilton. Design guidelines for prompt engineering text-to- image generative models. In S. D. J. Barbosa, C. Lampe, C. Appert, D. A. Shamma, S. M. Drucker, J. R. Williamson, and K. Yatani, eds.,��� ���� ��� ���������� �� ����� ������� �� ��������� �������� ���...

  19. [30]

    V . Liu, J. Vermeulen, G. W. Fitzmaurice, and J. Matejka. 3dall-e: Integrat- ing text-to-image AI in 3d design workflows. In D. Byrne, N. Martelaro, A. Boucher, D. J. Chatting, S. F. Alaoui, S. E. Fox, I. Nicenboim, and C. MacArthur, eds., ����������� �� ��� ���� ��� ���������...

  20. [32]

    Marks, B

    J. Marks, B. Andalman, P. A. Beardsley, W. T. Freeman, S. F. Gibson, J. K. Hodgins, T. Kang, B. Mirtich, H. Pfister, W. Ruml, K. Ryall, J. E. Seims, and S. M. Shieber. Design galleries: a general approach to setting param- eters for computer graphics and animation. In G. S. Ow...

  21. [33]

    S. R. Midway. Principles of effective data visualization. ��������, 1(9):100141, 2020. doi: 10.1016/j.patter.2020.100141 9

  22. [34]

    Mildenhall, P

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In A. Vedaldi, H. Bischof, T. Brox, and J. Frahm, eds.,�������� ������ � ���� ���� � ���� �������� ����������� �������� ...

  23. [35]

    Nichol, H

    A. Nichol, H. Jun, P. Dhariwal, P. Mishkin, and M. Chen. Point-e: A system for generating 3d point clouds from complex prompts. ���� , abs/2212.08751, 2022. doi: 10.48550/ARXIV.2212.08751 2

  24. [36]

    A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen. GLIDE: towards photorealistic image gen- eration and editing with text-guided diffusion models. In K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvári, G. Niu, and S. Sabato, eds., �...

  25. [37]

    Oppenlaender

    J. Oppenlaender. Prompt engineering for text-based generative art. ���� , abs/2204.13988, 2022. doi: 10.48550/ARXIV.2204.13988 2

  26. [38]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe. Training language models to follow instructio...

  27. [39]

    Pandey, F

    K. Pandey, F. Chevalier, and K. Singh. Juxtaform: interactive visual sum- marization for exploratory shape design. ��� ������ ������, 42(4):52:1– 52:14, 2023. doi: 10.1145/3592436 2

  28. [40]

    Podell, Z

    D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach. SDXL: improving latent diffusion models for high-resolution image synthesis. 2024. 4

  29. [41]

    Poole, A

    B. Poole, A. Jain, J. T. Barron, and B. Mildenhall. Dreamfusion: Text- to-3d using 2d diffusion. In ��� �������� ������������� ���������� �� �������� ���������������� ���� ����� ������� ������� ��� ���� ����. OpenReview.net, 2023. 2

  30. [43]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision. In M. Meila and T. Zhang, eds., ����������� �� ��� ���� ����������...

  31. [44]

    Ramesh, P

    A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen. Hier- archical text-conditional image generation with CLIP latents. ���� , abs/2204.06125, 2022. doi: 10.48550/ARXIV.2204.06125 2

  32. [45]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High- resolution image synthesis with latent diffusion models. In �������� ���������� �� �������� ������ ��� ������� ������������ ���� ����� ��� �������� ��� ���� ���� ������ ����, pp. 10674–10685. IEEE, 2022. doi: 1...

  33. [46]

    Y . Shi, P. Wang, J. Ye, L. Mai, K. Li, and X. Yang. Mvdream: Multi-view diffusion for 3d generation. In ��� ������� ������������� ���������� �� �������� ���������������� ���� ����� ������� �������� ��� ����� ����. OpenReview.net, 2024. 3

  34. [47]

    T. Shin, Y . Razeghi, R. L. L. IV , E. Wallace, and S. Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. In B. Webber, T. Cohn, Y . He, and Y . Liu, eds.,����������� �� ��� ���� ���������� �� ��������� ������� �� ������� ��������...

  35. [48]

    K. Son, D. Choi, T. S. Kim, Y . Kim, and J. Kim. Genquery: Support- ing expressive visual search with generative models. In F. F. Mueller, P. Kyburz, J. R. Williamson, C. Sas, M. L. Wilson, P. O. T. Dugas, and I. Shklovski, eds., ����������� �� ��� ��� ���������� �� ����� ����...

  36. [50]

    Swearngin, C

    A. Swearngin, C. Wang, A. Oleson, J. Fogarty, and A. J. Ko. Scout: Rapid exploration of interface layout alternatives through high-level de- sign constraints. In R. Bernhaupt, F. F. Mueller, D. Verweij, J. Andres, J. McGrenere, A. Cockburn, I. Avellino, A. Goguey, P. Bjøn, S. ...

  37. [51]

    J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu. LGM: large multi-view gaussian model for high-resolution 3d content creation. In A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, eds., �������� ������ � ���� ���� � ���� �������� ����������� ��...

  38. [52]

    J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng. Dreamgaussian: Generative gaussian splatting for efficient 3d content creation. In ��� ������� ������ �������� ���������� �� �������� ���������������� ���� ����� ������� �������� ��� ����� ����. OpenReview.net, 2024. 1

  39. [53]

    Z. Tang, S. Gu, C. Wang, T. Zhang, J. Bao, D. Chen, and B. Guo. V olumed- iffusion: Flexible text-to-3d generation with efficient volumetric encoder. ���� , abs/2312.11459, 2023. doi: 10.48550/ARXIV.2312.11459 2

  40. [54]

    the divergence and bhattacharyya distance measures in signal selection

    G. T. Toussaint. Comments on "the divergence and bhattacharyya distance measures in signal selection". ���� ������ �������, 20(3):485, 1972. doi: 10.1109/TCOM.1972.1091157 5

  41. [55]

    B. Wang. Democratizing content creation and consumption through human-ai copilot systems. In S. Follmer, J. Han, J. Steimle, and N. H. Riche, eds., ������� ����������� �� ��� ���� ������ ��� ��������� �� ���� ��������� �������� ��� ����������� ���� ����� ��� ���������� ��� ���...

  42. [56]

    H. Wang, X. Du, J. Li, R. A. Yeh, and G. Shakhnarovich. Score jaco- 11 bian chaining: Lifting pretrained 2d diffusion models for 3d generation. In �������� ���������� �� �������� ������ ��� ������� ������������ ���� ����� ���������� ��� ������� ���� ������ ����, pp. 12619–1262...

  43. [57]

    X. Wang, J. He, Z. Jin, M. Yang, Y . Wang, and H. Qu. M2lens: Visualizing and explaining multimodal models for sentiment analysis. ���� ������ ���� ������� ������, 28(1):802–812, 2022. doi: 10.1109/TVCG.2021. 3114794 2

  44. [58]

    Z. J. Wang, E. Montoya, D. Munechika, H. Yang, B. Hoover, and D. H. Chau. Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models. In A. Rogers, J. L. Boyd-Graber, and N. Okazaki, eds., ����������� �� ��� ���� ������ ������� �� ��� ����������� ���...

  45. [59]

    T. Wu, E. Jiang, A. Donsbach, J. Gray, A. Molina, M. Terry, and C. J. Cai. Promptchainer: Chaining large language model prompts through visual programming. In S. D. J. Barbosa, C. Lampe, C. Appert, and D. A. Shamma, eds., ��� ���� ��� ���������� �� ����� ������� �� ��������� �...

  46. [61]

    J. Xia, L. Huang, W. Lin, X. Zhao, J. Wu, Y . Chen, Y . Zhao, and W. Chen. Interactive visual cluster analysis by contrastive dimensionality reduction. ���� ������ ���� ������� ������, 29(1):734–744, 2023. doi: 10.1109/ TVCG.2022.3209423 2

  47. [62]

    L. Yang, C. Xiong, J. K. Wong, A. Wu, and H. Qu. Explaining with examples: Lessons learned from crowdsourced introductory description of information visualizations. ���� ������ ���� ������� ������, 29(3):1638– 1650, 2023. doi: 10.1109/TVCG.2021.3128157 7

  48. [63]

    H. Zeng, X. Wang, Y . Wang, A. Wu, T.-C. Pong, and H. Qu. Gesturelens: Visual analysis of gestures in presentation videos. ���� ������������ �� ������������� ��� �������� ��������, 2022. doi: 10.1109/TVCG.2022. 3169175 2

  49. [64]

    H. Zeng, X. Wang, A. Wu, Y . Wang, Q. Li, A. Endert, and H. Qu. Emoco: Visual analysis of emotion coherence in presentation videos. ���� ������ ������� �� ������������� ��� �������� ��������, 26(1):927–937, 2019. doi: 10.1109/TVCG.2019.2934656 2

  50. [65]

    X. Zeng, A. Vahdat, F. Williams, Z. Gojcic, O. Litany, S. Fidler, and K. Kreis. LION: latent point diffusion models for 3d shape generation. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, eds., �������� �� ������ ����������� ���������� ������� ��� ������...

  51. [66]

    Zhang, Y

    B. Zhang, Y . Cheng, J. Yang, C. Wang, F. Zhao, Y . Tang, D. Chen, and B. Guo. Gaussiancube: Structuring gaussian splatting using optimal transport for 3d generative modeling. ���� , abs/2403.19655, 2024. doi: 10.48550/ARXIV.2403.19655 2

  52. [67]

    Z. Zhao, W. Liu, X. Chen, X. Zeng, R. Wang, P. Cheng, B. Fu, T. Chen, G. Yu, and S. Gao. Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, eds., �������...

  53. [68]

    H. Zhu, M. Zhu, Y . Feng, D. Cai, Y . Hu, S. Wu, X. Wu, and W. Chen. Visualizing large-scale high-dimensional data via hierarchical embedding of knn graphs. ������ �����������, 5(2):51–59, 2021. doi: 10.1016/j.visinf .2021.06.002 2 12

  54. [2023]

    doi: 10.18653/V1/2023.ACL-LONG.51 3

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.