Pith. sign in

REVIEW 3 major objections 5 minor 130 references

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read ASV3D lets one extra image improve single-view 3D reconstruction by choosing the best condition for each generated view.

desk verdict A genuinely new per-view conditioning gate for adding an unposed extra image to single-view 3D reconstruction, with a real but narrow validation gap around gate correctness. read the letter →

arxiv 2608.08132 v1 pith:27M54SNX submitted 2026-08-08 cs.CV cs.GR

classification cs.CVcs.GR
keywords test-timeadaptationsingle-view3Dreconstructionmulti-viewgenerationconsistency-basedgatingcontrastivelearningdiffusionmodelszero-shotadditionalimage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ASV3D is a framework that adapts a pre-trained single-view 3D reconstruction model to test time by supplying one additional image of the same object, captured by a different camera and in a different context. It addresses the question of when an extra view helps: rather than fusing all images into one condition, ASV3D uses a consistency-based gate to choose, for each target view, whichever conditioning image the diffusion model denoises most consistently. The zero-shot variant does this without retraining; the optimised variant further fine-tunes the U-Net's attention layers with a denoising loss and a contrastive loss that pulls views generated from the additional image closer to that image. The authors report consistent gains over Wonder3D and Era3D on the GSO benchmark, and on a small real-world dataset, with the optimised Wonder3D reducing Chamfer distance from 0.0218 to 0.0120. The aim is to show that a single second image can resolve some of the ambiguity of single-view reconstruction without pose estimation or retraining from scratch.

What carries the argument

The consistency-based gate f(n) defined in Eq. (4) is the load-bearing mechanism: for each target view index n, it picks the condition c in {y,z} that minimises the expected squared deviation of the denoising-mean prediction mu from its own time-average over the late half of the diffusion trajectory, T = {T/2, ..., 1}. This is a pose-free proxy for which input image is most informative for that view, replacing the need to estimate camera poses or fuse conditions. In the optimised variant, a contrastive loss (NT-Xent, Eq. 9-10) anchored on the additional image z is added to the denoising loss, and only the cross- and self-attention layers of the U-Net are updated.

What would settle it

Take objects from a category outside GSO, such as glassware or symmetric toys, and for each target view compute both the gate's chosen condition and the actual reconstruction error under both conditions. If the gate's selections do not correlate with the lower-error condition, or if the Spearman correlation between gate consistency and generation quality drops below its GSO values, the proxy collapses.

Watch

Extended reading notes

Core claim

The paper's central claim is that a single additional image of the same object, even from an unposed, different camera, can be folded into an existing single-view generative reconstruction pipeline to improve both multi-view generation and the final 3D shape. The mechanism is a camera-pose-free consistency gate: for each target view n, it evaluates the variance of the diffusion model's denoising-mean predictions µ over the late half of the denoising trajectory under each candidate condition, and selects the condition with the smaller spread. Under the optimised strategy, the model is then adapted by minimising a denoising loss plus a contrastive (NT-Xent) loss that anchors views generated with the additional image to that image. On Google Scanned Objects, the optimised Wonder3D variant makes Chamfer distance 0.0120, IoU 0.5738, and improves PSNR/SSIM/LPIPS over the baseline; a 32-participant user study rates ASV3D higher on both 3D reconstruction and multi-view generation. The authors position this as the first method to adapt single-view reconstruction with additional imagery.

Load-bearing premise

The load-bearing premise is that the consistency-based gate of Eq. (4), which equates the variance of the model's denoising-mean predictions over the late half of the trajectory with the reliability of a conditioning image, is a valid proxy for which input produces the more accurate target view.

Editorial extensions

If this is right

  • Any existing conditional generative single-view reconstruction model can likely benefit from the same two-stage adaptation, since the gate only uses the model's own denoising predictions.
  • Because the gate decides per target view, the framework naturally handles cases where the extra view is helpful for some views (e.g., occluded back views) and redundant or misleading for others (e.g., views close to the primary image).
  • The optimised variant's reliance on contrastive learning to pull z-conditioned views toward z should transfer to improving cross-view consistency in other multi-view diffusion pipelines.
  • If the gate correctly identifies informative conditions, the method could be used to tell a user which additional view to acquire for a given object, a direction the authors mention as future work.
  • The improvements are obtained without pose estimation or calibration, opening test-time adaptation for casual multi-view inputs that are both unposed and captured under different conditions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • We infer that the gate's Spearman correlation with generation quality, verified only on GSO, may not hold for objects with strong symmetries or specular surfaces, where small denoising variance could arise from a confident but wrong condition; testing the gate on such categories would be a direct stress test.
  • Beyond the paper's explicit claim, the method suggests a simple acquisition policy: when generating a target view, the gate could rank candidate supplementary images by their consistency scores and pick the best one from a pool, turning the binary choice into a selection among many views.
  • We also conjecture that the contrastive loss could be replaced or augmented by attention-based feature alignment between the generated views and the additional image, which might reduce the need for a separate image encoder; this is an editorial suggestion, not a claim of the paper.
  • The reported real-world evaluation is qualitative, so a natural next step is a quantitative evaluation with ground-truth scans of objects in natural settings, which the paper does not provide.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes ASV3D, a framework to adapt pre-trained single-view 3D reconstruction models when an additional image of the same object is supplied. Two adaptation strategies are introduced: a zero-shot scheme that uses a consistency-based gate (Eq. 4) to select, per target view, whether the primary or the additional image conditions the diffusion generation, and an optimised scheme that further updates the U-Net with a denoising loss and a contrastive loss (Eqs. 8-11). The method is applied to Wonder3D and Era3D and evaluated on GSO and a small real-world dataset, with quantitative metrics for 3D reconstruction (CD, IoU) and multi-view generation (PSNR, SSIM, LPIPS), plus a user study. The reported results show consistent improvements over the baselines, with the optimised Wonder3D variant achieving the best numbers in both reconstruction and generation.

Significance. If the claimed gating mechanism reliably identifies which input image is more informative for each target view, the paper offers a practical and camera-pose-free way to leverage extra imagery for single-view 3D reconstruction, which is a relevant and timely problem. The zero-shot variant is particularly attractive because it requires no retraining, and the optimised variant demonstrates a promising direction for test-time adaptation. The paper also releases code and a real-world dataset, which are concrete contributions that would benefit the community. However, the central evidence for the gate is currently correlational rather than decision-oriented, and the headline quantitative results inherit this gap; the state-of-the-art claim is therefore not yet fully supported. The core idea is interesting and the experiments are extensive enough that the deficiencies are fixable within the scope of a revision.

major comments (3)
  1. [Sec. 3.1, Eq. (4) and Appendix A] The consistency-based gate f chooses, for each target view n, the condition c ∈ {y,z} that minimizes the dispersion of the denoising predictions over the late half of the diffusion trajectory. This gate is the core mechanism that converts the additional image into a reconstruction benefit: it determines the condition used in Eq. (5) for every view, and it also determines which views receive the contrastive loss in Eq. (10). The validation in Appendix A, however, only reports Spearman correlations between the consistency score g(c) and image-quality metrics pooled over all (condition, view) pairs on GSO (Figure 8). A pooled correlation does not establish that argmin_c g(c) selects the better condition for a specific view; a condition that consistently yields low dispersion but poor fidelity would win the argmin despite being the wrong choice. The paper does not report per-view binary selection accuracy against an oracle, object-level cross-validation, or any test of the gate on held-out objects or real-world data. Table 3 shows that the gate beats a multi-conditioning baseline, but that comparison does not separate the gate's decision accuracy from the general benefit of selective conditioning. Consequently, the zero-shot improvements in Table 1 and the condition assignments used in the optimised variant are not convincingly shown to arise from the gate correctly identifying which input image is more informative for each target view, which is the central claim of the paper.
  2. [Sec. 4.2 and Sec. 4.4] The hyperparameters of the method—the contrastive weight λ (set to 0.2), the temperature τ (set to 0.07), the number of diffusion steps T (set to 50), and the late-half step range T in Eq. (4)—are all selected on the same GSO benchmark that is used for the main quantitative evaluation. No cross-validation on held-out objects or a separate validation set is reported, and no sensitivity analysis is given for τ, T, or the late-half range. Because the paper frames the method as a test-time adaptation scheme, the absence of such analysis leaves open the possibility that the reported gains rely on hyperparameters that have been tuned to the evaluation distribution. At least a sensitivity table or a statement about the stability of the results under reasonable hyperparameter variations is needed for the claims in Table 1 to be fully load-bearing.
  3. [Sec. 4.5 and Figure 7] The user study is reported with mean ratings and standard deviations, but no significance testing is performed, and the standard-deviation values in Figure 7 appear to contradict the claim in the text that ASV3D has 'lower standard deviations' than Wonder3D: the figure shows ASV3D multi-view generation with ±0.53 and Wonder3D with ±0.12, which would imply the opposite. Please clarify the figure or correct the text, and add appropriate statistical tests (e.g., paired t-test or Wilcoxon signed-rank) on the ratings or the forced-choice preferences.
minor comments (5)
  1. [Tables 1, 3, 4] The tables report point estimates without error bars or significance tests; given that only 30 GSO objects are used, per-object standard deviations or confidence intervals would help assess the stability of the differences.
  2. [Figures 4 and 10] There is a recurring typo: 'FreeSplater' should be 'FreeSplatter'.
  3. [Sec. 4.1] The description of how the additional image is rendered for GSO ("another image in a random view") is underspecified; please state whether the view is uniformly sampled, whether it is always a different azimuth, and whether any objects are excluded due to near-duplicate views.
  4. [Eq. (4)] The symbol T is used both for the total number of diffusion steps and for the set of late-half step indices; using a different symbol for the set (e.g., T_late) would avoid confusion.
  5. [Real-world evaluation] The real-world evaluation is qualitative only (10 objects), which is understandable because ground-truth 3D is unavailable; however, this limitation should be stated more prominently in the main text rather than only in the supplementary.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central ASV3D result is validated against external GSO ground truth and does not reduce to the gate or to self-citations.

full rationale

The paper's core derivation is Eq. (4)-(11): a consistency-based gate selects per-view conditions, then either zero-shot generation or optimised adaptation (denoising plus contrastive losses) produces views for 3D reconstruction. No equation makes the output equal to the input by construction. The gate is a heuristic defined on the frozen model's denoising means; its utility is assessed in Appendix A by Spearman correlation with image quality and in Table 3 by reconstruction metrics against the multi-conditioning baseline, so it is not fitted to the final reconstruction target. The optimised variant does train on the model's own initial generated views, which is self-referential in architecture, but Table 1 reports CD/IoU against GSO ground-truth meshes and PSNR/SSIM/LPIPS against ground-truth renders, so the claimed improvement is externally anchored. The only overlapping-author citation (Shum et al. 2025, co-authored by D. T. Nguyen) supports a general statement about diffusion models and is not load-bearing. The weak per-view validation of the gate (pooled correlations rather than binary selection accuracy) is a robustness and correctness concern, not circularity under the criteria here.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The method's central claim depends on a test-time auxiliary image, on a heuristic gate that is validated empirically, and on the ability to reuse a frozen diffusion backbone with swapped conditions. The only fitted constants of consequence are the contrastive weight and the late-half trajectory range, both selected on the GSO benchmark.

free parameters (4)
  • lambda (contrastive loss weight) = 0.2
    Eq. (11). Chosen via ablation on GSO (Table 4), comparing 0.1 and 0.2; final metrics use the same benchmark, so this is fitted to the evaluation set.
  • Late-half diffusion-step range for gate = {T/2, ..., 1}
    Eq. (4) and Appendix A; the authors state the late-half setting 'perform[s] best' and validate it on GSO, making it a hand-selected design choice.
  • tau (contrastive temperature) = 0.07
    Eq. (9); taken from SimCLR (Chen et al. 2020), not tuned here, but still a chosen hyperparameter.
  • Diffusion steps T = 50
    Implementation detail in Section 4.2; affects runtime and quality, not a claim-specific fitted constant.
assumptions (4)
  • domain assumption An additional image of the same object instance is available at test time and is captured by a different camera/environment without known pose.
    The entire problem setting assumes this (Section 3.1); no mechanism validates that the extra image depicts the same object.
  • domain assumption The variance of the model's denoising-mean predictions over late diffusion steps for a candidate condition is a reliable proxy for the quality of views generated with that condition.
    Eq. (4) defines the gate on this premise; Appendix A reports a positive Spearman correlation on GSO, but the premise is not proven for out-of-distribution objects.
  • domain assumption A pre-trained single-view diffusion U-Net can be re-used with a per-view swapped conditioning image, without retraining the backbone, and still produce a coherent multi-view set.
    Section 3.2 applies epsilon_theta with f(n) in {y,z}; the model was never trained to mix conditions across views.
  • standard math Standard DDPM definitions and the pretrained checkpoints of Wonder3D/Era3D behave as described.
    Eqs. (1)-(3) import Ho et al. (2020); no modification to the diffusion formalism.

how reviews work

0 comments
Cite this review

Pith. "Pith review of When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery." pith.science (2026). https://pith.science/paper/27M54SNX

@misc{pith2026260808132,
  author       = {Pith},
  title        = {Pith review of: When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/27M54SNX}},
  note         = {Machine review of arXiv:2608.08132}
}
read the original abstract

Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical information from viewpoints to complete 3D structures. Using an additional view may help to resolve the issue. However, there is no mechanism that can integrate the extra view into the single-view 3D reconstruction principle. We address this challenge by proposing ASV3D, a framework for adapting single-view 3D object reconstruction to test-time data with support from one additional image. We introduce two adaptation strategies: (i) a zero-shot adaptation scheme that leverages the auxiliary image to improve the reconstruction quality of an object without retraining, and (ii) an optimised adaptation scheme that further enhances visual fidelity and cross-view consistency via contrastive learning. We apply our ASV3D to improve two state-of-the-art single-view 3D reconstruction pipelines on both benchmark and real-world datasets. Results demonstrate that our approach consistently improves reconstruction accuracy and robustness under unconstrained multi-view inputs, outperforming the baselines in both quantitative metrics and human preference. We publish our code and the real-world object dataset in our project page at https://github.com/YNhuHuynh/ASV3D/tree/main.

Figures

Figures reproduced from arXiv: 2608.08132 by the authors.

Figure 1
Figure 1. ASV3D vs. Wonder3D. (a) Wonder3D (Long et al. 2024) uses the same input image to condition the generation of all target views. (b) Our ASV3D drives the input and addi￾tional images into proper generations, highlighted in corre￾sponding colours (green: input image, red: additional one). • We propose ASV3D, a framework that adapts single-view generative 3D reconstruction models using a primary im￾age together with an … view at source ↗
Figure 2
Figure 2. Overview of ASV3D. Given an input image y, we apply a pre-trained generative model (e.g., Wonder3D) to generate an initial set of target views. We then adapt the model to an additional image z once supplied. Our method supports two adaptation strategies: zero-shot and optimised adaptation. The zero-shot setting determines the optimal condition for each target view using a consistency-based gate. The optimised adapta… view at source ↗
Figure 3
Figure 3. Qualitative comparison of ASV3D with Won￾der3D baseline in 3D reconstruction (a) and in multi-view image generation (b). Artifacts in the reconstruction and gen￾eration results are highlighted. Qualitative results [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (8 more)
Figure 5
Figure 5. Figure 5: Qualitative comparison of ASV3D (optimised version) with FreeSplatter under an extreme pose gap. Left: Input image in the front view (top) and additional image in an arbitrary view (bottom). Middle: Multi-view images. Right: 3D reconstruction results [PITH_FULL_IMAGE:…
Figure 6
Figure 6. Figure 6: Qualitative results of ASV3D (optimised ver￾sion) on real-world objects. (Left) Input front and auxiliary images. (Middle) Reconstructed mesh and rendered texture. (Right) Generated views. As shown, ASV3D yields consis￾tent geometry, accurate texture, and stable multi-…
Figure 7
Figure 7. Figure 7: summarises user preference. As shown, ASV3D is consistently favoured to Wonder3D in both tasks, evident by higher mean scores and lower standard deviations. Details of this study are presented in the supplementary material. ASV3D Wonder3D 3D Reconstruction Multi-View G…
Figure 8
Figure 8. Figure 8: As shown in the results, our proposed consistency [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 8
Figure 8. Figure 8: Validation of consistency-based gating. Among the tested measurements, our proposed consistency-based gate with late half diffusion steps shows the strongest observed pooled association with poorer condition quality for both negative PSNR and LPIPS. Error bars denote o…
Figure 9
Figure 9. Figure 9: Background removal on real-world inputs. Original images from our collected dataset (left) and foreground segmentation results produced by Qin et al. (2020) (right) [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]
Figure 10
Figure 10. Figure 10: Qualitative comparisons on real-world objects. Left: primary front-view image (top) and auxiliary image from an arbitrary viewpoint (bottom). Middle: generated multi-view images. Right: reconstructed 3D geometry. For ASV3D we present results of the optimised version …
Figure 11
Figure 11. Figure 11: Example user-study interfaces for evaluating (a) 3D reconstruction and (b) multi-view image generation. Method identities and display positions were randomised during the study [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

130 extracted references · 39 canonical work pages

  1. [1]

    Yihong Luo and Tianyang Hu and Jiacheng Sun and Yujun Cai and Jing Tang , title =

  2. [2]

    Davison , title =

    Xin Kong and Shikun Liu and Xiaoyang Lyu and Marwan Taher and Xiaojuan Qi and Andrew J. Davison , title =

  3. [3]

    Chenglizhao Chen and Ziyue Xue and Longyan Yang and Zhenyu Wu and Shanchen Pang and Hong Qin , title =

  4. [4]

    International Conference on Machine Learning , pages =

    Yang Song and Prafulla Dhariwal and Mark Chen and Ilya Sutskever , title =. International Conference on Machine Learning , pages =

  5. [5]

    Black-Box Test-Time Shape

    Brandon Leung and Chih. Black-Box Test-Time Shape

  6. [6]

    British Machine Vision Conference , year =

    Kim Yu. British Machine Vision Conference , year =

  7. [7]

    Advances in Neural Information Processing Systems , year =

    Yuheng Yuan and Qiuhong Shen and Shizun Wang and Xingyi Yang and Xinchao Wang , title =. Advances in Neural Information Processing Systems , year =

  8. [8]

    International Conference on Learning Representations , year =

    Xingyu Chen and Yue Chen and Yuliang Xiu and Andreas Geiger and Anpei Chen , title =. International Conference on Learning Representations , year =

Show all 130 references
  1. [9]

    Shangzhan Zhang and Jianyuan Wang and Yinghao Xu and Nan Xue and Christian Rupprecht and Xiaowei Zhou and Yujun Shen and Gordon Wetzstein , title =

  2. [10]

    No Pose, No Problem: Surprisingly Simple

    Botao Ye and Sifei Liu and Haofei Xu and Xueting Li and Marc Pollefeys and Ming. No Pose, No Problem: Surprisingly Simple. International Conference on Learning Representations , year =

  3. [11]

    Richter and Stefan Roth , title =

    Stephan R. Richter and Stefan Roth , title =

  4. [12]

    Kim and Bryan C

    Thibault Groueix and Matthew Fisher and Vladimir G. Kim and Bryan C. Russell and Mathieu Aubry , title =

  5. [13]

    Maxim Tatarchenko and Alexey Dosovitskiy and Thomas Brox , title =

  6. [14]

    Advances in Neural Information Processing Systems , year =

    Jiajun Wu and Chengkai Zhang and Tianfan Xue and Bill Freeman and Josh Tenenbaum , title =. Advances in Neural Information Processing Systems , year =

  7. [15]

    Fouhey and Mikel Rodriguez and Abhinav Gupta , title =

    Rohit Girdhar and David F. Fouhey and Mikel Rodriguez and Abhinav Gupta , title =. European Conference Computer Vision , pages =

  8. [16]

    Choy and Danfei Xu and JunYoung Gwak and Kevin Chen and Silvio Savarese , title =

    Christopher B. Choy and Danfei Xu and JunYoung Gwak and Kevin Chen and Silvio Savarese , title =. European Conference Computer Vision , pages =

  9. [17]

    Learning Category-Specific Deformable

    Shubham Tulsiani and Abhishek Kar and Jo. Learning Category-Specific Deformable

  10. [18]

    Oswald and Eno T

    Martin R. Oswald and Eno T. Fast and globally optimal single view reconstruction of curved objects , booktitle =

  11. [19]

    Chung and Andrew Y

    Ashutosh Saxena and Sung H. Chung and Andrew Y. Ng , title =. Advances in Neural Information Processing Systems , year =

  12. [20]

    Barron and Ben Mildenhall and Matthew Tancik and Peter Hedman and Ricardo Martin

    Jonathan T. Barron and Ben Mildenhall and Matthew Tancik and Peter Hedman and Ricardo Martin

  13. [21]

    Srinivasan and Matthew Tancik and Jonathan T

    Ben Mildenhall and Pratul P. Srinivasan and Matthew Tancik and Jonathan T. Barron and Ravi Ramamoorthi and Ren Ng , title =. Communications of the. 2022 , url =. doi:10.1145/3503250 , timestamp =

  14. [22]

    Language-driven Object Fusion into Neural Radiance Fields with Pose-Conditioned Dataset Updates , booktitle =

    Ka. Language-driven Object Fusion into Neural Radiance Fields with Pose-Conditioned Dataset Updates , booktitle =

  15. [23]

    Advances in Neural Information Processing Systems , year =

    Peng Li and Yuan Liu and Xiaoxiao Long and Feihu Zhang and Cheng Lin and Mengfei Li and Xingqun Qi and Shanghang Zhang and Wei Xue and Wenhan Luo and Ping Tan and Wenping Wang and Qifeng Liu and Yike Guo , title =. Advances in Neural Information Processing Systems , year =

  16. [24]

    Advances in Neural Information Processing Systems , year =

    Daniel Pfrommer and Zehao Dou and Christopher Scarvelis and Max Simchowitz and Ali Jadbabaie , title =. Advances in Neural Information Processing Systems , year =

  17. [25]

    Advances in Neural Information Processing Systems , year =

    Giannis Daras and Yuval Dagan and Alex Dimakis and Constantinos Daskalakis , title =. Advances in Neural Information Processing Systems , year =

  18. [26]

    Viet Nguyen and Giang Vu and Tung Nguyen Thanh and Khoat Than and Toan Tran , title =

  19. [27]

    Advances in Neural Information Processing Systems , year =

    Jonathan Ho and Ajay Jain and Pieter Abbeel , title =. Advances in Neural Information Processing Systems , year =

  20. [28]

    Yunhan Yang and Yukun Huang and Xiaoyang Wu and Yuan

  21. [29]

    Chenjie Cao and Chaohui Yu and Shang Liu and Fan Wang and Xiangyang Xue and Yanwei Fu , title =

  22. [30]

    Advances in Neural Information Processing Systems , pages =

    Peng Wang and Lingjie Liu and Yuan Liu and Christian Theobalt and Taku Komura and Wenping Wang , title =. Advances in Neural Information Processing Systems , pages =

  23. [31]

    A Survey on Quality Metrics for Text-to-Image Generation , journal =

    Sebastian Hartwig and Dominik Engel and Leon Sick and Hannah Kniesel and Tristan Payer and Poonam Poonam and Michael Gl. A Survey on Quality Metrics for Text-to-Image Generation , journal =

  24. [32]

    Hinton , title =

    Ting Chen and Simon Kornblith and Mohammad Norouzi and Geoffrey E. Hinton , title =. International Conference on Machine Learning , pages =

  25. [33]

    Guangcong Wang and Zhaoxi Chen and Chen Change Loy and Ziwei Liu , title =

  26. [34]

    Structure-from-Motion Revisited , booktitle=

    Sch\". Structure-from-Motion Revisited , booktitle=

  27. [35]

    Russell and P

    S. Russell and P. Norvig , title =

  28. [36]

    Color Alignment in Diffusion , booktitle =

    Ka. Color Alignment in Diffusion , booktitle =

  29. [37]

    High-Resolution Image Synthesis with Latent Diffusion Models , booktitle =

    Robin Rombach and Andreas Blattmann and Dominik Lorenz and Patrick Esser and Bj. High-Resolution Image Synthesis with Latent Diffusion Models , booktitle =

  30. [38]

    CoRR , volume =

    Tanveer Younis and Zhanglin Cheng , title =. CoRR , volume =

  31. [39]

    Journal of Real Time Image Processing , volume =

    Matteo Bortolon and Luca Bazzanella and Fabio Poiesi , title =. Journal of Real Time Image Processing , volume =

  32. [40]

    Li, Dongxu and Li, Junnan and Hoi, Steven , urldate =

  33. [41]

    Ruiz, Nataniel and Li, Yuanzhen and Jampani, Varun and Pritch, Yael and Rubinstein, Michael and Aberman, Kfir , year =

  34. [42]

    Conditional Text Image Generation With Diffusion Models , abstract =

    Zhu, Yuanzhi and Li, Zhaohai and Wang, Tianwei and He, Mengchao and Yao, Cong , langid =. Conditional Text Image Generation With Diffusion Models , abstract =

  35. [43]

    Advances in Neural Information Processing Systems , author =

    Denoising Diffusion Probabilistic Models , pages =. Advances in Neural Information Processing Systems , author =

  36. [44]

    Diffusion Models Beat

    Dhariwal, Prafulla and Nichol, Alexander , urldate =. Diffusion Models Beat. Advances in Neural Information Processing Systems , publisher =

  37. [46]

    and Mildenhall, Ben , urldate =

    Poole, Ben and Jain, Ajay and Barron, Jonathan T. and Mildenhall, Ben , urldate =

  38. [47]

    Interactive Sketch & Fill: Multiclass Sketch-to-Image Translation , rights =

    Ghosh, Arnab and Zhang, Richard and Dokania, Puneet and Wang, Oliver and Efros, Alexei and Torr, Philip and Shechtman, Eli , urldate =. Interactive Sketch & Fill: Multiclass Sketch-to-Image Translation , rights =. 2019. doi:10.1109/ICCV.2019.00126 , shorttitle =

  39. [48]

    It's All About Your Sketch: Democratising Sketch Control in Diffusion Models , url =

    Koley, Subhadeep and Bhunia, Ayan Kumar and Sekhri, Deeptanshu and Sain, Aneeshan and Chowdhury, Pinaki Nath and Xiang, Tao and Song, Yi-Zhe , urldate =. It's All About Your Sketch: Democratising Sketch Control in Diffusion Models , url =

  40. [49]

    Text-to-Image Diffusion Models are Great Sketch-Photo Matchmakers , abstract =

    Koley, Subhadeep and Bhunia, Ayan Kumar and Sain, Aneeshan and Chowdhury, Pinaki Nath and Xiang, Tao and Song, Yi-Zhe , langid =. Text-to-Image Diffusion Models are Great Sketch-Photo Matchmakers , abstract =

  41. [50]

    An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion , url =

    Gal, Rinon and Alaluf, Yuval and Atzmon, Yuval and Patashnik, Or and Bermano, Amit Haim and Chechik, Gal and Cohen-or, Daniel , urldate =. An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion , url =

  42. [52]

    Proceedings of the 40th International Conference on Machine Learning , publisher =

    Li, Junnan and Li, Dongxu and Savarese, Silvio and Hoi, Steven , urldate =. Proceedings of the 40th International Conference on Machine Learning , publisher =

  43. [53]

    Large Scale

    Brock, Andrew and Donahue, Jeff and Simonyan, Karen , urldate =. Large Scale

  44. [54]

    and Pouget-Abadie, Jean and Mirza, Mehdi and Xu, Bing and Warde-Farley, David and Ozair, Sherjil and Courville, Aaron and Bengio, Yoshua , urldate =

    Goodfellow, Ian J. and Pouget-Abadie, Jean and Mirza, Mehdi and Xu, Bing and Warde-Farley, David and Ozair, Sherjil and Courville, Aaron and Bengio, Yoshua , urldate =. Generative Adversarial Nets , volume =. Advances in Neural Information Processing Systems , publisher =

  45. [61]

    Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion , abstract =

    Lu, Yuanxun and Zhang, Jingyang and Li, Shiwei and Fang, Tian and. Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion , abstract =

  46. [65]

    Mantecon, Héctor Laria and Wang, Kai and Weijer, Joost van de and Raducanu, Bogdan and Wang, Yaxing , urldate =

  47. [70]

    Peng, Songyou and Genova, Kyle and Jiang, Chiyu and Tagliasacchi, Andrea and Pollefeys, Marc and Funkhouser, Thomas , urldate =. 2023. doi:10.1109/CVPR52729.2023.00085 , shorttitle =

  48. [74]

    Visual Chain-of-Thought Diffusion Models , url =

    Harvey, William and Wood, Frank , urldate =. Visual Chain-of-Thought Diffusion Models , url =

  49. [75]

    and Cao, Yuan and Narasimhan, Karthik R

    Yao, Shunyu and Yu, Dian and Zhao, Jeffrey and Shafran, Izhak and Griffiths, Thomas L. and Cao, Yuan and Narasimhan, Karthik R. , urldate =. Tree of Thoughts: Deliberate Problem Solving with Large Language Models , url =

  50. [78]

    Visual Programming: Compositional Visual Reasoning Without Training , url =

    Gupta, Tanmay and Kembhavi, Aniruddha , urldate =. Visual Programming: Compositional Visual Reasoning Without Training , url =

  51. [83]

    , urldate =

    Lei, Chao and Lipovetzky, Nir and Ehinger, Krista A. , urldate =. Generalized Planning for the Abstraction and Reasoning Corpus , url =

  52. [84]

    Xception: Deep Learning with Depthwise Separable Convolutions , isbn =

    Chollet, Francois , urldate =. Xception: Deep Learning with Depthwise Separable Convolutions , isbn =. 2017. doi:10.1109/CVPR.2017.195 , shorttitle =

  53. [87]

    Mamba: Linear-Time Sequence Modeling with Selective State Spaces , url =

    Gu, Albert and Dao, Tri , urldate =. Mamba: Linear-Time Sequence Modeling with Selective State Spaces , url =. doi:10.48550/arXiv.2312.00752 , shorttitle =

  54. [88]

    European Conference on Computer Vision , author =

    Make-Your-3. European Conference on Computer Vision , author =

  55. [89]

    Xiaoxiao Long and Yuan. Wonder3

  56. [90]

    and Nguyen, Quoc-Dung and Kingkan, Cherdsak , urldate =

    Huynh Ngoc Nhu, Y. and Nguyen, Quoc-Dung and Kingkan, Cherdsak , urldate =. A novel approach with vision-language models for custom e-commerce product listings , rights =. doi:10.1007/s11042-025-20873-4 , abstract =

  57. [92]

    Multi-View Image Fusion , rights =

    Trinidad, Marc Comino and Martin-Brualla, Ricardo and Kainz, Florian and Kontkanen, Janne , urldate =. Multi-View Image Fusion , rights =. 2019. doi:10.1109/ICCV.2019.00420 , abstract =

  58. [94]

    doi:10.48550/arXiv.2305.18766 , eprinttype =

    Zhu, Junzhe and Zhuang, Peiye and Koyejo, Sanmi , urldate =. doi:10.48550/arXiv.2305.18766 , eprinttype =. 2305.18766 [cs] , note =

  59. [95]

    doi:10.48550/arXiv.2312.11459 , eprinttype =

    Tang, Zhicong and Gu, Shuyang and Wang, Chunyu and Zhang, Ting and Bao, Jianmin and Chen, Dong and Guo, Baining , urldate =. doi:10.48550/arXiv.2312.11459 , eprinttype =. 2312.11459 [cs] , note =

  60. [96]

    Turbo3D: Ultra-fast Text-to-3D Generation , url =

    Hu, Hanzhe and Yin, Tianwei and Luan, Fujun and Hu, Yiwei and Tan, Hao and Xu, Zexiang and Bi, Sai and Tulsiani, Shubham and Zhang, Kai , urldate =. Turbo3D: Ultra-fast Text-to-3D Generation , url =. doi:10.48550/arXiv.2412.04470 , eprinttype =. 2412.04470 [cs] , note =

  61. [97]

    doi:10.48550/arXiv.2306.14685 , eprinttype =

    Xing, Ximing and Wang, Chuang and Zhou, Haitao and Zhang, Jing and Yu, Qian and Xu, Dong , urldate =. doi:10.48550/arXiv.2306.14685 , eprinttype =. 2306.14685 [cs] , note =

  62. [98]

    Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models , url =

    Ye, Jiacheng and Gong, Shansan and Chen, Liheng and Zheng, Lin and Gao, Jiahui and Shi, Han and Wu, Chuan and Jiang, Xin and Li, Zhenguo and Bi, Wei and Kong, Lingpeng , urldate =. Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models , url =. doi:10.4...

  63. [99]

    From Explicit

    Deng, Yuntian and Choi, Yejin and Shieber, Stuart , urldate =. From Explicit. doi:10.48550/arXiv.2405.14838 , eprinttype =. 2405.14838 [cs] , note =

  64. [100]

    On the Measure of Intelligence , url =

    Chollet, Fran�ois , urldate =. On the Measure of Intelligence , url =. doi:10.48550/arXiv.1911.01547 , eprinttype =. 1911.01547 [cs] , note =

  65. [101]

    , urldate =

    Ellis, Kevin and Wong, Catherine and Nye, Maxwell and Sable-Meyer, Mathias and Cary, Luc and Morales, Lucas and Hewitt, Luke and Solar-Lezama, Armando and Tenenbaum, Joshua B. , urldate =. doi:10.48550/arXiv.2006.08381 , eprinttype =. 2006.08381 [cs] , note =

  66. [102]

    Classifier-Free Diffusion Guidance , url =

    Ho, Jonathan and Salimans, Tim , urldate =. Classifier-Free Diffusion Guidance , url =. doi:10.48550/arXiv.2207.12598 , eprinttype =. 2207.12598 [cs] , note =

  67. [103]

    doi:10.48550/arXiv.2210.08933 , eprinttype =

    Gong, Shansan and Li, Mukai and Feng, Jiangtao and Wu, Zhiyong and Kong, Lingpeng , urldate =. doi:10.48550/arXiv.2210.08933 , eprinttype =. 2210.08933 [cs] , note =

  68. [104]

    Prompt-to-Prompt Image Editing with Cross Attention Control , url =

    Hertz, Amir and Mokady, Ron and Tenenbaum, Jay and Aberman, Kfir and Pritch, Yael and Cohen-Or, Daniel , urldate =. Prompt-to-Prompt Image Editing with Cross Attention Control , url =. doi:10.48550/arXiv.2208.01626 , eprinttype =. 2208.01626 [cs] , note =

  69. [105]

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , url =

    Wei, Jason and Wang, Xuezhi and Schuurmans, Dale and Bosma, Maarten and Ichter, Brian and Xia, Fei and Chi, Ed and Le, Quoc and Zhou, Denny , urldate =. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , url =. doi:10.48550/arXiv.2201.11903 , eprinttype =....

  70. [106]

    doi:10.48550/arXiv.2412.04604 , eprinttype =

    Chollet, Francois and Knoop, Mike and Kamradt, Gregory and Landers, Bryan , urldate =. doi:10.48550/arXiv.2412.04604 , eprinttype =. 2412.04604 [cs] , note =

  71. [107]

    and Tancik, Matthew and Barron, Jonathan T

    Mildenhall, Ben and Srinivasan, Pratul P. and Tancik, Matthew and Barron, Jonathan T. and Ramamoorthi, Ravi and Ng, Ren , urldate =. doi:10.48550/arXiv.2003.08934 , eprinttype =. 2003.08934 [cs] , note =

  72. [108]

    Mamba: Linear-Time Sequence Modeling with Selective State Spaces , url =

    Gu, Albert and Dao, Tri , urldate =. Mamba: Linear-Time Sequence Modeling with Selective State Spaces , url =

  73. [109]

    Raj, Amit and Kaza, Srinivas and Poole, Ben and Niemeyer, Michael and Ruiz, Nataniel and Mildenhall, Ben and Zada, Shiran and Aberman, Kfir and Rubinstein, Michael and Barron, Jonathan and Li, Yuanzhen and Jampani, Varun , urldate =. 2023. doi:10.1109/ICCV51070.2023.00223 , sh...

  74. [110]

    and Welling, Max , urldate =

    Kingma, Diederik P. and Welling, Max , urldate =. Auto-Encoding Variational Bayes , url =. doi:10.48550/arXiv.1312.6114 , eprinttype =. 1312.6114 [stat] , keywords =

  75. [111]

    doi:10.48550/arXiv.1908.03557 , eprinttype =

    Li, Liunian Harold and Yatskar, Mark and Yin, Da and Hsieh, Cho-Jui and Chang, Kai-Wei , urldate =. doi:10.48550/arXiv.1908.03557 , eprinttype =. 1908.03557 [cs] , note =

  76. [112]

    doi:10.48550/arXiv.1908.02265 , shorttitle =

    Lu, Jiasen and Batra, Dhruv and Parikh, Devi and Lee, Stefan , urldate =. doi:10.48550/arXiv.1908.02265 , shorttitle =. 1908.02265 [cs] , keywords =

  77. [113]

    doi:10.48550/arXiv.1908.07490 , shorttitle =

    Tan, Hao and Bansal, Mohit , urldate =. doi:10.48550/arXiv.1908.07490 , shorttitle =. 1908.07490 [cs] , keywords =

  78. [114]

    Learning Transferable Visual Models From Natural Language Supervision , url =

    Radford, Alec and Kim, Jong Wook and Hallacy, Chris and Ramesh, Aditya and Goh, Gabriel and Agarwal, Sandhini and Sastry, Girish and Askell, Amanda and Mishkin, Pamela and Clark, Jack and Krueger, Gretchen and Sutskever, Ilya , urldate =. Learning Transferable Visual Models Fr...

  79. [115]

    Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision , url =

    Jia, Chao and Yang, Yinfei and Xia, Ye and Chen, Yi-Ting and Parekh, Zarana and Pham, Hieu and Le, Quoc and Sung, Yun-Hsuan and Li, Zhen and Duerig, Tom , urldate =. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision , url =. Proceedings ...

  80. [116]

    doi:10.48550/arXiv.2311.01361 , abstract =

    Zhang, Xinlu and Lu, Yujie and Wang, Weizhi and Yan, An and Yan, Jun and Qin, Lianke and Wang, Heng and Yan, Xifeng and Wang, William Yang and Petzold, Linda Ruth , urldate =. doi:10.48550/arXiv.2311.01361 , abstract =. 2311.01361 [cs] , keywords =

  81. [117]

    Hierarchical Text-Conditional Image Generation with

    Ramesh, Aditya and Dhariwal, Prafulla and Nichol, Alex and Chu, Casey and Chen, Mark , urldate =. Hierarchical Text-Conditional Image Generation with. doi:10.48550/arXiv.2204.06125 , abstract =. 2204.06125 [cs] , keywords =

  82. [118]

    Denton and Seyed Kamyar Seyed Ghasemipour and Raphael Gontijo Lopes and Burcu Karagol Ayan and Tim Salimans and Jonathan Ho and David J

    Chitwan Saharia and William Chan and Saurabh Saxena and Lala Li and Jay Whang and Emily L. Denton and Seyed Kamyar Seyed Ghasemipour and Raphael Gontijo Lopes and Burcu Karagol Ayan and Tim Salimans and Jonathan Ho and David J. Fleet and Mohammad Norouzi , title =. Advances in...

  83. [119]

    and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , urldate =

    Hu, Edward J. and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , urldate =. doi:10.48550/arXiv.2106.09685 , shorttitle =. 2106.09685 [cs] , keywords =

  84. [120]

    and Mildenhall, Ben and Tancik, Matthew and Hedman, Peter and Martin-Brualla, Ricardo and Srinivasan, Pratul P

    Barron, Jonathan T. and Mildenhall, Ben and Tancik, Matthew and Hedman, Peter and Martin-Brualla, Ricardo and Srinivasan, Pratul P. , urldate =. Mip-. doi:10.48550/arXiv.2103.13415 , shorttitle =. 2103.13415 [cs] , keywords =

  85. [121]

    Instant Neural Graphics Primitives with a Multiresolution Hash Encoding , volume =

    Müller, Thomas and Evans, Alex and Schied, Christoph and Keller, Alexander , urldate =. Instant Neural Graphics Primitives with a Multiresolution Hash Encoding , volume =. doi:10.1145/3528223.3530127 , abstract =. 2201.05989 [cs] , keywords =

  86. [122]

    Text2Mesh: Text-Driven Neural Stylization for Meshes , rights =

    Michel, Oscar and Bar-On, Roi and Liu, Richard and Benaim, Sagie and Hanocka, Rana , urldate =. Text2Mesh: Text-Driven Neural Stylization for Meshes , rights =. 2022. doi:10.1109/CVPR52688.2022.01313 , shorttitle =

  87. [123]

    Minghua Liu and Ruoxi Shi and Linghao Chen and Zhuoyang Zhang and Chao Xu and Xinyue Wei and Hansheng Chen and Chong Zeng and Jiayuan Gu and Hao Su , title =

  88. [124]

    Ruoshi Liu and Rundi Wu and Basile Van Hoorick and Pavel Tokmakov and Sergey Zakharov and Carl Vondrick , title =

  89. [125]

    International Conference on Learning Representations , year =

    Yuan Liu and Cheng Lin and Zijiao Zeng and Xiaoxiao Long and Lingjie Liu and Taku Komura and Wenping Wang , title =. International Conference on Learning Representations , year =

  90. [126]

    doi:10.48550/arXiv.2308.16512 , shorttitle =

    Shi, Yichun and Wang, Peng and Ye, Jianglong and Long, Mai and Li, Kejie and Yang, Xiao , urldate =. doi:10.48550/arXiv.2308.16512 , shorttitle =. 2308.16512 [cs] , keywords =

  91. [127]

    On Scaling Up a Multilingual Vision and Language Model , rights =

    Chen, Xi and Djolonga, Josip and Padlewski, Piotr and Mustafa, Basil and Changpinyo, Soravit and Wu, Jialin and Ruiz, Carlos Riquelme and Goodman, Sebastian and Wang, Xiao and Tay, Yi and Shakeri, Siamak and Dehghani, Mostafa and Salz, Daniel and Lucic, Mario and Tschannen, Mi...

  92. [128]

    Advances in Neural Information Processing Systems , year =

    Kailu Wu and Fangfu Liu and Zhihan Cai and Runjie Yan and Hanyang Wang and Yating Hu and Yueqi Duan and Kaisheng Ma , title =. Advances in Neural Information Processing Systems , year =

  93. [129]

    Luke Melas. Real. 2023 , pages =

  94. [130]

    CoRR , volume =

    Alex Nichol and Heewoo Jun and Prafulla Dhariwal and Pamela Mishkin and Mark Chen , title =. CoRR , volume =

  95. [131]

    CoRR , volume =

    Heewoo Jun and Alex Nichol , title =. CoRR , volume =

  96. [132]

    Zero-1-to-3: Zero-shot One Image to 3

    Liu, Ruoshi and Wu, Rundi and Hoorick, Basile Van and Tokmakov, Pavel and Zakharov, Sergey and Vondrick, Carl , langid =. Zero-1-to-3: Zero-shot One Image to 3

  97. [133]

    Hong, Yicong and Zhang, Kai and Gu, Jiuxiang and Bi, Sai and Zhou, Yang and Liu, Difan and Liu, Feng and Sunkavalli, Kalyan and Bui, Trung and Tan, Hao , year=

  98. [134]

    CoRR , volume =

    Jiale Xu and Weihao Cheng and Yiming Gao and Xintao Wang and Shenghua Gao and Ying Shan , title =. CoRR , volume =

  99. [135]

    European Conference on Computer Vision , pages =

    Zhengyi Wang and Yikai Wang and Yifei Chen and Chendong Xiang and Shuo Chen and Dajiang Yu and Chongxuan Li and Hang Su and , title =. European Conference on Computer Vision , pages =

  100. [136]

    Downs, Laura and Francis, Anthony and Koenig, Nate and Kinman, Brandon and Hickman, Ryan and Reymann, Krista and. Google. International Conference on Robotics and Automation , year =

  101. [137]

    Nataniel Ruiz and Yuanzhen Li and Varun Jampani and Yael Pritch and Michael Rubinstein and Kfir Aberman , title =

  102. [138]

    Srinivasan and Matthew Tancik and Jonathan T

    Ben Mildenhall and Pratul P. Srinivasan and Matthew Tancik and Jonathan T. Barron and Ravi Ramamoorthi and Ren Ng , title =. European Conference on Computer Vision , pages =

  103. [139]

    2021 , pages =

    Alex Yu and Vickie Ye and Matthew Tancik and Angjoo Kanazawa , title =. 2021 , pages =

  104. [140]

    Barron and Ben Mildenhall and Mehdi S

    Michael Niemeyer and Jonathan T. Barron and Ben Mildenhall and Mehdi S. M. Sajjadi and Andreas Geiger and Noha Radwan , title =. 2022 , pages =

  105. [141]

    Barron and Ben Mildenhall and Pratul P

    Barbara Roessle and Jonathan T. Barron and Ben Mildenhall and Pratul P. Srinivasan and Matthias Nie. Dense Depth Priors for Neural Radiance Fields from Sparse Input Views , booktitle =. 2022 , pages =

  106. [142]

    Advances in Neural Information Processing Systems , pages =

    Lior Yariv and Jiatao Gu and Yoni Kasten and Yaron Lipman , title =. Advances in Neural Information Processing Systems , pages =

  107. [143]

    Srinivasan and Jonathan T

    Konstantinos Rematas and Andrew Liu and Pratul P. Srinivasan and Jonathan T. Barron and Andrea Tagliasacchi and Thomas A. Funkhouser and Vittorio Ferrari , title =

  108. [144]

    Multi-view to Novel View: Synthesizing Novel Views With Self-learned Confidence , booktitle =

    Shao. Multi-view to Novel View: Synthesizing Novel Views With Self-learned Confidence , booktitle =

  109. [145]

    Nanyang Wang and Yinda Zhang and Zhuwen Li and Yanwei Fu and Wei Liu and Yu. Pixel2. European Conference on Computer Vision , year =

  110. [146]

    Mescheder and Michael Oechsle and Michael Niemeyer and Sebastian Nowozin and Andreas Geiger , title =

    Lars M. Mescheder and Michael Oechsle and Michael Niemeyer and Sebastian Nowozin and Andreas Geiger , title =. 2019 , pages =

  111. [147]

    2022 , pages=

    Deng, Yu and Yang, Jiaolong and Xiang, Jianfeng and Tong, Xin , booktitle=. 2022 , pages=

  112. [148]

    2024 , pages =

    Zi‑Ting Chou and Sheng‑Yu Huang and I‑Jieh Liu and Yu‑Chiang Frank Wang , title =. 2024 , pages =

  113. [149]

    Rui Chen and Yongwei Chen and Ningxin Jiao and Kui Jia , title =

  114. [150]

    and Frahm, Jan-Michael , title =

    Schonberger, Johannes L. and Frahm, Jan-Michael , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , month =

  115. [151]

    Qin, Xuebin and Zhang, Zichen and Huang, Chenyang and Dehghan, Masood and Zaiane, Osmar and Jagersand, Martin , journal =. U\(

  116. [152]

    Public opinion quarterly , volume=

    Effects of questionnaire length on participation and indicators of response quality in a web survey , author=. Public opinion quarterly , volume=

  117. [153]

    Jiale Xu and Shenghua Gao and Ying Shan , title =

  118. [154]

    Advances in Neural Information Processing Systems , year=

    Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis , author=. Advances in Neural Information Processing Systems , year=

  119. [155]

    Bernhard Kerbl and Georgios Kopanas and Thomas Leimk. 3

  120. [156]

    CoRR , year =

    Hu Ye and Jun Zhang and Sibo Liu and Xiao Han and Wei Yang , title =. CoRR , year =

  121. [157]

    Lvmin Zhang and Anyi Rao and Maneesh Agrawala , title =

  122. [158]

    International Conference on Learning Representations , year =

    Jiahao Chang and Chongjie Ye and Yushuang Wu and Yuantao Chen and Yidan Zhang and Zhongjin Luo and Chenghong Li and Yihao Zhi and Xiaoguang Han , title =. International Conference on Learning Representations , year =

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.