Pith. sign in

REVIEW 3 major objections 3 minor 44 references

scI2CL claims that simultaneously contrasting cells within each omics and across omics yields fused cellular representations that outperform eight existing methods in clustering and uniquely recover a developmental trajectory and latent cel

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

The abstract claims a state-of-the-art single-cell multi-omics integration method with new cell-subtype and trajectory findings, but the supplied full text is a different paper.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection Unreviewable as-is: the supplied full text is an unrelated 3D reconstruction paper, so the abstract's strong benchmark and biological-discovery claims rest on nothing we can inspect. the 3 major comments →

arxiv 2508.18304 v1 pith:LHOPYJDW submitted 2025-08-23 q-bio.GN cs.AIcs.LGq-bio.CB

scI2CL: Effectively Integrating Single-cell Multi-omics by Intra- and Inter-omics Contrastive Learning

classification q-bio.GN cs.AIcs.LGq-bio.CB
keywords single-cell multi-omicscontrastive learningrepresentation learningmulti-omics integrationcell clusteringtrajectory inferencecell subtypingcellular heterogeneity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

scI2CL is a framework for fusing single-cell multi-omics data—multiple molecular measurements taken from the same cells—into one representation that can be used for clustering, cell subtyping, trajectory reconstruction, and classification. The central claim is that contrasting cells both within each omics and across omics (intra- and inter-omics contrastive learning) captures cross-omics relationships that standard integration methods miss. On this basis the paper reports that scI2CL outperforms eight state-of-the-art methods on four real-world datasets for clustering, is the only tested method to reconstruct the developmental trajectory from hematopoietic stem and progenitor cells to Memory B cells, and resolves latent monocyte subpopulations and misclassified CD4+ T cell subtypes. The significance, if these claims hold, is that a single learned representation can support several downstream analyses without task-specific tuning. It should be noted that the supplied full text belongs to a different manuscript on 3D reconstruction, so the abstract's claims cannot be checked against methods or experiment tables in this document.

Core claim

The paper's central discovery, stated in its abstract, is that a fusion model built from two contrastive objectives—one that learns cellular similarities inside each omics modality and one that aligns cells with their counterparts across modalities—produces cellular representations that are 'comprehensive and discriminative'. The authors report that scI2CL surpasses eight state-of-the-art methods on four widely-used real-world datasets in clustering; correctly constructs a cell developmental trajectory from hematopoietic stem and progenitor cells to Memory B cells, where existing methods fail; distinguishes three latent monocyte subpopulations not found by existing methods; and corrects the

What carries the argument

The central object is the contrastive learning objective operating at two levels. Intra-omics contrastive learning builds a representation in which cells that are similar within a single modality (for example, within the transcriptome alone) are pulled together and dissimilar cells are pushed apart. Inter-omics contrastive learning then aligns the same cell's representations across modalities, so that complementary information from each omics is combined into a shared embedding. The paper's claim is that this paired intra/inter objective is what yields representations that preserve fine-grained biological structure—latent subpopulations and continuous developmental order—rather than merely a

Load-bearing premise

The load-bearing premise is that the eight baseline methods were run and tuned under configurations representative of their best published performance on identical data splits and metrics, and that the reference labels and trajectory used to judge 'correct' reconstruction are valid; this cannot be checked because the supplied full text is a different manuscript.

What would settle it

Recompute the four clustering benchmarks with all nine methods under identical, publicly documented hyperparameter settings; if any baseline matches scI2CL's clustering performance or reconstructs the HSPC-to-Memory B trajectory on the same data, the paper's universal claims ('only method', 'not discovered by existing methods') are falsified. A simpler check is availability: if the datasets, code, and exact splits are not recoverable, the claims cannot be independently confirmed.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the clustering claim holds, scI2CL's fused representation can serve as a ready-made input for cell-type annotation and downstream differential analyses on multi-omics datasets.
  • If the trajectory claim holds, the embedding preserves continuous differentiation order, enabling pseudotime and lineage analysis without a separate trajectory-inference step.
  • If the subpopulation claim holds, contrastive fusion can expose rare or intermediate cell states that are averaged away by other integration methods, useful for discovering disease-relevant cell types.
  • If the misclassification result holds, the method can separate closely related immune cell subsets, improving resolution in immunology studies.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the contrastive formulation does not depend on specific omics types, the same intra/inter scheme could plausibly extend to other modality pairs, such as methylation or chromatin accessibility combined with transcriptomics, provided matched single-cell data exist.
  • The trajectory result suggests the differentiation axis emerges from the embedding itself; a testable extension would be checking whether the same representation reconstructs branching trajectories with multiple lineages, rather than a single stem-to-memory path.
  • If the claimed advantages are real, benchmark designers could expect future integration methods to be judged on whether they recover known biology (trajectories, minor subsets) and not only on clustering agreement with reference labels.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The submission consists of an abstract for scI2CL, a single-cell multi-omics integration method based on intra- and inter-omics contrastive learning, and a full text that is a different paper entirely: SERES, a sparse-view 3D reconstruction method from computer vision (arXiv:2508.18314). The abstract claims that scI2CL outperforms eight state-of-the-art methods on four real-world datasets for clustering, uniquely reconstructs an HSPC-to-Memory B cell trajectory, resolves CD4+ T cell subtyping ambiguities, and discovers monocyte subpopulations missed by existing methods. None of these claims are supported by the supplied full text, which contains no method description, equations, datasets, baseline names, metrics, or results for scI2CL.

Significance. If the claims were true, scI2CL would be a potentially valuable contribution to single-cell multi-omics integration. However, the manuscript as submitted provides no verifiable technical or empirical content for scI2CL. The central results, including all comparative and biological-discovery claims, are bare assertions. No code, data, or reproducibility artifacts are present. Consequently, the paper has no assessable significance in its current form.

major comments (3)
  1. [Abstract and Full Text (all)] The supplied full text is the SERES paper on semantic-aware 3D reconstruction from sparse views (arXiv:2508.18314), not the scI2CL manuscript. This is a load-bearing defect: the abstract's central claims—'surpasses eight state-of-the-art methods on four widely-used real-world datasets,' 'the only method that correctly constructs the cell developmental trajectory,' and 'distinguishes three latent monocyte cell subpopulations'—are made without any accompanying equations, architecture description, training details, dataset identifiers, or result tables. The technical content of scI2CL cannot be inspected.
  2. [Abstract (empirical comparisons)] The comparative claims rest on unstated evaluation conditions. There is no specification of which eight baselines were used, how their hyperparameters were selected, whether data splits and preprocessing were identical, which clustering metrics were computed, or whether any statistical significance testing was performed. The universal negatives ('the only method,' 'not discovered by existing methods') require a transparent comparison protocol to be meaningful. As submitted, no such protocol exists.
  3. [Full Text (limitations paragraph, p. 11)] The only limitation discussion in the supplied text concerns 3D reconstruction artifacts—reflective surfaces and textureless areas—and is irrelevant to scI2CL. The manuscript provides no scI2CL-specific limitations, such as sensitivity to modality completeness, dependence on reference cell-type labels, computational cost, or failure modes. This further confirms that the submitted document does not contain the claimed contribution.
minor comments (3)
  1. [Full Text (footer, p. 1)] The full text is labeled arXiv:2508.18314, whereas the review target is arXiv:2508.18304. The identifiers do not match, and the full text's title, abstract, and references are all from the unrelated SERES paper.
  2. [Full Text (references)] The reference list contains only computer-vision works (NeuS, VolRecon, SAM, Neuralangelo, etc.) and no single-cell or multi-omics references. This is inconsistent with the scI2CL abstract and makes the manuscript internally incoherent.
  3. [Full Text (all)] No reproducibility materials or data-availability statement for scI2CL are provided. Even if the correct full text were available, the absence of code or data would hamper verification of the clustering and trajectory claims.

Circularity Check

0 steps flagged

No circularity detectable: supplied full text is the unrelated SERES paper, so scI2CL's derivation and experiments are absent.

full rationale

The abstract describes scI2CL, a single-cell multi-omics contrastive learning method with claims of superior clustering, trajectory reconstruction, and subpopulation discovery. However, the supplied full text is arXiv:2508.18314 (SERES), a completely different computer-vision paper on sparse-view 3D reconstruction. No equations, experimental protocols, dataset descriptions, baseline tuning details, or derivation chain for scI2CL are present in the provided material. A circularity analysis requires quoting a specific reduction in the paper's own argument, such as a fitted parameter renamed as a prediction or a self-citation used as the sole justification for a central premise. No such passage exists in the supplied text because the scI2CL manuscript itself is absent. The mismatch is an evidence gap and a reviewability failure, not a demonstrated circularity. Therefore, on the evidence provided, no circular step can be identified, and the appropriate circularity score is 0.

Axiom & Free-Parameter Ledger

1 free parameters · 3 axioms · 0 invented entities

This ledger is minimal by necessity: the supplied full text is an unrelated manuscript, so the actual scI2CL objective, hyperparameters, and architectures are not in view. The entries listed are the assumptions that any fusion-via-contrastive-learning claim of this type would rest on, inferred from the abstract alone.

free parameters (1)
  • Model hyperparameters (embedding dimensions, contrastive temperature, relative weighting of intra- vs inter-omics losses = undisclosed (abstract only)
    The abstract provides no objective function or training details; the reported superiority over eight baselines depends on these choices, which are not auditable from the provided material.
axioms (3)
  • domain assumption Across-omics profiles from the same cell are meaningful correspondences to align in embedding space
    The inter-omics contrastive objective presupposes that modalities are coupled at the single-cell level; batch and technical noise in multi-omics assays can break this assumption, but the abstract does not address it.
  • domain assumption Benchmark annotations and reference trajectories used for evaluation are correct
    Claims of 'correct' trajectory construction and subtype discovery depend on the ground truth used; the abstract names no datasets or validation strategy.
  • domain assumption Contrastive objectives produce embeddings that capture true biological relatedness
    The framework relies on the standard contrastive-learning premise (alignment of positives, separation of negatives) without proof; common practice in the field, but a genuine modeling assumption.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of scI2CL: Effectively Integrating Single-cell Multi-omics by Intra- and Inter-omics Contrastive Learning." pith.science (2026). https://pith.science/paper/LHOPYJDW

@misc{pith2026250818304,
  author       = {Pith},
  title        = {Pith review of: scI2CL: Effectively Integrating Single-cell Multi-omics by Intra- and Inter-omics Contrastive Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LHOPYJDW}},
  note         = {Machine review of arXiv:2508.18304}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Single-cell multi-omics data contain huge information of cellular states, and analyzing these data can reveal valuable insights into cellular heterogeneity, diseases, and biological processes. However, as cell differentiation \& development is a continuous and dynamic process, it remains challenging to computationally model and infer cell interaction patterns based on single-cell multi-omics data. This paper presents scI2CL, a new single-cell multi-omics fusion framework based on intra- and inter-omics contrastive learning, to learn comprehensive and discriminative cellular representations from complementary multi-omics data for various downstream tasks. Extensive experiments of four downstream tasks validate the effectiveness of scI2CL and its superiority over existing peers. Concretely, in cell clustering, scI2CL surpasses eight state-of-the-art methods on four widely-used real-world datasets. In cell subtyping, scI2CL effectively distinguishes three latent monocyte cell subpopulations, which are not discovered by existing methods. Simultaneously, scI2CL is the only method that correctly constructs the cell developmental trajectory from hematopoietic stem and progenitor cells to Memory B cells. In addition, scI2CL resolves the misclassification of cell types between two subpopulations of CD4+ T cells, while existing methods fail to precisely distinguish the mixed cells. In summary, scI2CL can accurately characterize cross-omics relationships among cells, thus effectively fuses multi-omics data and learns discriminative cellular representations to support various downstream analysis tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

44 extracted references · 34 canonical work pages

  1. [1]

    Neuris: Neural reconstruction of indoor scenes using normal priors,

    J. Wang, P . Wang, X. Long, C. Theobalt, T. Komura, L. Liu, and W. Wang, “Neuris: Neural reconstruction of indoor scenes using normal priors,” in Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXII. Berlin, Heidelberg: Springer-Verlag, 2022, p. 139–155. [Online]. Available: https://doi.org/...

  2. [2]

    Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,

    P . Wang, L. Liu, Y. Liu, C. Theobalt, T. Komura, and W. Wang, “Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,” NeurIPS, 2021

  3. [3]

    Nerf++: Analyz- ing and improving neural radiance fields,

    K. Zhang, G. Riegler, N. Snavely, and V . Koltun, “Nerf++: Analyz- ing and improving neural radiance fields,” arXiv:2010.07492, 2020

  4. [4]

    Nerfingmvs: Guided optimization of neural radiance fields for indoor multi- view stereo,

    Y. Wei, S. Liu, Y. Rao, W. Zhao, J. Lu, and J. Zhou, “Nerfingmvs: Guided optimization of neural radiance fields for indoor multi- view stereo,” in ICCV, 2021

  5. [5]

    Vdn- nerf: Resolving shape-radiance ambiguity via view-dependence normalization,

    B. Zhu, Y. Yang, X. Wang, Y. Zheng, and L. Guibas, “Vdn- nerf: Resolving shape-radiance ambiguity via view-dependence normalization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 35– 45

  6. [6]

    pixelNeRF: Neural radiance fields from one or few images,

    A. Yu, V . Ye, M. Tancik, and A. Kanazawa, “pixelNeRF: Neural radiance fields from one or few images,” in CVPR, 2021

  7. [7]

    Sparseneus: Fast generalizable neural surface reconstruction from sparse views,

    X. Long, C. Lin, P . Wang, T. Komura, and W. Wang, “Sparseneus: Fast generalizable neural surface reconstruction from sparse views,” ECCV, 2022

  8. [8]

    Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo,

    A. Chen, Z. Xu, F. Zhao, X. Zhang, F. Xiang, J. Yu, and H. Su, “Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14 124–14 133

  9. [9]

    MonoSDF: Exploring monocular geometric cues for neural implicit surface reconstruction,

    Z. Yu, S. Peng, M. Niemeyer, T. Sattler, and A. Geiger, “MonoSDF: Exploring monocular geometric cues for neural implicit surface reconstruction,” in Advances in Neural Information Processing Systems, A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022. [Online]. Available: https://openreview.net/forum? id=dMK7EwoTYp

  10. [10]

    Sparsenerf: Distilling depth ranking for few-shot novel view synthesis,

    Guangcong, Z. Chen, C. C. Loy, and Z. Liu, “Sparsenerf: Distilling depth ranking for few-shot novel view synthesis,” in ICCV, 2023

  11. [11]

    Volrecon: Volume rendering of signed ray distance functions for generalizable multi-view reconstruction,

    Y. Ren, F. Wang, T. Zhang, M. Pollefeys, and S. S ¨usstrunk, “Volrecon: Volume rendering of signed ray distance functions for generalizable multi-view reconstruction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 16 685–16 695

  12. [12]

    Nerfactor: Neural factorization of shape and reflectance under an unknown illumination,

    X. Zhang, P . P . Srinivasan, B. Deng, P . Debevec, W. T. Freeman, and J. T. Barron, “Nerfactor: Neural factorization of shape and reflectance under an unknown illumination,” ACM Trans. Graph. , vol. 40, no. 6, dec 2021. [Online]. Available: https://doi.org/10.1145/3478513.3480496

  13. [13]

    A theory of shape by space carving,

    K. N. Kutulakos and S. M. Seitz, “A theory of shape by space carving,” International Journal of Computer Vision , vol. 38, no. 3, pp. 199–218, 2000. [Online]. Available: https://doi.org/10.1023/a: 1008191222954

  14. [14]

    Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs,

    M. Niemeyer, J. T. Barron, B. Mildenhall, M. S. M. Sajjadi, A. Geiger, and N. Radwan, “Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs,” in Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2022

  15. [15]

    SparsePose: Sparse-view camera pose regression and refinement,

    S. Sinha, J. Y. Zhang, A. Tagliasacchi, I. Gilitschenski, and D. B. Lindell, “SparsePose: Sparse-view camera pose regression and refinement,” in Computer Vision and Pattern Recognition (CVPR) , 2023

  16. [16]

    Sc-neus: Consistent neural surface reconstruction from sparse and noisy views,

    S.-S. Huang, Z. Zou, Y. Zhang, Y.-P . Cao, and Y. Shan, “Sc-neus: Consistent neural surface reconstruction from sparse and noisy views,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 3, pp. 2357–2365, Mar. 2024. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/28010

  17. [17]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P . P . Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in ECCV, 2020

  18. [18]

    Learning signed distance field for multi-view surface reconstruction,

    J. Zhang, Y. Yao, and L. Quan, “Learning signed distance field for multi-view surface reconstruction,” International Conference on Computer Vision (ICCV), 2021

  19. [19]

    Volume rendering of neural implicit surfaces,

    L. Yariv, J. Gu, Y. Kasten, and Y. Lipman, “Volume rendering of neural implicit surfaces,” Advances in Neural Information Processing Systems, vol. 34, pp. 4805–4815, 2021

  20. [20]

    Improving neural implicit surfaces geometry with patch warp- ing,

    F. Darmon, B. Bascle, J.-C. Devaux, P . Monasse, and M. Aubry, “Improving neural implicit surfaces geometry with patch warp- ing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022, pp. 6260–6269

  21. [21]

    Neusurf: On-surface priors for neural surface reconstruction from sparse input views,

    H. Huang, Y. Wu, J. Zhou, G. Gao, M. Gu, and Y.-S. Liu, “Neusurf: On-surface priors for neural surface reconstruction from sparse input views,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 3, 2024, pp. 2312–2320

  22. [22]

    Multiview neural surface reconstruction by disentangling geometry and appearance,

    L. Yariv, Y. Kasten, D. Moran, M. Galun, M. Atzmon, B. Ronen, and Y. Lipman, “Multiview neural surface reconstruction by disentangling geometry and appearance,” Advances in Neural Information Processing Systems, vol. 33, 2020

  23. [23]

    Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision,

    M. Niemeyer, L. Mescheder, M. Oechsle, and A. Geiger, “Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Los Alamitos, CA, USA: IEEE Computer Society, jun 2020, pp. 3501–3512. [Online]. Available: https://doi.ieeecomputersociety...

  24. [24]

    Ref-NeRF: Structured view-dependent appearance for neural radiance fields,

    D. Verbin, P . Hedman, B. Mildenhall, T. Zickler, J. T. Barron, and P . P . Srinivasan, “Ref-NeRF: Structured view-dependent appearance for neural radiance fields,” CVPR, 2022

  25. [25]

    Nope- nerf: Optimising neural radiance field with no pose prior,

    W. Bian, Z. Wang, K. Li, J.-W. Bian, and V . A. Prisacariu, “Nope- nerf: Optimising neural radiance field with no pose prior,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 4160–4169

  26. [26]

    Uforecon: Generalizable sparse-view surface reconstruction from arbitrary and unfavorable sets,

    Y. Na, W. J. Kim, K. B. Han, S. Ha, and S.-E. Yoon, “Uforecon: Generalizable sparse-view surface reconstruction from arbitrary and unfavorable sets,” in Proceedings of the IEEE/CVF Conference 12 on Computer Vision and Pattern Recognition (CVPR) , June 2024, pp. 5094–5104

  27. [27]

    Geo-neus: Geometry- consistent neural implicit surfaces learning for multi-view reconstruction,

    Q. Fu, Q. Xu, Y.-S. Ong, and W. Tao, “Geo-neus: Geometry- consistent neural implicit surfaces learning for multi-view reconstruction,” in Advances in Neural Information Processing Systems, A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022. [Online]. Available: https://openreview.net/forum? id=JvIFpZOjLF4

  28. [28]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P . Dollar, and R. Girshick, “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 4015–4026

  29. [29]

    SAD: segment any RGBD,

    J. Cen, Y. Wu, K. Wang, X. Li, J. Yang, Y. Pei, L. Kong, Z. Liu, and Q. Chen, “SAD: segment any RGBD,” CoRR, vol. abs/2305.14207,

  30. [30]

    Segment any point cloud sequences by distilling vision foundation models,

    Y. Liu, L. Kong, J. Cen, R. Chen, W. Zhang, L. Pan, K. Chen, and Z. Liu, “Segment any point cloud sequences by distilling vision foundation models,” in Advances in Neural Information Processing Systems, 2023

  31. [31]

    Sam3d: Segment anything in 3d scenes,

    Y. Yang, X. Wu, T. He, H. Zhao, and X. Liu, “Sam3d: Segment anything in 3d scenes,” ArXiv, vol. abs/2306.03908, 2023

  32. [32]

    Sam3d: Zero-shot 3d object detection via segment anything model,

    D. Zhang, D. Liang, H. Yang, Z. Zou, X. Ye, Z. Liu, and X. Bai, “Sam3d: Zero-shot 3d object detection via segment anything model,” ArXiv, vol. abs/2306.02245, 2023

  33. [33]

    Panoptic lifting for 3d scene understanding with neural fields,

    Y. Siddiqui, L. Porzi, S. R. Bul `o, N. M ¨uller, M. Nießner, A. Dai, and P . Kontschieder, “Panoptic lifting for 3d scene understanding with neural fields,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2023, pp. 9043–9052

  34. [34]

    Anything-3d: Towards single-view anything reconstruction in the wild,

    Q. Shen, X. Yang, and X. Wang, “Anything-3d: Towards single-view anything reconstruction in the wild,” ArXiv, vol. abs/2304.10261, 2023

  35. [35]

    High-resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P . Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022, pp. 10 684–10 695

  36. [36]

    Segment anything in 3d with nerfs,

    J. Cen, Z. Zhou, J. Fang, C. Yang, W. Shen, L. Xie, X. Zhang, and Q. Tian, “Segment anything in 3d with nerfs,” NeurIPS, 2023

  37. [37]

    Open-nerf: Towards open vocabulary nerf decomposition,

    H. Zhang, F. Li, and N. Ahuja, “Open-nerf: Towards open vocabulary nerf decomposition,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , January 2024, pp. 3456–3465

  38. [38]

    One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization,

    M. Liu, C. Xu, H. Jin, L. Chen, T. MukundVarma, Z. Xu, and H. Su, “One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization,” NeurIPS, 2023

  39. [39]

    Im- plicit geometric regularization for learning shapes,

    A. Gropp, L. Yariv, N. Haim, M. Atzmon, and Y. Lipman, “Im- plicit geometric regularization for learning shapes,” Proceedings of Machine Learning and Systems 2020, 2020

  40. [40]

    Large scale multi-view stereopsis evaluation,

    R. Jensen, A. Dahl, G. Vogiatzis, E. Tola, and H. Aanæs, “Large scale multi-view stereopsis evaluation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2014, pp. 406– 413

  41. [41]

    Blendedmvs: A large-scale dataset for generalized multi-view stereo networks,

    Y. Yao, Z. Luo, S. Li, J. Zhang, Y. Ren, L. Zhou, T. Fang, and L. Quan, “Blendedmvs: A large-scale dataset for generalized multi-view stereo networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1790– 1799

  42. [42]

    Neuralangelo: High-fidelity neural surface reconstruction,

    Z. Li, T. M ¨uller, A. Evans, R. H. Taylor, M. Unberath, M.-Y. Liu, and C.-H. Lin, “Neuralangelo: High-fidelity neural surface reconstruction,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023

  43. [43]

    Structure-from-motion re- visited,

    J. L. Sch ¨onberger and J.-M. Frahm, “Structure-from-motion re- visited,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4104–4113

  44. [2023]

    Available: https://doi.org/10.48550/arXiv.2305

    [Online]. Available: https://doi.org/10.48550/arXiv.2305. 14207

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.