Pith. sign in

REVIEW 3 major objections 7 minor 113 references

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation

T0 review · 3 major / 7 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Mosaic3D claims the largest open-vocabulary 3D mask-text dataset—5.6M pairs across 29K scenes—and uses it to reach state-of-the-art open-vocabulary 3D semantic and instance segmentation.

desk verdict A genuinely large 3D mask-text dataset and a solid model, but the OV3D baseline and the missing human evaluation of data quality need fixing before the headline claims are fully trustworthy. read the letter →

arxiv 2502.02548 v2 pith:QVN2ADES submitted 2025-02-04 cs.CV

classification cs.CV
keywords open-vocabulary3Dsegmentationmask-textdatasetdatagenerationpipelinecontrastivelearningregion-awarevision-languagemodelsinstancepointcloudfoundationmodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to remove the data bottleneck in open-vocabulary 3D scene understanding. The authors build an automatic pipeline that turns 2D foundation-model masks and region-aware captions into 3D mask-text pairs, and apply it to five indoor scene datasets to produce Mosaic3D-5.6M, with 5.6 million pairs across roughly 29,000 scenes. They then train a language-aligned 3D encoder with contrastive learning and a lightweight mask decoder on this data. The result, they argue, is a single open-vocabulary model that sets new state-of-the-art numbers on ScanNet20, ScanNet200, ScanNet++, and Matterport3D semantic segmentation, and achieves the first single-stage open-vocabulary 3D instance segmentation without ground-truth labels.

What carries the argument

The load-bearing machinery is the data engine: RAM++ tags objects, Grounding-DINO proposes boxes, SAM2 and SEEM produce precise foreground and panoptic masks, Osprey writes region-specific captions, and a projection-plus-depth-inclusion test transfers each 2D mask onto 3D points. On the model side, the contrastive per-point loss aligns point features with text embeddings, and the mask decoder's caption loss aligns mask embeddings with captions, supporting open-vocabulary instance segmentation without ground-truth labels.

What would settle it

Take a random sample of Mosaic3D-5.6M mask-text pairs, show a human the masked RGB region and the caption, and measure agreement; if captions frequently describe content outside the mask, the dataset's core premise fails. A cheaper quantitative probe: retrain the encoder with Osprey captions replaced by class-name-only labels from RAM++, and check whether the ScanNet200 gains collapse; if they do not collapse, the detailed region captions are not carrying the claimed benefit.

Watch

Extended reading notes

Core claim

The paper claims that large-scale, precisely masked, richly captioned 3D training data can be produced automatically—no human annotation—by combining open-vocabulary image segmentation (Grounded-SAM and SEEM) with a region-aware vision-language model (Osprey) and projecting the resulting 2D masks to 3D points through a depth inclusion test. On this data, contrastive per-point training aligns 3D geometry with text embeddings, and a Mask3D-style decoder, trained with an added caption loss, turns those aligned features into open-vocabulary instance predictions. Training on the full 5.6M-pair dataset yields state-of-the-art f-mIoU across four semantic segmentation benchmarks and a single-stage 3D-only instance segmenter that runs in about one second per scene, in contrast to prior methods that require multi-view CLIP inference.

Load-bearing premise

The whole dataset's value rests on the 2D teachers: if Osprey's captions do not actually describe the masked region, or if the projection and inclusion test attaches masks to the wrong 3D points, then the scale and apparent gains are artifacts of the teachers rather than genuine 3D understanding.

Editorial extensions

If this is right

  • If the data engine transfers to new scene datasets, scaling 3D open-vocabulary understanding becomes a matter of running 2D foundation models on more RGB-D scans rather than hiring annotators.
  • The single-stage 3D instance segmentation model shows that language-aligned 3D features can carry open-vocabulary instance prediction directly, making multi-view CLIP inference at test time unnecessary.
  • The monotonic gains as datasets are added suggest that further scaling of mask-text pairs will keep improving open-vocabulary semantic segmentation.
  • The zero-shot results with anonymized class names suggest that the model learns region semantics beyond memorized class labels, which is closer to true open-vocabulary behavior.
  • The caption-loss-aligned mask decoder could be reused as a proposal-free 3D vision-language interface for referring segmentation and other language-grounded tasks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the pipeline's quality holds, the same teacher-student recipe could extend to outdoor, dynamic, or object-centric 3D data without new annotation, since the projection step only needs posed RGB-D frames.
  • The paper's own zero-shot anonymization experiment implies that part of previous methods' apparent success came from class names leaking into training captions; Mosaic3D-5.6M appears less reliant on that leak, but the residual drop still leaves much of the open-vocabulary gain tied to caption content.
  • A direct human audit of a random sample of mask-caption pairs would be a cheap, decisive test of whether Osprey's descriptions are grounded in the masked region rather than hallucinated from context.
  • Combining the mask decoder with a proposal method like Segment3D suggests a modular path: any improved 3D proposal network could slot in and lift instance segmentation further.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper introduces an automatic data-generation pipeline that combines Grounded-SAM2/SEEM mask proposals with the Osprey region-aware VLM to create 3D mask-text pairs, and applies it to ScanNet, ARKitScenes, Matterport3D, ScanNet++, and Structured3D to construct Mosaic3D-5.6M (about 5.6M captions across about 29K scenes). It then trains a SparseUNet encoder with a point-text contrastive loss and a Mask3D-style decoder with a mask-caption loss, reporting state-of-the-art open-vocabulary semantic segmentation on ScanNet20, ScanNet200, Matterport3D, and ScanNet++, plus a single-stage open-vocabulary 3D instance segmentation method that does not require ground-truth labels. The paper includes data-scaling and model-scaling studies, component ablations, a class-name anonymization experiment, and an extensive supplementary appendix.

Significance. If the claims hold, Mosaic3D-5.6M would be a valuable community asset: it is substantially larger than existing mask-text datasets, the pipeline is automatic, and the authors provide a project page and extensive ablations that isolate the contributions of mask generators, captioners, frame sampling, data scale, model scale, and text encoders. The anonymization study is a useful attempt to separate class-name memorization from open-vocabulary generalization. The main weaknesses are the absence of direct human validation of the generated captions and masks, the internally inconsistent OV3D baseline, and the lack of clarity on whether evaluation scenes were excluded from the generated training data; these issues must be resolved before the central benchmark claims can be accepted.

major comments (3)
  1. [Section 3.2 / Table A1 / Section 5.3 / Table 1] The paper never states that the data-generation pipeline was restricted to the official training splits of the source datasets. Table A1 reports 1,513 ScanNet scenes, 2,194 Matterport3D scenes, and 380 ScanNet++ scenes, which appear to be the full datasets, while Table 1 evaluates Mosaic3D-5.6M on ScanNet20/200 validation, Matterport3D test, and ScanNet++ validation. If captions from evaluation scenes were included in training, the benchmark results are contaminated by scene-level leakage, because the model has seen the same point clouds and their captions during training. Please specify the exact train/validation/test split used for each source dataset, state whether any evaluation scene appears in Mosaic3D-5.6M, and re-run the affected benchmarks if leakage occurred.
  2. [Section 3.2 / Table A2 / Figure 2] The 'highest-quality dataset' claim is not directly validated. The metrics reported in Table A2, namely noun count, coverage, and entropy, cannot detect hallucinated captions or cross-view mask misalignment, and mask entropy rewards a partial mask covering a subset of one ground-truth instance as much as a complete mask of that instance, since both have low entropy. No human evaluation of caption correctness or mask-boundary quality is reported in the main text or the supplementary. Please add a human study on a random sample of mask-text pairs, reporting e.g. caption relevance/accuracy and mask IoU against ground-truth or manually refined masks, and state how the qualitative examples in Figure A2 were selected.
  3. [Table 1 / Appendix B.3 / Table A7] The OV3D comparison is internally inconsistent. Table 1 reports the original OV3D numbers of 64.0 f-mIoU on ScanNet20 and 8.7 on ScanNet200, while Table A7 shows that the authors' own reimplementation, OV3D-rep with DenseAlign, obtains only 34.7/4.6, and the improved OV3D++ with Contrastive obtains 58.4/9.2. The text's claim of 'surpassing OV3D by 1.0p' on ScanNet20 therefore depends on a baseline that the paper itself cannot reproduce. Please present the original-paper number and the same-protocol reproduction, ideally OV3D++ Contrastive under the shared SPUNet34C/Recap-CLIP setting, side by side in Table 1, and make the claims in the text consistent with the chosen baseline.
minor comments (7)
  1. [Abstract / Section 3.2 / Table A1] The abstract and Section 3.2 say 'over 30K scenes', but Table A1 sums to 29,197 scenes; please correct the count or rephrase to 'about 29K scenes'.
  2. [Section 3.2 / Table A1] Section 3.2 states 'approximately 1M RGB-D frames', while Table A1 reports a total of 7.1M frames; please reconcile these numbers.
  3. [Figure 2 / Table A2] Figure 2 reports Entropy 60.7 for Mosaic3D-5.6M, but Table A2 reports the same value for Mosaic3D-SN and leaves the entropy blank for Mosaic3D-5.6M; the figure caption should state the subset used for the entropy statistic.
  4. [Section 3.1 / Appendix A.3 / Table 3] The main text says 'Grounded-SAM' for mask generation, while Appendix A.3 and Table 3 use Grounded-SAM2 with SAM2 checkpoints; please unify the terminology throughout the paper.
  5. [Section 4.2 / Equation (5)] Equation (5) uses the subscript k in the numerator text embedding \bar ztext_k after defining the mask embeddings with subscript m; replace k with m for consistency.
  6. [Section 3.1 / Equation (1)] Equation (1) does not state that the projected pixel must lie inside the image bounds, and no sensitivity analysis is given for the depth threshold epsilon; please clarify the bounds check and, ideally, include an ablation over epsilon.
  7. [Tables 1-4] All benchmark tables report a single training run; given margins as small as 1.0 f-mIoU in Table 1, please report at least three seeds with mean and standard deviation for the key comparisons.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: benchmarks are external, and the anonymization experiment directly breaks the class-name loop.

full rationale

The paper's derivation chain is: (i) generate 3D mask-text pairs from 2D open-vocabulary segmentation models and a region-aware VLM via projection and inclusion test (Eq. 1); (ii) train a 3D encoder with a contrastive loss against a frozen text encoder (Eq. 2); (iii) train a mask decoder using caption-merged Segment3D masks; and (iv) evaluate on external benchmarks (ScanNet, ScanNet++, Matterport3D, ScanNet200). No equation reduces to a fitted parameter or to the benchmark numbers. The training objective does align point features to CLIP text embeddings, and evaluation also uses CLIP text embeddings of class names, but that is the definition of open-vocabulary segmentation, not a forced reduction: the 3D encoder could fail to learn the alignment, and the benchmarks provide independent human-annotated labels. The paper explicitly addresses the concern that captions contain evaluation class names with an anonymization experiment (Table 4), replacing class names with 'object' and showing that Mosaic3D-5.6M still retains the strongest performance; this is a direct control against the main potential circularity. The dataset-quality claim relies on the accuracy of teacher models (Grounded-SAM, SEEM, Osprey) without human evaluation of the generated pairs, and the mask-entropy metric in Table A2 measures homogeneity against GT instance IDs rather than caption correctness. These are validation gaps and correctness risks, not circularity by construction. Self-citations (e.g., MinkowskiNet for the sparse-conv backbone) are standard architectural references and are not load-bearing for the paper's central data-scaling or state-of-the-art claims. Consequently, the analysis finds no significant circularity.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new theoretical entities. Its claims rest on empirical assumptions about the quality of existing 2D models and the reliability of automated captioning and mask projection. These are domain assumptions, not fitted parameters in a theory.

free parameters (5)
  • Depth threshold epsilon in Eq. (1) = Not reported
    Used in the inclusion test to decide whether a projected 2D mask pixel matches the 3D point depth. Chosen by hand; the paper does not state its value.
  • IoU threshold tau in Algorithm 1 = Not reported
    Controls which generated mask-text pairs are merged with Segment3D masks. The value is not specified in the main text.
  • Number of frames per scene K = 25 or 125 in ablations
    Affects the diversity and redundancy of captions. The paper's ablation varies this, but the final choice is a hyperparameter.
  • Loss weights lambda_obj, lambda_dice, lambda_bce, lambda_cap = 2, 5, 2, 1
    Set manually in Section 4.2 for the mask decoder training objective.
  • Grounding-DINO thresholds and NMS = box score 0.25, text score 0.2, NMS IoU 0.5, max box area 95%
    Chosen by hand for filtering object proposals in the data engine, reported in Appendix A.3.
assumptions (5)
  • domain assumption 2D segmentation models Grounded-SAM and SEEM produce accurate open-vocabulary masks for the pipeline.
    Invoked in Section 3.1; the quality of the 3D dataset depends on these models' boundary precision.
  • domain assumption Osprey region captioning produces accurate, diverse, and contextually correct captions for masked regions.
    Invoked in Section 3.1; the richness of the dataset and downstream performance rely on caption quality.
  • domain assumption The projection and inclusion test in Eq. (1) correctly associates 2D masks to 3D points given accurate camera poses and depth.
    Invoked in Section 3.1; any systematic misalignment would generate noisy mask-text pairs.
  • domain assumption CLIP or Recap-CLIP text embeddings adequately represent caption semantics for contrastive learning.
    Invoked in Section 4.1; the training objective assumes these embeddings are a faithful semantic space.
  • domain assumption Segment3D class-agnostic masks provide a good proposal set for instance segmentation training.
    Invoked in Section 4.2 and Algorithm 1; the instance decoder is trained on these proposals.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation." pith.science (2026). https://pith.science/paper/QVN2ADES

@misc{pith2026250202548,
  author       = {Pith},
  title        = {Pith review of: Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QVN2ADES}},
  note         = {Machine review of arXiv:2502.02548}
}
read the original abstract

We tackle open-vocabulary 3D scene understanding by introducing a novel data generation pipeline and training framework. Our method addresses three critical requirements for effective training: precise 3D region segmentation, comprehensive textual descriptions, and sufficient dataset scale. By leveraging state-of-the-art open-vocabulary image segmentation models and region-aware Vision-Language Models, we develop an automatic pipeline that generates high-quality 3D mask-text pairs. Applying this pipeline to multiple 3D scene datasets, we create Mosaic3D-5.6M, a dataset of over 30K annotated scenes with 5.6M mask-text pairs, significantly larger than existing datasets. Building upon this data, we propose Mosaic3D, a foundation model combining a 3D encoder trained with contrastive learning and a lightweight mask decoder for open-vocabulary 3D semantic and instance segmentation. Our approach achieves state-of-the-art results on open-vocabulary 3D semantic and instance segmentation tasks including ScanNet200, Matterport3D, and ScanNet++, with ablation studies validating the effectiveness of our large-scale training data.

Figures

Figures reproduced from arXiv: 2502.02548 by the authors.

Figure 1
Figure 1. Mosaic3D-5.6M. Mosaic3D-5.6M is a large-scale dataset generated from a collection of existing datasets [7, 14, 24, 99, 107], consisting of 5.6M mask-text pairs, providing fine-grained masks (black outline in the figure) and detailed captions (text with matching color) pairs. Using this large-scale dataset, we propose Mosaic3D, a foundation model for open-vocabulary 3D segmentation. Abstract We tackle open-vocabulary… view at source ↗
Figure 2
Figure 2. Dataset comparison. We compare datasets using three metrics: # Nouns (the total number of unique normalized nouns in captions; higher is better), Coverage (the percentage of 3D points with associated captions per scene; higher is better), and Entropy (the entropy of GT instance ID distribution within masks; lower means more homogeneity - hense better). (a) Mosaic3D-5.6M uses precise masks (Entropy: 60.7) with region… view at source ↗
Figure 3
Figure 3. Mosaic3D-5.6M data engine. Our data generation process consists of three key steps: (a) We predict object segments for each RGB frame using state-of-the-art image segmentation models [50, 75, 110]. (b) We pass the images and predicted masks to a region-aware vision-language model [103] to generate descriptive captions for each region. (c) We project the 2D segmentation masks onto 3D points using camera parameters to… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Statistics of 3D mask-text datasets. We show the total number of scenes, tokens for generated captions. Our Mosaic3D￾5.6M significantly surpasses previous datasets in scale, combining multiple datasets to create the largest 3D mask-text dataset to date. frames, yieldin…
Figure 5
Figure 5. Figure 5: Mosaic3D model. Mosaic3D model is a SparseUNet [21] trained with our Mosaic3D-5.6M dataset to extract language-aligned features from 3D point clouds. A mask decoder with positional encodings (P.E) is trained on top to enable instance segmentation. 4.1. Mosaic3D: Langua…
Figure 6
Figure 6. Figure 6: Model performance scales with training data. We observe consistent improvements in open-vocabulary semantic seg￾mentation on ScanNet200 [78] as we increase the amount of train￾ing data. This shows the value of our large-scale data generation pipeline in improving open-…
Figure 7
Figure 7. Figure 7: Attention visualization of Mosaic3D as a 3D foundational model. From left to right, the first two examples show results on ScanNet [24], while the next two examples show results on ETH3D [80]. More examples are provided in the supplmentary material. level captioning ag…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

113 extracted references · 51 canonical work pages

  1. [1]

    https://huggingface

    Vit-gpt2 image captioning. https://huggingface. co / nlpconnect / vit - gpt2 - image - captioning, 2022. 2, 3

  2. [2]

    Gpt-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ah- mad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774,

  3. [3]

    Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas J. Guibas. ReferIt3D: Neural lis- teners for fine-grained 3d object identification in real-world scenes. In 16th European Conference on Computer Vision (ECCV), 2020. 2

  4. [4]

    Scanqa: 3d question answering for spatial scene understanding

    Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, and Motoaki Kawanabe. Scanqa: 3d question answering for spatial scene understanding. In proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19129– 19139, 2022. 2

  5. [5]

    Qwen technical report

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023. 2

  6. [6]

    Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 2023. 2

  7. [7]

    Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data

    Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Yuri Fei- gin, Peter Fu, Thomas Gebauer, Daniel Kurz, Tal Dimry, Brandon Joffe, Arik Schwartz, et al. Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), ...

  8. [8]

    Audiolm: A language modeling approach to audio generation

    Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Rob- lek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al. Audiolm: A language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2523–2533, 2023. 2

Show all 113 references
  1. [9]

    Large-scale machine learning with stochas- tic gradient descent

    Léon Bottou. Large-scale machine learning with stochas- tic gradient descent. In Proceedings of COMPSTAT’2010: 19th International Conference on Computational Statistic- sParis France, August 22-27, 2010 Keynote, Invited and Contributed Papers, pages 177–186. Springer, 2010. 6

  2. [10]

    Language models are few-shot learners

    Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. 2

  3. [11]

    Coyo-700m: Image-text pair dataset

    Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset. https : / / github . com / kakaobrain/coyo-dataset, 2022. 1

  4. [12]

    End- to-end object detection with transformers

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End- to-end object detection with transformers. In European conference on computer vision , pages 213–229. Springer,

  5. [13]

    Emerging properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jé- gou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021. 2

  6. [14]

    Matterport3d: Learning from rgb-d data in indoor environments

    Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Hal- ber, Matthias Niebner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. In 2017 International Confer- ence on 3D Vision (3DV), pages 667–676. IEEE Comp...

  7. [15]

    Scanrefer: 3d object localization in rgb-d scans using natural language

    Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner. Scanrefer: 3d object localization in rgb-d scans using natural language. In European conference on computer vision, pages 202–221. Springer, 2020. 2, 4, 5

  8. [16]

    Pali: A jointly-scaled multilingual language-image model

    Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergio- vanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. Pali: A jointly-scaled multilingual language-image model. In The Eleventh International Conference on Learning Represent...

  9. [17]

    Per-pixel classification is not all you need for semantic segmentation

    Bowen Cheng, Alexander G Schwing, and Alexander Kir- illov. Per-pixel classification is not all you need for semantic segmentation. In 35th Conference on Neural Information Processing Systems, NeurIPS 2021 , pages 17864–17875. Neural information processing systems foundation, ...

  10. [18]

    Masked-attention mask transformer for universal image segmentation

    Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexan- der Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1290–1299, 2022. 2, 5

  11. [19]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 2

  12. [20]

    Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation

    Seokju Cho, Heeseong Shin, Sunghwan Hong, Anurag Arnab, Paul Hongsuck Seo, and Seungryong Kim. Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4113–4123,

  13. [21]

    4d spatio-temporal convnets: Minkowski convolutional neural networks

    Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4d spatio-temporal convnets: Minkowski convolutional neural networks. In Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition , pages 3075–3084,

  14. [22]

    Scaling instruction- finetuned language models

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction- finetuned language models. Journal of Machine Learning Research, 25(70):1–53, 2024. 2

  15. [23]

    Pointcept: A codebase for point cloud perception research

    Pointcept Contributors. Pointcept: A codebase for point cloud perception research. https://github.com/ Pointcept/Pointcept, 2023. 2

  16. [24]

    Scannet: Richly-annotated 3d reconstructions of indoor scenes

    Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017. 1, 2, 4, ...

  17. [25]

    Procthor: Large-scale embodied ai using procedural generation

    Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. Procthor: Large-scale embodied ai using procedural generation. Ad- vances in Neural Information Processing Systems, 35:5982...

  18. [26]

    Pengi: An audio language model for audio tasks

    Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang. Pengi: An audio language model for audio tasks. Advances in Neural Information Processing Systems, 36:18090–18108, 2023. 2

  19. [27]

    Pla: Language-driven open- vocabulary 3d scene understanding

    Runyu Ding, Jihan Yang, Chuhui Xue, Wenqing Zhang, Song Bai, and Xiaojuan Qi. Pla: Language-driven open- vocabulary 3d scene understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 2, 3, 5, 6, 7

  20. [28]

    The llama 3 herd of models

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783,

  21. [29]

    Efficient graph-based image segmentation

    Pedro F Felzenszwalb and Daniel P Huttenlocher. Efficient graph-based image segmentation. International journal of computer vision, 59:167–181, 2004. 7

  22. [30]

    Dat- acomp: In search of the next generation of multimodal datasets

    Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al. Dat- acomp: In search of the next generation of multimodal datasets. Advances in Neural Information Processing Sys...

  23. [31]

    Scal- ing open-vocabulary image segmentation with image-level labels

    Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. Scal- ing open-vocabulary image segmentation with image-level labels. In European Conference on Computer Vision, pages 540–557. Springer, 2022. 1, 2

  24. [32]

    Imagebind one embedding space to bind them all

    Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. Imagebind one embedding space to bind them all. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, ...

  25. [33]

    3d semantic segmentation with submanifold sparse convolutional networks

    Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. 3d semantic segmentation with submanifold sparse convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 9224–9232, 2018. 5, 6

  26. [34]

    Open- vocabulary object detection via vision and language knowl- edge distillation

    Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open- vocabulary object detection via vision and language knowl- edge distillation. In International Conference on Learning Representations, 2021. 1

  27. [35]

    Regiongpt: Towards region understanding vision lan- guage model

    Qiushan Guo, Shalini De Mello, Hongxu Yin, Wonmin Byeon, Ka Chun Cheung, Yizhou Yu, Ping Luo, and Sifei Liu. Regiongpt: Towards region understanding vision lan- guage model. arXiv preprint arXiv:2403.02330, 2024. 2

  28. [36]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 3

  29. [37]

    Denoising diffu- sion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 2

  30. [38]

    Scaling up vision-language pre-training for image captioning

    Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang. Scaling up vision-language pre-training for image captioning. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17980–17989, 2022. 1

  31. [39]

    An Embodied Generalist Agent in 3D World, 2023

    Jiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu, Puhao Li, Yan Wang, Qing Li, Song-Chun Zhu, Baoxiong Jia, and Siyuan Huang. An Embodied Generalist Agent in 3D World, 2023. 8

  32. [40]

    Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels

    Rui Huang, Songyou Peng, Ayca Takmaz, Federico Tombari, Marc Pollefeys, Shiji Song, Gao Huang, and Francis Engel- mann. Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels. In European Confer- ence on Computer Vision, 2024. 2, 5, 7

  33. [41]

    Open-set image tagging with multi-grained text supervision

    Xinyu Huang, Yi-Jie Huang, Youcai Zhang, Weiwei Tian, Rui Feng, Yuejie Zhang, Yanchun Xie, Yaqian Li, and Lei Zhang. Open-set image tagging with multi-grained text supervision. arXiv e-prints, pages arXiv–2310, 2023. 2, 3, 8, 4

  34. [42]

    Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation

    Zhening Huang, Xiaoyang Wu, Xi Chen, Hengshuang Zhao, Lei Zhu, and Joan Lasenby. Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation. In European Conference on Computer Vision, 2024. 2, 5, 6, 7

  35. [43]

    Oneformer: One transformer to rule universal image segmentation

    Jitesh Jain, Jiachen Li, Mang Tik Chiu, Ali Hassani, Nikita Orlov, and Humphrey Shi. Oneformer: One transformer to rule universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2989–2998, 2023. 2

  36. [44]

    Scen- eVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding, 2024

    Baoxiong Jia, Yixin Chen, Huangyue Yu, Yan Wang, Xuesong Niu, Tengyu Liu, Qing Li, and Siyuan Huang. Scen- eVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding, 2024. 2, 6, 7, 8

  37. [45]

    Scaling up visual and vision-language representa- tion learning with noisy text supervision

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representa- tion learning with noisy text supervision. In International conference on machine learning, pages 4904–4916. PMLR,

  38. [46]

    Mistral 7b

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lam- ple, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023. 2

  39. [47]

    Mixtral of experts

    Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Deven- dra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024. 2

  40. [48]

    Open-vocabulary 3d semantic segmentation with foundation models

    Li Jiang, Shaoshuai Shi, and Bernt Schiele. Open-vocabulary 3d semantic segmentation with foundation models. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21284–21294, 2024. 2, 3, 6, 7, 8, 1, 4

  41. [49]

    In defense of lazy visual grounding for open-vocabulary semantic segmentation

    Dahyun Kang and Minsu Cho. In defense of lazy visual grounding for open-vocabulary semantic segmentation. In European Conference on Computer Vision, 2024. 2

  42. [50]

    Segment any- thing

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 4015–4026, 202...

  43. [51]

    Language-driven semantic seg- mentation

    Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic seg- mentation. In International Conference on Learning Repre- sentations, 2022. 1, 2

  44. [52]

    Semantic-sam: Segment and recognize anything at any gran- ularity

    Feng Li, Hao Zhang, Peize Sun, Xueyan Zou, Shilong Liu, Jianwei Yang, Chunyuan Li, Lei Zhang, and Jianfeng Gao. Semantic-sam: Segment and recognize anything at any gran- ularity. In European Conference on Computer Vision, 2024. 2

  45. [53]

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, pages 128...

  46. [54]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA , pages 19730...

  47. [55]

    What if we recaption billions of web images with llama-3? arXiv preprint arXiv:2406.08478,

    Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu, Huangjie Zheng, et al. What if we recaption billions of web images with llama-3? arXiv preprint arXiv:2406.08478,

  48. [56]

    Open-vocabulary semantic segmentation with mask-adapted clip

    Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-vocabulary semantic segmentation with mask-adapted clip. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pag...

  49. [57]

    Improved baselines with visual instruction tuning

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024. 8, 2

  50. [58]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Advances in neural information processing systems, 2024. 2, 3, 7

  51. [59]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023. 2, 3

  52. [60]

    Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annotations

    Ruiyuan Lyu, Jingli Lin, Tai Wang, Xiaohan Mao, Yilun Chen, Runsen Xu, Haifeng Huang, Chenming Zhu, Dahua Lin, and Jiangmiao Pang. Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annotations. Advances in Neural Information Processing Systems , 37: 50...

  53. [61]

    Multiscan: Scalable rgbd scanning for 3d environments with articulated objects

    Yongsen Mao, Yiming Zhang, Hanxiao Jiang, Angel Chang, and Manolis Savva. Multiscan: Scalable rgbd scanning for 3d environments with articulated objects. Advances in neural information processing systems, 35:9058–9071, 2022. 7

  54. [62]

    V-net: Fully convolutional neural networks for volumetric medical image segmentation

    Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV), pages 565–571. Ieee, 2016. 6

  55. [63]

    Silc: Improving vision language pretraining with self-distillation

    Muhammad Ferjad Naeem, Yongqin Xian, Xiaohua Zhai, Lukas Hoyer, Luc Van Gool, and Federico Tombari. Silc: Improving vision language pretraining with self-distillation. In European Conference on Computer Vision, 2024. 2

  56. [64]

    Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution

    Tuan Duc Ngo, Binh-Son Hua, and Khoi Nguyen. Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13550–13559, 2023. 7

  57. [65]

    Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance

    Phuc Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan, Anh Tran, Cuong Pham, and Khoi Nguyen. Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 4018–402...

  58. [66]

    Improved denoising diffusion probabilistic models

    Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International conference on machine learning, pages 8162–8171. PMLR,

  59. [67]

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rab...

  60. [68]

    Openscene: 3d scene understanding with open vocabularies

    Songyou Peng, Kyle Genova, Chiyu "Max" Jiang, An- drea Tagliasacchi, Marc Pollefeys, and Thomas Funkhouser. Openscene: 3d scene understanding with open vocabularies. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 6, 7, 5

  61. [69]

    Kosmos-2: Ground- ing multimodal large language models to the world

    Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Ground- ing multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023. 2, 3, 8

  62. [70]

    Language models are un- supervised multitask learners

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are un- supervised multitask learners. OpenAI blog, 1(8):9, 2019. 2

  63. [71]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of th...

  64. [72]

    Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai

    Santhosh Kumar Ramakrishnan, Aaron Gokaslan, Erik Wi- jmans, Oleksandr Maksymets, Alexander Clegg, John M Turner, Eric Undersander, Wojciech Galuba, Andrew West- bury, Angel X Chang, et al. Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai....

  65. [73]

    Denseclip: Language-guided dense prediction with context- aware prompting

    Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. Denseclip: Language-guided dense prediction with context- aware prompting. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 1808...

  66. [74]

    Glamm: Pixel grounding large multimodal model

    Hanoona Rasheed, Muhammad Maaz, Sahal Shaji, Abdel- rahman Shaker, Salman Khan, Hisham Cholakkal, Rao M Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S Khan. Glamm: Pixel grounding large multimodal model. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Patter...

  67. [75]

    Sam 2: Segment anything in images and videos

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024. 2, 3, 4, 8

  68. [76]

    Grounded sam: Assembling open-world models for diverse visual tasks

    Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, et al. Grounded sam: Assembling open-world models for diverse visual tasks. arXiv preprint arXiv:2401.14159,

  69. [77]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2

  70. [78]

    Language- grounded indoor 3d semantic segmentation in the wild

    David Rozenberszki, Or Litany, and Angela Dai. Language- grounded indoor 3d semantic segmentation in the wild. In European Conference on Computer Vision, pages 125–141. Springer, 2022. 6, 7, 8, 2, 4, 5

  71. [79]

    Audiopalm: A large language model that can speak and listen

    Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al. Audiopalm: A large language model that can speak and listen. arXiv preprint arXiv:2306.12925,

  72. [80]

    A multi-view stereo benchmark with high- resolution images and multi-camera videos

    Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and An- dreas Geiger. A multi-view stereo benchmark with high- resolution images and multi-camera videos. In CVPR, 2017. 8

  73. [81]

    Laion-5b: An open large-scale dataset for train- ing next generation image-text models

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, et al. Laion-5b: An open large-scale dataset for train- ing next generation image-text models. Advances in Neural Inf...

  74. [82]

    Mask3d: Mask trans- former for 3d semantic instance segmentation

    Jonas Schult, Francis Engelmann, Alexander Hermans, Or Litany, Siyu Tang, and Bastian Leibe. Mask3d: Mask trans- former for 3d semantic instance segmentation. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 8216–8223. IEEE, 2023. 2, 5, 7

  75. [83]

    Clip-fields: Weakly supervised semantic fields for robotic memory

    Nur Muhammad Mahi Shafiullah, Chris Paxton, Lerrel Pinto, Soumith Chintala, and Arthur Szlam. Clip-fields: Weakly supervised semantic fields for robotic memory. InICRA2023 Workshop on Pretraining for Robotics (PT4R), 2023. 1

  76. [84]

    Super-convergence: Very fast training of neural networks using large learning rates

    Leslie N Smith and Nicholay Topin. Super-convergence: Very fast training of neural networks using large learning rates. In Artificial intelligence and machine learning for multi-domain operations applications, pages 369–386. SPIE,

  77. [85]

    Denois- ing diffusion implicit models

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models. In International Conference on Learning Representations, 2021. 2

  78. [86]

    Score-based generative modeling through stochastic differential equa- tions

    Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions. In International Conference on Learning Representa- tions, 2021. 2

  79. [87]

    Open- mask3d: open-vocabulary 3d instance segmentation

    Ayça Takmaz, Elisabetta Fedele, Robert W Sumner, Marc Pollefeys, Federico Tombari, and Francis Engelmann. Open- mask3d: open-vocabulary 3d instance segmentation. In Proceedings of the 37th International Conference on Neural Information Processing Systems, pages 68367–68390, 20...

  80. [88]

    Gemini: a family of highly capable multimodal models

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. 2

  81. [89]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text. arXiv preprint arXiv:2403.05530, 2024. 2

  82. [90]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Mar- tinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 2

  83. [91]

    Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023. 2

  84. [92]

    Rio: 3d object instance re- localization in changing indoor environments

    Johanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari, and Matthias Nießner. Rio: 3d object instance re- localization in changing indoor environments. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 7658–7667, 2019. 7

  85. [93]

    Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learn- ing framework

    Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learn- ing framework. In International Conference on Machine Learni...

  86. [94]

    EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI, 2023

    Tai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu, Ruiyuan Lyu, Peisen Li, Xiao Chen, Wenwei Zhang, Kai Chen, Tianfan Xue, Xihui Liu, Cewu Lu, Dahua Lin, and Jiangmiao Pang. EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI, 2023. 2, 8

  87. [95]

    Towards large- scale 3d representation learning with multi-dataset point prompt training

    Xiaoyang Wu, Zhuotao Tian, Xin Wen, Bohao Peng, Xihui Liu, Kaicheng Yu, and Hengshuang Zhao. Towards large- scale 3d representation learning with multi-dataset point prompt training. arXiv preprint arXiv:2308.09718, 2023. 6, 4

  88. [96]

    Groupvit: Semantic segmentation emerges from text supervision

    Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18134–18144, 2022. 1

  89. [97]

    Qwen2 technical report.arXiv preprint arXiv:2407.10671, 2024

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. Qwen2 technical report.arXiv preprint arXiv:2407.10671, 2024. 2

  90. [98]

    Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding

    Jihan Yang, Runyu Ding, Weipeng Deng, Zhe Wang, and Xi- aojuan Qi. Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 2, 3, 5, 6, 7, 8

  91. [99]

    Scannet++: A high-fidelity dataset of 3d in- door scenes

    Chandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, and Angela Dai. Scannet++: A high-fidelity dataset of 3d in- door scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12–22, 2023. 1, 2, 4, 6, 7, 5

  92. [100]

    Sai3d: Segment any instance in 3d scenes

    Yingda Yin, Yuzheng Liu, Yang Xiao, Daniel Cohen-Or, Jingwei Huang, and Baoquan Chen. Sai3d: Segment any instance in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3292–3302, 2024. 5, 7

  93. [101]

    Ferret: Refer and ground anything anywhere at any granularity

    Haoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du, Bowen Zhang, Zirui Wang, Liangliang Cao, Shih-Fu Chang, and Yinfei Yang. Ferret: Refer and ground anything anywhere at any granularity. arXiv preprint arXiv:2310.07704, 2023. 2, 3, 8

  94. [102]

    Convolutions die hard: open-vocabulary seg- mentation with single frozen convolutional clip

    Qihang Yu, Ju He, Xueqing Deng, Xiaohui Shen, and Liang- Chieh Chen. Convolutions die hard: open-vocabulary seg- mentation with single frozen convolutional clip. In Pro- ceedings of the 37th International Conference on Neural Information Processing Systems, pages 32215–32234, 2023. 2

  95. [103]

    Osprey: Pixel understanding with visual instruction tuning

    Yuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, and Jianke Zhu. Osprey: Pixel understanding with visual instruction tuning. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 28202–28211, 2024. 2, 3, 4, 8

  96. [104]

    Sigmoid loss for language image pre-training

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11975–11986, 2023. 1, 4

  97. [105]

    Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities

    Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023. 2

  98. [106]

    Recognize anything: A strong image tagging model

    Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li, Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo, Yaqian Li, Shilong Liu, et al. Recognize anything: A strong image tagging model. arXiv preprint arXiv:2306.03514 ,

  99. [107]

    Structured3d: A large photo-realistic dataset for structured 3d modeling

    Jia Zheng, Junfei Zhang, Jing Li, Rui Tang, Shenghua Gao, and Zihan Zhou. Structured3d: A large photo-realistic dataset for structured 3d modeling. In Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, Au- gust 23–28, 2020, Proceedings, Part IX 16, pages 519–53...

  100. [108]

    Detecting twenty-thousand classes using image-level supervision

    Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähen- bühl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. In European Conference on Computer Vision, pages 350–368. Springer, 2022. 8

  101. [109]

    Generalized decoding for pixel, image, and language

    Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jianfeng Wang, Lu Yuan, et al. Generalized decoding for pixel, image, and language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 151...

  102. [110]

    3D Object Detection (3DOD)

    Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee. Segment everything everywhere all at once. Advances in Neural Information Processing Systems, 36, 2024. 2, 3, 4, 7, 8 Mosaic3D: Foundation Dataset and Model for...

  103. [111]

    In the first round, LLaV A-1.5 is prompted to generate an image caption describing the overall scene

  104. [112]

    In the second round, LLaV A-1.5 is prompted to extract entity names from the generated image caption

  105. [113]

    entity name A

    In the final round, LLaV A-1.5 is prompted to generate de- tailed entity descriptions for each extracted entity name. During our implementation, we encountered inconsistencies in LLaV A-1.5’s response formats. To ensure structured and consistent entity-level text descriptions,...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.