Pith. sign in

Paper Citation Record · LEDGER

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation

As of 10 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 0 inbound Pith citation observations for arXiv:2502.02548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02548 v2

Coverage vector

measured 100 of 113 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:52:14.978847Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 113 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 00e19b1d-8dcb-4811-945b-7b9a620ece04 · outbound

This paper cites https://huggingface.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation https://huggingface

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.562340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.562340Z digest=sha256:44565fb0ff15d8b448075489f76d4b6608d3acf7c63a62c0d0cc9a5a273a7e1c

Observation 0de5cbea-bd04-4fda-8e9d-fbbf712568fe · outbound

This paper cites GPT-4 Technical Report.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.567886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.567886Z digest=sha256:f9f6aa0a213c6036cd61984e2dbe0089ebcf6d5f352fceafe86177aaa2f80c08

Observation 0e29b170-02c0-4df9-bfb0-671e370e82f9 · outbound

This paper cites an unresolved cited work.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.572923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.572923Z digest=sha256:03ebecc4c35e32c45bc1e886357c2a36435fb62fea403b71aa2678bf329b7b31

Observation 3e1d639c-f057-45ae-9757-e97df963a3fd · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scanqa: 3d question answering for spatial scene understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.577648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.577648Z digest=sha256:2b7d18a2919c560a40f21f67f0742ff322b43b8c662ed2a294ec5368ff45f77f

Observation 4c4b582d-deda-451e-a389-2e911d10581c · outbound

This paper cites Qwen Technical Report.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.581601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.581601Z digest=sha256:d6b653c833109507a6cdf1ece75eefb8b1bfa47e73506c5823971985212af2f9

Observation 0ba80f78-5e9c-4d59-b1a4-d70e6a00b12e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.586476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.586476Z digest=sha256:b55422ea9cb19ce01815189d32d6e86ebdb041aa71ee0b5606d40953aba092bc

Observation 3639b949-16f6-4b95-9744-eef372a2b615 · outbound

This paper cites Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Arkitscenes: A diverse real-world dataset for 3d indoor scene understanding using mobile rgb-d data

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.590946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.590946Z digest=sha256:098c4e5148555dbad8e708ca29b4fae839b2730206aa872e42282026fe39e721

Observation 29b0a2c2-4039-427f-8956-128928764e1c · outbound

This paper cites Audiolm: A language modeling approach to audio generation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Audiolm: A language modeling approach to audio generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.595991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.595991Z digest=sha256:b0586d3fab2a49fcde079ec8c639f9f54ebd34f387d8eecfb5f0dc3dfd91d5a8

Observation d8d940d8-b557-43f7-b49d-8964fbdf988e · outbound

This paper cites Large-scale machine learning with stochas- tic gradient descent.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Large-scale machine learning with stochas- tic gradient descent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.600157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.600157Z digest=sha256:f085f1663d38cc73a8cc5c791801956bd7eb7a2ca6e25e9e389db2c5ff54bc46

Observation ad369f0b-c399-48fa-95d1-ffc270cccd52 · outbound

This paper cites Language Models are Few-Shot Learners.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Language Models are Few-Shot Learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.604539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.604539Z digest=sha256:6ceda0fb31d9f5ba0795bfadac480faf0a8fca1b14382aa652eb020a5918e746

Observation 14e98da3-0684-481e-b882-6cfd28da27f6 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Coyo-700m: Image-text pair dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.609153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.609153Z digest=sha256:5e2b6b89a4bb90b7e178ebde9324d6f64bbe4d1b55382bf630fc74e2005a9353

Observation 221ae02b-dd6b-4cb6-81c1-a3bb34e3218c · outbound

This paper cites End- to-end object detection with transformers.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation End- to-end object detection with transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.613366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.613366Z digest=sha256:8adb025d9c6fa2c1ed4726c9d800668ebc7284c3de6f312ff1d17f9fd91a834e

Observation 52057327-0ffd-42f7-ae4d-bb2d7947d518 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Emerging properties in self-supervised vision transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.617847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.617847Z digest=sha256:4edde3132f8f44f3938f3e3cf9a8f81684fce3a88efd28823012e92ed033630b

Observation a12d3e7a-41c0-4542-990d-1f922873e342 · outbound

This paper cites Matterport3d: Learning from rgb-d data in indoor environments.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Matterport3d: Learning from rgb-d data in indoor environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.622107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.622107Z digest=sha256:db902988b7a153a6687fd7cd76aa0314d237fd43960213082865c27459fab373

Observation 0b88250b-1c5f-434a-9913-e08b0c1daf13 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.626298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.626298Z digest=sha256:35347e2a7f387c79e40676a0664814adc57160cd0db8ae20555c5a71964dcd99

Observation c003712d-e347-4713-aeb7-707cc7e791a3 · outbound

This paper cites Pali: A jointly-scaled multilingual language-image model.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Pali: A jointly-scaled multilingual language-image model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.630797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.630797Z digest=sha256:6c73f95f8dd20363c158e634be1b92e2a2d29178cf904a3aa0013ea483c7cc38

Observation bfe2f05f-8f82-4f3c-94e2-f244efeae5c5 · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Per-pixel classification is not all you need for semantic segmentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.634940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.634940Z digest=sha256:55fdc0e19ba100a474338b6c4231e073be41e30138e8917adf7d1d953669eb1b

Observation f96db358-1b11-4f9d-b955-f74a13d91fa5 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Masked-attention mask transformer for universal image segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.639318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.639318Z digest=sha256:b9f4e394c987746260310cf2014a414ace9b001d027eea89125815ed3ee63012

Observation 2f5ad14a-f21c-4b58-9ee0-b4ff22b0adde · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Gonzalez, Ion Stoica, and Eric P

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.643885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.643885Z digest=sha256:d6837cf0f90949828b98d6515322c92b2b20495053703dde8d6a4eb99fca1c0f

Observation 6dd0943b-d02f-406f-a495-24fdc474956f · outbound

This paper cites Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Cat- seg: Cost aggregation for open-vocabulary semantic seg- mentation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.648054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.648054Z digest=sha256:2a4276c9ea0122e4a3546361842ba5fcc59487ce9d517e560dcefab75852cd94

Observation 0cd8e49c-106c-4e32-8721-6a31597b6b33 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.652521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.652521Z digest=sha256:aa81e5d17b748f2f3426412e53b4bbd033f323ed50f70cb0c1bbdb0e81687862

Observation 4fef109b-e5a2-4649-af72-6c41e31eaafb · outbound

This paper cites Scaling instruction- finetuned language models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scaling instruction- finetuned language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.656602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.656602Z digest=sha256:c6e0798f41d39d42f3f8c5a2563d74b8757198fdae53d53f87adc92ea114202a

Observation cf04d106-fd56-40a3-a7d0-198382c3c0c7 · outbound

This paper cites Pointcept: A codebase for point cloud perception research.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Pointcept: A codebase for point cloud perception research

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.660694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.660694Z digest=sha256:593a95f951daab07a88c526380ac36c1b7217a9313ba5404030329e433b0eb03

Observation dcbaa040-18d9-4e97-9961-aa7248af709c · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.664581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.664581Z digest=sha256:2501695468f19a625819154ab8274fe2f624ac156b463f5fe15eef84789afa52

Observation 853ff7e2-95c8-42b9-907e-7b55d17ea351 · outbound

This paper cites Procthor: Large-scale embodied ai using procedural generation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Procthor: Large-scale embodied ai using procedural generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.668572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.668572Z digest=sha256:59403f71f5d73c454461516d0c484f081c3c6d66e1026118b49744c4f6d6fbc4

Observation a9b89b0c-746c-44c2-9241-9152b147eca2 · outbound

This paper cites Pengi: An audio language model for audio tasks.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Pengi: An audio language model for audio tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.672837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.672837Z digest=sha256:a932f1f58546c06f244204de6eaaf99d531d621279fa5e63c89c179176f38498

Observation 874ef941-752f-4003-b50c-849466114b63 · outbound

This paper cites Pla: Language-driven open- vocabulary 3d scene understanding.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Pla: Language-driven open- vocabulary 3d scene understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.677012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.677012Z digest=sha256:251f7876ed71ab0fb63741b98676a1ca17fc0b8d762f7967a886b8becb9bb016

Observation 90edc446-c1b2-4171-a472-a8a6645caadf · outbound

This paper cites The Llama 3 Herd of Models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation The Llama 3 Herd of Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.681025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.681025Z digest=sha256:906aa8d1ae650430113f6cba559e7fbdae7041eab24e314cf10858926e763ebc

Observation d2a0e2a6-2a47-4af7-a4d8-2eaa6b1782f0 · outbound

This paper cites Efficient graph-based image segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Efficient graph-based image segmentation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.685595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.685595Z digest=sha256:f4bf9b3b3962ff4451c4c923b88d183abd2ad7a866f17216a0ce1fd911438899

Observation e7bb280a-4088-4615-a105-42f71715cb38 · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Dat- acomp: In search of the next generation of multimodal datasets

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.689680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.689680Z digest=sha256:b17988182667bf410aed7846ea1ba04fad8bdb3c7a8c334a2fb916e5ce5b8504

Observation affb818f-864c-437a-8742-acb3439b38eb · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scal- ing open-vocabulary image segmentation with image-level labels

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.693601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.693601Z digest=sha256:0154b85cb2be2e84fdbd57ccb63fe57d613934effe92504571064661a431dce4

Observation 82c286eb-35b4-447f-9fd1-96c11635699c · outbound

This paper cites Imagebind one embedding space to bind them all.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Imagebind one embedding space to bind them all

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.697384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.697384Z digest=sha256:6f529f48036896cc04eddcdb1b0a68f625a97e86d06fa9f5348edd6b0426cdb9

Observation 80b722c7-3dac-4852-8e14-757d560df0e8 · outbound

This paper cites 3d semantic segmentation with submanifold sparse convolutional networks.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation 3d semantic segmentation with submanifold sparse convolutional networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.701027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.701027Z digest=sha256:193aec9b0d18a3599858a9380b8167b868a4bf72e72541216fe9aa64c2f3bf55

Observation 275aa17b-4b30-422c-8f52-a1ebc99847df · outbound

This paper cites Open- vocabulary object detection via vision and language knowl- edge distillation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open- vocabulary object detection via vision and language knowl- edge distillation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.705118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.705118Z digest=sha256:24fb4796eb5057ec13b02ecb53dd3977622f1527e1e9dfc57436dbcf27cc390e

Observation 29f12a23-69f1-4df5-9e6f-7afb377fbebb · outbound

This paper cites RegionGPT: Towards Region Understanding Vision Language Model.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation RegionGPT: Towards Region Understanding Vision Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.708786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.708786Z digest=sha256:4bc07c343c5c2534851658bd9911c1a71ff0d8552278fbf1621983132bedc9a3

Observation 06895320-52ee-4f98-90a0-bc1db054f058 · outbound

This paper cites Deep residual learning for image recognition.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Deep residual learning for image recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.712783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.712783Z digest=sha256:e31f0782d092e6b3643cdff4008e8518b1e5c4a93856ae7dbeefe703460bc7a8

Observation a3949f46-5039-46ba-9463-e09527f2f633 · outbound

This paper cites Denoising diffu- sion probabilistic models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Denoising diffu- sion probabilistic models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.716714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.716714Z digest=sha256:867415aafda1f6c157ea7ff9eb476df990f462d1816bb0e0616dfa6e15d6c5cc

Observation cecbf8e7-2c25-4899-93dc-25671c5a93bb · outbound

This paper cites Scaling up vision-language pre-training for image captioning.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scaling up vision-language pre-training for image captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.720613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.720613Z digest=sha256:78260b3936580f3b8204ec489ba1c046d73d4b46f03b2e0612d48e55a654cf81

Observation a74ad816-a918-4526-bef7-94ef9b703b70 · outbound

This paper cites An Embodied Generalist Agent in 3D World, 2023.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation An Embodied Generalist Agent in 3D World, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.724428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.724428Z digest=sha256:ae7046cad209c3f321c4031831d512b7438e0b9658989f0fc0811739a314e607

Observation cfc6e54a-bb57-437c-9fab-32911b613dc9 · outbound

This paper cites Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Segment3d: Learning fine-grained class-agnostic 3d segmentation without manual labels

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.268078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.728313Z digest=sha256:c81ff3f0b14fb8ac6a88803031ac59cc960174d8d05519337aeb726efb45c14e

Observation bfa33195-a310-44db-ab51-a169911857d5 · outbound

This paper cites Open-set image tagging with multi-grained text supervision.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open-set image tagging with multi-grained text supervision

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.252677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.732287Z digest=sha256:7d7d9060f8f6c6cee9fa6b96f747af184551ef256492ee61ff22d4ab1e33b6f8

Observation daf2bf4e-38a0-4e7a-8daa-9dfb26f4a3d0 · outbound

This paper cites Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Openins3d: Snap and lookup for 3d open-vocabulary instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.237410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.737048Z digest=sha256:cf70deb855490451ffc0cf8c52dfbd2578c62ce7df913ced44e8ab604954785d

Observation b9c9593d-abe8-4039-91b1-fee9aa4083a0 · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Oneformer: One transformer to rule universal image segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.741058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.741058Z digest=sha256:8000e2cb2a387463a72818fc8f623c7703db8f48ed27089397f3514e12c70ca9

Observation 4a45b972-ec9e-4c94-aa03-1ec35f0ccc6b · outbound

This paper cites Scen- eVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding, 2024.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scen- eVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.210733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.745098Z digest=sha256:7cc856b14051bc02d3466c8a4d692b3ba7860447d6655017f0811d1c327c77dd

Observation e412da27-9b92-4f97-a127-2174cc369ab5 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.749054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.749054Z digest=sha256:ea31acc6db4a2427b2d68f52578b0de9aabfc28036f8fc25226bc890f1c6ff57

Observation 97cd7bfa-79b8-40f6-83a8-69edb7645e72 · outbound

This paper cites Mistral 7B.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Mistral 7B

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.753930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.753930Z digest=sha256:a96d317f9d620c50d729d01a398cd268c4e0bf14411c91c7c5a0e35567ef94eb

Observation a96f6ac5-9a3b-4aa2-b53e-1f4f34c88b0c · outbound

This paper cites Mixtral of Experts.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Mixtral of Experts

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.758083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.758083Z digest=sha256:e212ee1a90d2ea4779c57a44ad3678c55625f830eee8f2eff3ab247a396260dc

Observation 42d0e6dc-74c8-426e-a0cf-35e1a991544c · outbound

This paper cites Open-vocabulary 3d semantic segmentation with foundation models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open-vocabulary 3d semantic segmentation with foundation models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.183318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.762576Z digest=sha256:77933f9247fe858e5953053054f89c8429b55e5a6069121b29c08a5439c86f75

Observation 917ac727-ee91-446c-8584-4546bcfa6273 · outbound

This paper cites In defense of lazy visual grounding for open-vocabulary semantic segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation In defense of lazy visual grounding for open-vocabulary semantic segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.167880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.766649Z digest=sha256:971ced71b3880f09f04c6f94ee9de425783174487cb3414f196bf578b547c811

Observation 8d77640e-e6b6-4ee5-a438-891fe4fa787c · outbound

This paper cites Segment any- thing.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Segment any- thing

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.153388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.770955Z digest=sha256:75035bd11bb828cbd221ac411bfc5aab381452b8b5fc7a89c51b11521b606f38

Observation 4c119a91-cee5-4b8d-ac9f-154d3a7c83aa · outbound

This paper cites Language-driven semantic seg- mentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Language-driven semantic seg- mentation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.138563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.775337Z digest=sha256:8476f534f723e7cb48d92bd782135533d02d587be9698260ec97e8d76b477ca4

Observation 66a26ccf-570b-4106-bad7-c82817705a69 · outbound

This paper cites Semantic-sam: Segment and recognize anything at any gran- ularity.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Semantic-sam: Segment and recognize anything at any gran- ularity

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.122887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.779553Z digest=sha256:6a01fc2629b14d90ab4517599354655b5947adca56d219e70ef2db49e8b49f9e

Observation bd214a08-26b4-432c-a489-3fb65d964819 · outbound

This paper cites an unresolved cited work.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:52:16.107193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.783713Z digest=sha256:2a7dfa6455eeaca8b945abe3388d2e6666a5037282e88f6e9365c1a11a16e19f

Observation 4d3bdca8-f72a-4e4d-8435-bdcf36ef16e2 · outbound

This paper cites an unresolved cited work.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:52:16.090482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.788044Z digest=sha256:a6749ec7e135d9b496496a377dec1491f9efbd3d24a7957d542f9e8bb1c5ac2e

Observation 7d1262d9-2858-43e1-ad4d-6566f6f7459e · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.792125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.792125Z digest=sha256:38df75a3d5fee1b4953c56d060dc0806abf7ab8152b482000cbff43547f0a4dd

Observation 8d8d16bd-341f-4c4e-84f8-b04f7a8fb787 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open-vocabulary semantic segmentation with mask-adapted clip

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.073384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.796758Z digest=sha256:70f57bb3ff413a876105230e6614363599f35593670996091e63e98a27b43d50

Observation cafe5ede-5d40-403f-86f5-78a1cc20a36f · outbound

This paper cites Improved baselines with visual instruction tuning.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Improved baselines with visual instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.057811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.801329Z digest=sha256:a88374f09d1f01f7cdfc097ee79e4b9fb7a13b52e951d32c14e1d16a88749950

Observation 62176497-bbee-443c-b9d7-63956413a7b2 · outbound

This paper cites Visual instruction tuning.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Visual instruction tuning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.042655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.805634Z digest=sha256:fa8a0d9516a6b03d737da2ab43dfea6ce93144f4f510e4bcebd5e697be424a23

Observation 944f78de-385d-4098-9ecf-8240a8a9d096 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.809674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.809674Z digest=sha256:37af8ca97e9ff85783cd43528a006e58baec9a02b4c3a2d9275dc7db9298d798

Observation a038eb75-8d16-46f8-a685-26915a8ddbf2 · outbound

This paper cites Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annotations.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Mmscan: A multi-modal 3d scene dataset with hierarchical grounded language annotations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.027246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.813845Z digest=sha256:168a96136b9dff0f496bf6238886785632268b96601228e67c5ac3f0eb5c54b7

Observation 8a65f695-9d47-4f26-9e95-0e56413c5246 · outbound

This paper cites Multiscan: Scalable rgbd scanning for 3d environments with articulated objects.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Multiscan: Scalable rgbd scanning for 3d environments with articulated objects

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:16.011839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.817951Z digest=sha256:43a81c2dc32a090877d86ebf7e5f13699ad8cd82dc90fd0b84d5b59027972126

Observation 0ef770a6-86a3-494f-bfaa-2c53ab59a942 · outbound

This paper cites V-net: Fully convolutional neural networks for volumetric medical image segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation V-net: Fully convolutional neural networks for volumetric medical image segmentation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.821919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.821919Z digest=sha256:b0ae36ece47792ddb3eb487f2c43b539fe023aaa43a82d9cb9d18873d1c6564a

Observation 4f8d6c61-81f5-4fba-bf9e-c89c1417b0c3 · outbound

This paper cites Silc: Improving vision language pretraining with self-distillation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Silc: Improving vision language pretraining with self-distillation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.983110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.826052Z digest=sha256:8c1d2f7be2557205838c02139fffe6bcb8749888ca3cc44eb287b2e483080635

Observation d331ff8d-a1fd-4598-8940-d1f297e3f0e6 · outbound

This paper cites Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Isbnet: a 3d point cloud instance segmentation network with instance- aware sampling and box-aware dynamic convolution

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.967578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.830000Z digest=sha256:824f92fc7572f7589d770b6451b3287b1975617a86d83c6c254be9dbab4e62ab

Observation 52f7004c-a463-42a1-bf5f-11bcb7e40aa8 · outbound

This paper cites Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.951605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.833920Z digest=sha256:cdc786755b5c4fd7de3120cfa972835e3275d9eb9c7704e570ed75268b82724b

Observation db49205e-1102-46ab-ae86-1f6b00356dcc · outbound

This paper cites Improved denoising diffusion probabilistic models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Improved denoising diffusion probabilistic models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.838076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.838076Z digest=sha256:de84a3dd9193772ad0d1a349193a9de8b0342b6ac8b15c5fe2245b7dcdf8f6db

Observation 437c7dd3-6051-45ba-9b30-a6969133c011 · outbound

This paper cites an unresolved cited work.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:52:15.926916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.842102Z digest=sha256:d208264eb07c5b1937e353e433593708585774201dfa8cff74d9899af7042893

Observation 227a1564-54f1-4295-aa9e-21da05878fa6 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Openscene: 3d scene understanding with open vocabularies

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.911926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.846000Z digest=sha256:6b70320874e8e78bc9e2e57b030a4c849c10951c6b9634d8ab23f49d0b055e5e

Observation e3969332-b124-442c-9ea3-239664801938 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.850056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.850056Z digest=sha256:d1d81c44fae0ff21903408ef3c7b520bdb7678036e5d6cb77baf3124d654c97f

Observation 22966ade-2886-49b3-abe2-24aae0b57603 · outbound

This paper cites Language models are un- supervised multitask learners.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Language models are un- supervised multitask learners

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.895454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.854046Z digest=sha256:e984762d5fe39b8f54252b2b9584773f4f86fa7cd04c57af9e3ad00d02c81cbb

Observation de4e563d-3480-4b5d-8558-0a78c389876e · outbound

This paper cites Learning transferable visual models from natural language supervision.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Learning transferable visual models from natural language supervision

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.858027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.858027Z digest=sha256:1fd38afdabbac415013ec722e8caa1b304b75d133bd71b787c5b9aabe707726c

Observation dbc9ddaf-8ff0-46df-b640-d0ec3ed21df0 · outbound

This paper cites Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.868396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.862319Z digest=sha256:da3e1ea1b12b9bec5503f2059a4486a2d6fef5d6ac43709e0ece147241748bdd

Observation 7e015112-44c9-4cd7-b4a4-ab2e5fb1ec0d · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Denseclip: Language-guided dense prediction with context- aware prompting

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.852879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.866163Z digest=sha256:b76dfbadaaec4e5c00d4cc985e00da8dd5d708a00619c9c9ec6194b82812bf86

Observation 680b86ee-8ea0-4327-b090-2a045849b70e · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Glamm: Pixel grounding large multimodal model

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.838090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.870527Z digest=sha256:cf81af2f908bafd654abd31cb636552aceef597985d2aecf568ab9ae5b8e690e

Observation da088572-6e7b-41b3-bd14-4a7eebcd8b9d · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation SAM 2: Segment Anything in Images and Videos

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.874411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.874411Z digest=sha256:f8b84f71cb55ce292b09f16e4c7c9902ef9ccc85b4f3012aab1a7eb9c258ef6e

Observation bea92ce4-0e2f-48d5-bfc5-03a1add15019 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.878648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.878648Z digest=sha256:308b3d13d64b3ed9866535403fefaa56a8f1f13db192d0a86c25633855f80f50

Observation 7479ea9b-c09b-45e0-9c16-0c470f089aaf · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation High-resolution image synthesis with latent diffusion models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.882773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.882773Z digest=sha256:d8b2fd3008436f59a0a911b6448625367f82086166bebfe322f035f20ffa5ad3

Observation 0ac63485-9b44-4471-86aa-71f807630c69 · outbound

This paper cites Language- grounded indoor 3d semantic segmentation in the wild.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Language- grounded indoor 3d semantic segmentation in the wild

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.812509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.887128Z digest=sha256:d3d3d4532ee1b54902b52d1c00525b9923b03ead86ae80140a3c9595ad011564

Observation b75cad6e-adb1-4754-95cd-024c8708cfb9 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.891226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.891226Z digest=sha256:7497944e39a4891a6e8222dab657b9c599f7ac6a48064ffe53a1fed94a94ebe3

Observation eba8f216-4af8-4c68-a3f2-efbf85eb62a0 · outbound

This paper cites A multi-view stereo benchmark with high- resolution images and multi-camera videos.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation A multi-view stereo benchmark with high- resolution images and multi-camera videos

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.796445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.895671Z digest=sha256:53ef88a594e8b770549b45beded6f9268bf1f827580559d87b848b4eaad39b92

Observation bd752179-a75a-4e49-ba6f-49596dfcaea6 · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Laion-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.781293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.899890Z digest=sha256:d44e1d06152becabc239a2352ecd59e9c96f51df4ae24a3a627fe1fdb57b4161

Observation d72dade2-eaf3-41b5-99fb-0a94abe84dc8 · outbound

This paper cites Mask3d: Mask trans- former for 3d semantic instance segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Mask3d: Mask trans- former for 3d semantic instance segmentation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.766568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.904387Z digest=sha256:ea01c647e5cdd090b6db426462705d2b26d2f9e6716dcf19a16f48556d6a4ef4

Observation 23570bcc-5127-4d60-a5d3-c548bbd66bc3 · outbound

This paper cites Clip-fields: Weakly supervised semantic fields for robotic memory.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Clip-fields: Weakly supervised semantic fields for robotic memory

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.751438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.908510Z digest=sha256:15f1bb83372b377660e5f55462f5e0d29600106d7557b116fd41b1879802bc84

Observation 2ba6bcdc-5af8-442d-adf5-0b97d7c0c9fc · outbound

This paper cites Super-convergence: Very fast training of neural networks using large learning rates.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Super-convergence: Very fast training of neural networks using large learning rates

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.736995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.912701Z digest=sha256:1346712ce28628fda0d53d72af8cec0554c5df3e5a6a355aaf2008971d7c757a

Observation 9053103e-006e-4379-8a24-b4af73fafbb2 · outbound

This paper cites Denois- ing diffusion implicit models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Denois- ing diffusion implicit models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.916736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.916736Z digest=sha256:c5ff4ef63cadb763eac05a4b51ed4de1d3d5afd6193fc43afcedc01112652f2d

Observation 6e7b02f5-b163-4c87-a68f-c672a457b773 · outbound

This paper cites Score-based generative modeling through stochastic differential equa- tions.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Score-based generative modeling through stochastic differential equa- tions

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.712154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.920752Z digest=sha256:9f1e3a479808d6c3d96e80087339c77a555159fe8dfb41116b4f7f52e44c6882

Observation dc84ceaa-8ca9-4668-ae2d-b27557d85d5b · outbound

This paper cites Open- mask3d: open-vocabulary 3d instance segmentation.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Open- mask3d: open-vocabulary 3d instance segmentation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.697981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.924759Z digest=sha256:adbe279281472288f1718a1ea5fb5682a2b94e622a6de819bbe70552ab17e1a1

Observation 90173931-af52-49d2-824e-73679f5b0cc6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Gemini: A Family of Highly Capable Multimodal Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.928740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.928740Z digest=sha256:f50a05803f897965c0c943e50e511f9f37cf31e20d24277a0ad11e789893311d

Observation af2fa7ac-9477-476a-a3d5-343a25669593 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.933150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.933150Z digest=sha256:c402d6b00b29ad21d283f0b56b9a0f84e73af0ec3a2bae4d375444524bd2366c

Observation 3713a9ab-650f-44a5-9ac9-244e55a82ef9 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation LLaMA: Open and Efficient Foundation Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.937675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.937675Z digest=sha256:aca23227b2707aabb67c15f80ee525f67014d659ff422ec2d0fffa8db3cc62c1

Observation f4414da6-319f-434a-9e94-603cb6c7a9c9 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.942007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.942007Z digest=sha256:67d390e2d4e44391e3f0e7cf3260e3e18c9be329910e7345b9211cecf3716959

Observation 188efec1-78d5-45b2-b1e5-ce0fcf114127 · outbound

This paper cites Rio: 3d object instance re- localization in changing indoor environments.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Rio: 3d object instance re- localization in changing indoor environments

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.684118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.946399Z digest=sha256:bd1b016faa9366493b11357c40e4c96dce7bf1819da86e281560f00bca7fee34

Observation 6a5fa897-59a9-4aa5-9e79-71927bfd4a50 · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learn- ing framework.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learn- ing framework

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.670541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.950455Z digest=sha256:9ebff2445fb37e690339056475465cf565b22e6c4aa63977689e2653f4119700

Observation 36b75abf-d93b-40da-a930-dae7958998e8 · outbound

This paper cites EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI, 2023.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI, 2023

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.655917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.954462Z digest=sha256:69d6702045e51d9f093c2a5d9f22c1b48b2c7f49af2b5a585b90cf68bb98f8e3

Observation 1e765161-49db-4cfa-b061-16ce0c0d2cd3 · outbound

This paper cites Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.958546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.958546Z digest=sha256:d5d3d217f47db98aa46f044c5770943f5579e999a87bcc8ae380628117c7bf8d

Observation 6d02b364-431b-4fcb-9a34-24d44812b2f1 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Groupvit: Semantic segmentation emerges from text supervision

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.641360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.962797Z digest=sha256:d2889491515c938c02765a0a4cf56fe2c5f3e799c70c356943f7c0771c78e7d1

Observation 8efbfbf8-8e95-46fb-9a5c-67288c2bc6d8 · outbound

This paper cites Qwen2 Technical Report.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Qwen2 Technical Report

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.966702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.966702Z digest=sha256:e0a293398cefad0d1e7da9fa4de295600fe50308fb5c0b0593479707f8e5be8b

Observation 9b00930e-bdc8-4255-a703-030886c09744 · outbound

This paper cites Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Regionplc: Regional point-language contrastive learning for open-world 3d scene understanding

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.627270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.971103Z digest=sha256:abb67c9f172d8bd6522421096814447713749f9050330c535655857660cf20b8

Observation 3b50bde3-04dd-4cc0-8a10-3c8987d93792 · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d in- door scenes.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Scannet++: A high-fidelity dataset of 3d in- door scenes

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.612383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.975066Z digest=sha256:4a7c48c59299b78de3114e7144845bcc2758f0907b630b3046ed6978f3ae55c6

Observation 5879e755-4333-4f1e-968f-6e636d3a7a22 · outbound

This paper cites Sai3d: Segment any instance in 3d scenes.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation Sai3d: Segment any instance in 3d scenes

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:52:15.598253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T11:52:14.978847Z digest=sha256:d90bdfd11ad7f74fce43958dec1bbde0bf16d92580c6159ca82b8bc886346c90

Pith citing papers

No inbound Pith citation observations are available.