Pith. sign in

Paper Citation Record · LEDGER

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation

As of 13 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2411.13243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13243 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:46:10.072459Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf1a2954-5bf4-4875-b748-477ce9e0d64c · outbound

This paper cites 3d semantic parsing of large-scale indoor spaces.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation 3d semantic parsing of large-scale indoor spaces

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.835898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.835898Z digest=sha256:05e300bd17c07aaf37dfadd7d2a2765361827742f35642c5d79d733824b77f74

Observation 78a22aa5-8226-45e1-8b61-3f5f734538fb · outbound

This paper cites Emerging properties in self-supervised vision transformers.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Emerging properties in self-supervised vision transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.841512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.841512Z digest=sha256:9ca5d710bc3fca2b626818036ed4e1904f6fde3dd74de44fc555d2e2f18069e7

Observation 7f36bd1a-307a-4595-8e54-35d1243464f8 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Masked-attention mask transformer for universal image segmentation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.786369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.846666Z digest=sha256:e485f55815b10983fd5928d97dccff7bd110ab05854f95886c2992d8cbe37ee1

Observation 558d4a84-abc6-47f0-b362-5249f61db796 · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Per-pixel classification is not all you need for semantic segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.769872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.851979Z digest=sha256:590287544cb7d0625f5c10db975e6232727fb9bece7c8818cb617972cb41cc4d

Observation 857ed3f8-de2b-469a-b7f1-25dbbf8bdc2b · outbound

This paper cites Transductive zero-shot learning for 3d point cloud classification.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Transductive zero-shot learning for 3d point cloud classification

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.754268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.857085Z digest=sha256:ee59657931cc6bcd80190509a544a8ae97b26c678323ecb4185f6331556d878f

Observation 11d8c0be-a93f-452b-acb4-bfe4445bf919 · outbound

This paper cites CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.861897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.861897Z digest=sha256:e7002b0ab7185021ad481b86fcc920b678709c21e3eca3ad7b0a94f7cff1273f

Observation 721146ed-8144-491d-9086-aaa2a1988be3 · outbound

This paper cites 4d spatio-temporal convnets: Minkowski convolutional neural networks.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation 4d spatio-temporal convnets: Minkowski convolutional neural networks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.738908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.867561Z digest=sha256:cb97cbc5244e51d1c0902e4e12e206d283beb27552b2f5db1feea1020c9f9c29

Observation 83eb9f6b-6e09-4c81-a1ba-d9d7c3f3e2b1 · outbound

This paper cites Spconv: Spatially sparse convolution library.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Spconv: Spatially sparse convolution library

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.871855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.871855Z digest=sha256:a0d0d8f1df93352d90c317e524387a9036117ae621d0872672c27bae2c39ec1d

Observation b822299a-bca6-46ac-ba91-9d1bbdf56dd3 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.876484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.876484Z digest=sha256:c326daf2532df22f1e3ffa1209ff8a4b8a98cdaa0346976b8390861b111ca6cf

Observation efd47bee-b905-467c-9dc8-1e812a919cf2 · outbound

This paper cites Decoupling zero-shot semantic segmen- tation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Decoupling zero-shot semantic segmen- tation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.704441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.881197Z digest=sha256:bcf0601b2b1716fa2b98abd8064c0837941c1cf1252b5b29509d474f2c7d60b1

Observation dcf081a5-4de0-44a2-b913-ffb064558750 · outbound

This paper cites Pla: Language-driven open-vocabulary 3d scene understanding.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Pla: Language-driven open-vocabulary 3d scene understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.885601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.885601Z digest=sha256:e11d2b67e54c228909578007ae20e0a34471c4ee1288e936c55ea11f8f169040

Observation ea3da3bf-6a19-47d2-b36d-189d3cc877cb · outbound

This paper cites Open-vocabulary universal image segmentation with maskclip.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Open-vocabulary universal image segmentation with maskclip

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.678562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.890230Z digest=sha256:431a0289ba2658c37c3b1749e6b254a7f2c648fcd1fdb3c97041793a341590f5

Observation d060d7ef-3829-4316-9136-70ade5a9db14 · outbound

This paper cites Scaling open-vocabulary image segmen- tation with image-level labels.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Scaling open-vocabulary image segmen- tation with image-level labels

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.662409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.894688Z digest=sha256:07e8d7c4111a489e273666c3dd9f9aae9a130d3fde952dce73db768d605eeeb8

Observation e5f36373-22e4-4d0a-8927-2f1f0ca5b3ce · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.898818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.898818Z digest=sha256:45186626ad9708d30f63b1803addaea56bfdc74a0e1d14c5f09a7ec96b9c4568

Observation 9a355519-75d1-4bb9-9e98-f979e2d83829 · outbound

This paper cites Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.903856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.903856Z digest=sha256:c7e246ecf8e9bbd8dd3094d02eb03655fd627e9f45b9012e78a2abe096b92aa0

Observation 1ad8db9a-f275-43a1-af8e-67e9ed00dcbc · outbound

This paper cites UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.908857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.908857Z digest=sha256:a33352e9f3aaee5fb70c954d1d8b5bb2833baf0821471e4f25f7606f98172437

Observation 0f93fb09-3cd1-4cad-958b-cfe34be54ea7 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.913908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.913908Z digest=sha256:a7ae30a1a290300c1bcabd2cf5d76caf612e9e460b72d6a11243e74b0985aada

Observation 80fb5555-0805-41b3-b501-bfb86c64254c · outbound

This paper cites Denoising diffusion probabilistic models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Denoising diffusion probabilistic models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.647307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.919063Z digest=sha256:5149a22a3596f407fb6a1f0c1743e751a0fbee1d63dc81830998c7885e1d954f

Observation dab2d84f-8365-4fc7-a82a-dbc632566b50 · outbound

This paper cites Clip2point: Transfer clip to point cloud classification with image-depth pre-training.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Clip2point: Transfer clip to point cloud classification with image-depth pre-training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.923653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.923653Z digest=sha256:5b66447d725b0c44f3424f869e73ba1787516c5965ec78b505222f455af79dc5

Observation 62cfd89d-462a-499a-b6f1-6203bb873b54 · outbound

This paper cites Open-vocabulary 3d semantic segmentation with foundation models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Open-vocabulary 3d semantic segmentation with foundation models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.622689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.928234Z digest=sha256:7a70b250f8de390e4f562df6faf71aab5ea639476bd6510c943ce33c01abc619

Observation 609eb4ee-0363-47ab-915a-ce8ad7489197 · outbound

This paper cites Diffusion Models for Open-Vocabulary Segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Diffusion Models for Open-Vocabulary Segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.932975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.932975Z digest=sha256:255829c7398d8d5ed3d975cf11bdc2b57095439a95049a774fd7f29976d8a096

Observation ce77b63d-a1f3-4478-8f2a-52061a56d790 · outbound

This paper cites Segment anything.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Segment anything

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.937672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.937672Z digest=sha256:2e58a5bbd0c5c707d1e5760334e1e23c5d61432b87c3d3a55a674f5122134804

Observation 39e53374-2591-465a-91eb-1659c74ed2d4 · outbound

This paper cites Language-driven Semantic Segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Language-driven Semantic Segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.941912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.941912Z digest=sha256:6262bba1c7f82967bcd44cdafe16040115dae800adaa57cfdfe222b840628ede

Observation 52d6224d-8473-404d-976d-a8d9d99b8466 · outbound

This paper cites TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic Segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.946613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.946613Z digest=sha256:b3724c2ff4e7de95bee75edc1124c9975bf18153f24eee3b3ddc0d369d30decd

Observation fc44eab2-f161-4e25-aa9f-037048dc1e9d · outbound

This paper cites Guiding text-to-image diffusion model towards grounded generation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Guiding text-to-image diffusion model towards grounded generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.598419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.951597Z digest=sha256:3a4e186cf91bc1ed3fb23e8d3c624b2b1e20a0ea6b33b2075d2b9b9fb589e12c

Observation 64846e98-c1f8-4e76-b350-034dbd85ba83 · outbound

This paper cites Magic3d: High-resolution text-to-3d content creation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Magic3d: High-resolution text-to-3d content creation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.582927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.955911Z digest=sha256:0bd8f65eefc3914c1aecf9bce067570325c61839f38562acba528ad52ab0546b

Observation 19c529cf-e808-47ab-b667-9314e3f8ea13 · outbound

This paper cites Weakly Supervised 3D Open-vocabulary Segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Weakly Supervised 3D Open-vocabulary Segmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.960220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.960220Z digest=sha256:585f00b77560543de027a4024e1066406c5aec65df2d4cde3672015c789c476c

Observation 7077cfea-7048-4b40-b00e-047fd27ca418 · outbound

This paper cites Segment any point cloud sequences by distilling vision foundation models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Segment any point cloud sequences by distilling vision foundation models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.568087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.964909Z digest=sha256:dbb80f7cc82cbc6d3016bb46599ce890ea43d5d5df3504539d070d238af10b0e

Observation 27448eb0-aeb7-43a9-93c4-9fbc36512e56 · outbound

This paper cites Decoupled Weight Decay Regularization.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.969515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.969515Z digest=sha256:26bedeca667dd173277db9984df9f3289faae2cc8b64e919deb3d11ff391f1c2

Observation 395aba57-d766-42d7-873b-cb4d2912e21b · outbound

This paper cites Open-vocabulary Semantic Segmentation with Frozen Vision-Language Models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Open-vocabulary Semantic Segmentation with Frozen Vision-Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:46:10.182714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.974385Z digest=sha256:a77e681388f9b5dd846edddc1446d9a6770994e60df500fe41b632c9d6953a52

Observation eac43e42-c907-41be-928d-0f8f12efa93e · outbound

This paper cites Generative zero-shot learning for semantic segmentation of 3d point clouds.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Generative zero-shot learning for semantic segmentation of 3d point clouds

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.553002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.978968Z digest=sha256:ecdf7e7d7ab3050d110a50355cb97277dba08a7ac533f09f3a71370d5c4052e3

Observation 5dd0dbd7-89bc-4da2-bd26-15f2e8b8e4bb · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.983560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.983560Z digest=sha256:1748522c490380d8cffe5c1135078ec33545ae449aa6e7fd7ea8c3fe0c4b279d

Observation 7ba131b5-fb58-4b05-8bdc-28e47a8432b9 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation DINOv2: Learning Robust Visual Features without Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.987758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.987758Z digest=sha256:a7a63cba34995191d398702386b0126396446d38a3dd47ab353e00e0e6b47187

Observation 0b2204d1-82d8-403f-9f6d-1d93b0da5614 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Openscene: 3d scene understanding with open vocabularies

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:09.993708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:09.993708Z digest=sha256:56e1546a502c9e67d431d55b288ad35f5556a2d3bec3ce2813875e45d717f13e

Observation eff91478-3e75-408c-bdb3-659c1841506e · outbound

This paper cites Dreamfusion: Text-to-3d using 2d diffusion.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Dreamfusion: Text-to-3d using 2d diffusion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.519250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:09.998127Z digest=sha256:326f13aaff7672d45262c593fbf09e7d4cf3c132fa58222ff37ce700a6edb100

Observation a5f85f7f-74aa-48fa-8f51-5fd6ca11687b · outbound

This paper cites Learning transferable visual models from natural language supervision.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Learning transferable visual models from natural language supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.003200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.003200Z digest=sha256:9f462c0ddb7b1914effe07b296fe5c5cda6d2d6dfcdea050f1a320ed470aef9a

Observation 93103c53-289e-4781-9f83-f5c2f706ebd1 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation High-resolution image synthesis with latent diffusion models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.007716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.007716Z digest=sha256:3219050e2a81f4aa26d11f4007c0d24cb58759872f42d5273811e0fbf8dd0dfc

Observation 5f773dd6-48f4-4386-964b-6e205659d548 · outbound

This paper cites Language-grounded indoor 3d semantic segmentation in the wild.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Language-grounded indoor 3d semantic segmentation in the wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.012269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.012269Z digest=sha256:b10ebd6071549f60205c87d9940175e2449e7531941ea847e994ec3ba0f27f04

Observation 5828134d-e411-40d1-97c9-7e6bb1c6729e · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Photorealistic text-to-image diffusion models with deep language understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.474877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:10.017134Z digest=sha256:f687eed74275bd023888bf171f65db2298f5ee12c5630ec0c3734a152324b70a

Observation c446b89e-62ed-4580-aed1-8e2b4e6153ed · outbound

This paper cites Dreamgaussian: Generative gaussian splatting for efficient 3d content creation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Dreamgaussian: Generative gaussian splatting for efficient 3d content creation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.459203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:10.022245Z digest=sha256:72942f7dc6292831f32abd7b280ec787a27e26d66368f16d59533452de8016d2

Observation 0a64c565-b35d-4415-8c8b-44523533772f · outbound

This paper cites Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.443737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:10.026793Z digest=sha256:8e6f9a9dddfaf235a788d28b6f138286b6e6d2733ef6d1af2bc36484b5787a7b

Observation 1a8875c1-c52a-43e6-ab82-b6701ee2bbdf · outbound

This paper cites Point Transformer V3: Simpler, Faster, Stronger.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Point Transformer V3: Simpler, Faster, Stronger

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.031420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.031420Z digest=sha256:f68b52b812909f61ac03aa3d946c6a3c684f5ce66081a444b19bbdeaba6d4cbb

Observation 3d8920c5-7953-46fd-9a1e-21730a93cf52 · outbound

This paper cites Point transformer v2: Grouped vector attention and partition-based pooling.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Point transformer v2: Grouped vector attention and partition-based pooling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.428777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:10.036429Z digest=sha256:e862b5efbc9db9d0bd2ed713a113386a66440b4e848f8c83162a2ef6f2352121

Observation ef8baa7b-3fbc-47c3-9165-7d2b2eee86a2 · outbound

This paper cites 3D Open-Vocabulary Panoptic Segmentation with 2D-3D Vision-Language Distillation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation 3D Open-Vocabulary Panoptic Segmentation with 2D-3D Vision-Language Distillation

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:46:10.131526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:10.041152Z digest=sha256:9e587e3cdf9878247cb2974a6f8462968a89e9165bc193793a4461c33fea92dc

Observation dac75b9e-69b9-4853-92eb-f52db788d026 · outbound

This paper cites Open-vocabulary panoptic segmentation with text-to-image diffusion models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Open-vocabulary panoptic segmentation with text-to-image diffusion models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.045841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.045841Z digest=sha256:98bcc7d8ea462c86499e82b8a3dc14de964eb575bec30b596d606a04eeff6b5b

Observation 03402bea-bfa2-402e-ae9a-2c662e1c0345 · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Side adapter network for open-vocabulary semantic segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.403769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:10.050407Z digest=sha256:db37f98ee5b43ba8549af10e22084ccf81dee7750d29cccb52237146f73a48e7

Observation 689b8e2e-4542-48bd-8dce-93e5a4082c8c · outbound

This paper cites RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.054878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.054878Z digest=sha256:c9501c433428b1f70dee1b6e8a8fbb9714e3c9059b20c0ad6b564ac205494917

Observation b7a64b4e-a6dd-4b62-ad84-357088974853 · outbound

This paper cites Vit-gpt2 image captioning.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Vit-gpt2 image captioning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.388705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:10.059407Z digest=sha256:587160dddee64318f5c7339a5e88cf13edb588eeb45558225fc97cea6b3f198c

Observation ab28d98a-0014-4cbd-9097-9372a76f5ca7 · outbound

This paper cites Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.373333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:10.063707Z digest=sha256:c2e62b899f2acb2cc5b5f60d530bbeb984e42053ade4a95256db5aeb406469b5

Observation c7d84e53-af3b-4033-a054-292f1cff11ed · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Adding conditional control to text-to-image diffusion models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T16:46:10.067942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:46:10.067942Z digest=sha256:0ffb706ee2c0d86afc4f5ab0fec46292008d3cbc0af1c26518182577f7b09898

Observation 247efaf2-45f9-4e37-913c-a97bce88aaf8 · outbound

This paper cites Point transformer.

XMask3D: Cross-modal Mask Reasoning for Open Vocabulary 3D Semantic Segmentation Point transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:46:10.348859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T16:46:10.072459Z digest=sha256:784ad6fb761b26fb0d593cb3cbaa90d9d7deded8d18f7f165f2e06ca581cbee8

Pith citing papers

No inbound Pith citation observations are available.