Pith. sign in

Paper Citation Record · LEDGER

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding

As of 21 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 1 inbound Pith citation observation for arXiv:2505.18819.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18819 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:07.374604Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:34:14.621355Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T17:34:15.117487Z

Reference resolution

95 of 95 outbound references displayed

  • verified exact3
  • verified fuzzy62
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f601a621-b8f1-436c-b778-145e8e648eba · outbound

This paper cites Segment anything.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Segment anything

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:57.493569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:57.493569Z digest=sha256:fceec7acf7e8ffc9fb741b48b3686b873138f643655cedbea06f9cb30ca040bc

Observation 3091c7cb-beb7-4184-8dc0-f1ab3ac97057 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding SAM 2: Segment Anything in Images and Videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:57.544665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:57.544665Z digest=sha256:2c25c4e1f55eeaee5399c22b1779e2e04d8d18bf34d56c7e653a21d6b2592305

Observation 0da3052e-4c77-4225-8593-6419a2bfca14 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Learning transferable visual models from natural language supervision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:57.624160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:57.624160Z digest=sha256:2057f50a6974a4d688f0d88c10de55cb3684191ccfc3b7a2ab9a5961ae081fce

Observation 542f138f-57ea-4b7f-bccc-dea0c7b817b8 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Emerging properties in self-supervised vision transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:57.687059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:57.687059Z digest=sha256:34f25172781f1314d9de2f5a8b196cc7b3687c415a3a66317560e67bcb3d1548

Observation 16647114-b818-4f5b-9545-95bf2881885a · outbound

This paper cites an unresolved cited work.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:57.740016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:57.740016Z digest=sha256:359ee88de47016954ea82eafe5c144cdce32fb6ef68a99215f35463764a91ecc

Observation 49b504cb-8c35-4970-9d64-fe4a41d23991 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:57.803925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:57.803925Z digest=sha256:a15f35fec31e735334ed45ebb100733f85caf7d637a146163fbb7ebb971d2b80

Observation 2f61ef11-4b75-49f3-b783-4005e377f692 · outbound

This paper cites Perla: Perceptive 3d language assistant.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Perla: Perceptive 3d language assistant

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:57.861294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:57.861294Z digest=sha256:5185f89815abeba041150b0b82922b30a0dbfd80690dbb22bad9e32e1ceaa6fc

Observation 8f4994a1-a916-4fed-b541-e32d49f4b879 · outbound

This paper cites Shapesplat: A large-scale dataset of gaussian splats and their self-supervised pretraining.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Shapesplat: A large-scale dataset of gaussian splats and their self-supervised pretraining

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:57.915221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:57.915221Z digest=sha256:4c7a9ef97601b3d3e189235c1e9759ef7143c26c18e3d82555a8cd4387dd911d

Observation f79611db-3256-4c8e-b75d-bf0bdc7c4aba · outbound

This paper cites SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:57.987125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:57.987125Z digest=sha256:be2b4480ff44470cceb6dc35d85a66f03c965d8dd4f2d101c7bdc5c8d72e9da5

Observation 760f88c6-8039-457b-a77e-253ec5fcd8d6 · outbound

This paper cites Unsupervised deep probabilistic approach for partial point cloud registration.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Unsupervised deep probabilistic approach for partial point cloud registration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:58.067391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:58.067391Z digest=sha256:a61268b938659268b392f7726dcf5b3a9a8bf5de6c27aa6249c7f99fa2167814

Observation 1b5031e9-1e10-4b84-9a7c-f7452697f650 · outbound

This paper cites Pvafn: Point-voxel attention fusion network with multi-pooling enhancing for 3d object detection.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pvafn: Point-voxel attention fusion network with multi-pooling enhancing for 3d object detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:58.140788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:58.140788Z digest=sha256:30f2dda24a717303513dc75533f970c8964932e6e3befa38ee7e483b00ff643e

Observation ac7e62f0-7dc8-4b6b-b1fd-1435e367d0dc · outbound

This paper cites ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:58.227222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:58.227222Z digest=sha256:4045abacb82f09c3cdfb58dd59bb8b07ee63f4b8f7289d7e61ea9cf0f310adad

Observation 29e71aa9-3b02-428e-afbb-1958cf4ca6ec · outbound

This paper cites Attention is all you need.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Attention is all you need

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:58.265388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:58.265388Z digest=sha256:d680a00633d85f004755aaafebbb11f9991c0502a46fc9bb1312015c29be1409

Observation 34842a55-e66c-4bd2-a0d3-a518b3b9b48b · outbound

This paper cites Masked jigsaw puzzle: A versatile position embedding for vision transformers.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Masked jigsaw puzzle: A versatile position embedding for vision transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:58.336861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:58.336861Z digest=sha256:8fb462a12c0415d85feea5227eb91bd953da3e5fbde2cde71ec48f7f94051f08

Observation 5c275d38-db8c-4d56-ba6f-e3b7facec82a · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding An image is worth 16x16 words: Transformers for image recognition at scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:58.413884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:58.413884Z digest=sha256:8c15c74c1187245cbe3a9141206a0d8bc6fd107a5eeef237bb8d97ba457a00dc

Observation ec305381-9e04-4924-9614-35cff6cee61d · outbound

This paper cites Point transformer.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Point transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:58.500441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:58.500441Z digest=sha256:4c2bdaab6914d9e24780580cc2387cfb1c068a31ec04dd817a310b8db7b201c9

Observation 343dde7d-6275-43c1-8597-56ab67415b1e · outbound

This paper cites Point transformer v2: Grouped vector attention and partition-based pooling.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Point transformer v2: Grouped vector attention and partition-based pooling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:58.561462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:58.561462Z digest=sha256:e5d1fb7531c4f5acf46e1aa76bfcd77728f723148eab34b4cfbb64a4f0b1d8b3

Observation 5f2cb2c8-d54c-4dc7-9897-9428332de9ee · outbound

This paper cites Point transformer v3: Simpler faster stronger.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Point transformer v3: Simpler faster stronger

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:58.643925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:58.643925Z digest=sha256:16e48aacec04bd3a4acec9f36604e4310d265ad30b1a9e97da9a63aa56e60edd

Observation ce3abd1b-e9fe-43dc-8bd4-014a095e7f27 · outbound

This paper cites Bringing masked autoencoders explicit contrastive properties for point cloud self-supervised learning.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Bringing masked autoencoders explicit contrastive properties for point cloud self-supervised learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:58.718171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:58.718171Z digest=sha256:2d14676d325dd62db031921ddd8477997f2d6daed498d7e98be5a8575e17245e

Observation 114a15ae-ef4d-4d31-8d01-848fdb2ed881 · outbound

This paper cites Point-DAE: Denoising Autoencoders for Self-supervised Point Cloud Learning.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Point-DAE: Denoising Autoencoders for Self-supervised Point Cloud Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:31:08.802223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:58.799285Z digest=sha256:afd08279a48d78d82e98c1c4782a0a55e18483838383c637f69346357eceacdf

Observation 4eb5dbf1-a71b-47c1-995d-ceed029b1531 · outbound

This paper cites Geomae: Masked geometric target prediction for self-supervised point cloud pre-training.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Geomae: Masked geometric target prediction for self-supervised point cloud pre-training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:21.307762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:58.895984Z digest=sha256:5f782760bacb7115eb61df8d5b39447d908efca5760d12e5d662dc7cebd6d845

Observation f1e2362a-dae5-4d4e-9f7f-e966245e8d7d · outbound

This paper cites Masked autoencoders for point cloud self-supervised learning.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Masked autoencoders for point cloud self-supervised learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:20.993693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:58.975517Z digest=sha256:4680f6075ab97ff8f5290a16f322c6fc4bbb240a01af870ee23b052b8c753048

Observation ce473609-e824-4e84-b030-d5c440340a28 · outbound

This paper cites Bootstrap your own latent: A new approach to self-supervised learning.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Bootstrap your own latent: A new approach to self-supervised learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:20.493998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:59.156559Z digest=sha256:086595f3a158514f8c9caf2f819dc39a3e774ebda117499df05f90e549df75fd

Observation c28429f7-220a-41e9-82f9-456d7b3ddeb7 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Momentum contrast for unsupervised visual representation learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:20.281665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:59.243829Z digest=sha256:e08a50c8868b29c967efc7be0f0e0d776265c0f5d4985877fc62522a194a1880

Observation d2e3f62f-6ad0-4b1f-b83d-9f7ef26212e9 · outbound

This paper cites Point-m2ae: multi-scale masked autoencoders for hierarchical point cloud pre-training.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Point-m2ae: multi-scale masked autoencoders for hierarchical point cloud pre-training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:20.022788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:59.317784Z digest=sha256:aae669dbd6f34e93a5ed11144d5873041266bb1c70641f2d0baaf4af13b0024b

Observation 18eb51ed-0036-4e49-97be-dcf66a15d4e5 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Masked autoencoders are scalable vision learners

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:19.862763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:59.425041Z digest=sha256:b547204121ce7dfe357404c938c9173ffad2144b24a817a224f1a3bc8b633cb3

Observation ea258da3-3ec5-44dc-818f-14ffdd3bb96b · outbound

This paper cites ShapeNet: An Information-Rich 3D Model Repository.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding ShapeNet: An Information-Rich 3D Model Repository

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:59.519899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:59.519899Z digest=sha256:bf1373f9909446b9b2fb55845b94a06f900cabe8284fe0167ea4bb7d2eba3aa3

Observation 6aebc049-a809-4157-b2d9-e86cb7ea6b10 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:19.734983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:59.565921Z digest=sha256:b027e47367705e1e782200a25cbbe2e94192fe089a3559dad9e0cfd25539f43a

Observation 4af60b0d-d21f-4050-9bd5-e64017e0901a · outbound

This paper cites Con- trast with reconstruct: Contrastive 3d representation learning guided by generative pretraining.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Con- trast with reconstruct: Contrastive 3d representation learning guided by generative pretraining

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:19.612751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:59.646335Z digest=sha256:2866278b5bf62ad3769518e74281dcbbadf61f8c42acf561c15481a28041ca6b

Observation 6dd51cc3-a9c2-4417-865f-2ce40df02a8c · outbound

This paper cites Frozen clip transformer is an efficient point cloud encoder.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Frozen clip transformer is an efficient point cloud encoder

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:19.444272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:59.721675Z digest=sha256:4b2f1d0c0ac106c51f1b59625cbd8515a309075fe39c9081d3367aecb667c172

Observation 58f66f49-378d-40a7-ac64-23db889454a1 · outbound

This paper cites Pimae: Point cloud and image interactive masked autoencoders for 3d object detection.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pimae: Point cloud and image interactive masked autoencoders for 3d object detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:19.312052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:30:59.839603Z digest=sha256:f37ea605ab7d12aaf36105bcc44d58a3b08f8fd2536f1cdf207fbc6b339163d7

Observation 0a10ceb8-d9ce-4f61-859a-aaad75086616 · outbound

This paper cites Pointmamba: A simple state space model for point cloud analysis.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pointmamba: A simple state space model for point cloud analysis

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:19.200186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:00.014468Z digest=sha256:7cb30d189f542dc68c2b93866e3920f12917ef7fc1e075db49f0ca399b69d5dd

Observation ca2fc097-a46a-4922-8d8e-ddacd6c9f7da · outbound

This paper cites Pix4point: image pretrained standard transformers for 3d point cloud understanding.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pix4point: image pretrained standard transformers for 3d point cloud understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:19.029943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:00.084134Z digest=sha256:5e1fb18ad5a7f561f93a19c85bdcd3b65b7980f08ee1900e8e3a4ff7b1922ef5

Observation 71e4f928-2f15-47ee-8e47-037bb4db8d74 · outbound

This paper cites Can We Solve 3D Vision Tasks Starting from A 2D Vision Transformer?.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Can We Solve 3D Vision Tasks Starting from A 2D Vision Transformer?

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:31:08.536346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:00.150624Z digest=sha256:f035e8915e36e2b3dc6bdc9dce718c1ab1a1ee61e3bf37556c7d0687d04c28c6

Observation 0498613c-fa76-4360-a121-b6b1fab7f1d5 · outbound

This paper cites Pointcontrast: Unsupervised pre-training for 3d point cloud understanding.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pointcontrast: Unsupervised pre-training for 3d point cloud understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:18.795470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:00.250447Z digest=sha256:9a7e763141f579dbdd0ce49e4199277b7457726751e054f0da97c1003301360e

Observation bf8d5040-b357-4467-95ca-227767b43892 · outbound

This paper cites Spatio-temporal self-supervised representation learning for 3d point clouds.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Spatio-temporal self-supervised representation learning for 3d point clouds

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:18.632037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:00.368322Z digest=sha256:bcb0f62aa99d30d4f054c16912fb71c97e01d9ee5920c09b6f7c347b29b63da3

Observation ef262909-64fe-41e5-81e8-08b00408ceb2 · outbound

This paper cites Global-local bidirectional reasoning for unsupervised representation learning of 3d point clouds.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Global-local bidirectional reasoning for unsupervised representation learning of 3d point clouds

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:18.480563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:00.493839Z digest=sha256:3df93a662ae55c43c7a5f12d523397439cb212681f57811d7ac8da3117e4a9a2

Observation 86b64999-99cb-4918-aa7b-15084c3f4f40 · outbound

This paper cites Point discrimi- native learning for data-efficient 3d point cloud analysis.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Point discrimi- native learning for data-efficient 3d point cloud analysis

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:18.256799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:00.582331Z digest=sha256:9af6ba4d5cd631cbceff30c6253209414586fa26293fdef7e0d0af11c85908c6

Observation aebf0e7b-c1c1-417a-a0a6-2199ab4e23de · outbound

This paper cites Data augmentation-free unsupervised learning for 3d point cloud understanding.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Data augmentation-free unsupervised learning for 3d point cloud understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:18.134556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:00.721849Z digest=sha256:bb85876fd588cff0dda02f00cfe19e9c4790a39548b8e5f8ad66b59d6a7bdd7d

Observation a2e0def2-a139-48ce-af44-39555e7f42c6 · outbound

This paper cites Unsupervised point cloud representation learning by clustering and neural rendering.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Unsupervised point cloud representation learning by clustering and neural rendering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:17.938629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:00.833646Z digest=sha256:e51e39f6e2566ec6f94ef9ad91dd945dfa98122b5f5ea7fa5046466d5a7cb2d4

Observation 444bd5b0-85d8-44d6-af2f-02be00988940 · outbound

This paper cites Pointclustering: Unsupervised point cloud pre-training using transformation invariance in clustering.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pointclustering: Unsupervised point cloud pre-training using transformation invariance in clustering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:17.718583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:00.925509Z digest=sha256:3f68252ccb8faaedbb4d8edf63a84493264e42832410c797e119ad9f7f1cde4e

Observation 3980f775-c2f9-4d3d-813e-29ee31b199c3 · outbound

This paper cites Gd-mae: generative decoder for mae pre-training on lidar point clouds.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Gd-mae: generative decoder for mae pre-training on lidar point clouds

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:17.431440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:01.018647Z digest=sha256:f56840464a0514495f93b75928c8bbb58c6900b504a69c8fbe1a27c580362e44

Observation 7977a6e6-e4c7-458a-ba7e-82b5e6072d14 · outbound

This paper cites Pointgpt: Auto- regressively generative pre-training from point clouds.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pointgpt: Auto- regressively generative pre-training from point clouds

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:17.207800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:01.156154Z digest=sha256:2da1aabf9dc4f174df8de308182726d9c9fee6123f12cf50acfd6ba4862530ed

Observation c93181fb-3bca-4156-9d84-8f11ef0e3520 · outbound

This paper cites Point cloud pre-training with diffusion models.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Point cloud pre-training with diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:16.976813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:01.275628Z digest=sha256:91ed937a576442cca1e6db7f7ab354f605910d384e596452da66aa508439dd32

Observation eac6e71e-5e4b-4223-ab54-fd26cf812987 · outbound

This paper cites Denoising diffusion probabilistic models.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Denoising diffusion probabilistic models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:16.671042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:01.410702Z digest=sha256:270365523a24a0ad49615344ccba57a1fd096394aba21d380a438ca33a0bfced

Observation 911764cd-a211-4ef1-a5f1-3cfdeadc89d8 · outbound

This paper cites Spatio-temporal graph diffusion for text-driven human motion generation.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Spatio-temporal graph diffusion for text-driven human motion generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:16.427327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:01.537635Z digest=sha256:7565c6d8adecaf0ea4b45210ee8b53d7d683447954068559c15c9748c478a47d

Observation f2de19b8-ae44-4d8e-a1d3-dbc559960e65 · outbound

This paper cites Denoising diffusion probabilistic models for action-conditioned 3d motion generation.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Denoising diffusion probabilistic models for action-conditioned 3d motion generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:16.139201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:01.693831Z digest=sha256:160331ec46df5f23b161a4738ca9920d82fe34dca0df936c6d245fff3192f4e7

Observation ffc74efe-c326-4579-9189-d59f841077e5 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Efficiently modeling long sequences with structured state spaces

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:15.940840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:01.811312Z digest=sha256:9da80a2b00320b7ffb2d520f6c75d86dfe831a7336ccbe14c0f1b541b5e81b0b

Observation 61625a85-baa3-4226-bdbe-02a0a0a9849c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding LoRA: Low-Rank Adaptation of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:01.915982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:01.915982Z digest=sha256:a4bd2f40ace23abada4ab2e61029bff845ad4091ebd4e8602da2d42e5b7afcef

Observation 9e9c116f-f003-4c08-8360-566912aeb108 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Parameter-efficient transfer learning for nlp

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:15.751618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:02.025816Z digest=sha256:b0c5dd137912d9310501ce32ac119890c8e5005b085eaae559c5f604494e549b

Observation 05e90377-a106-4812-9f7b-a64acc369604 · outbound

This paper cites Adaptformer: Adapting vision transformers for scalable visual recognition.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Adaptformer: Adapting vision transformers for scalable visual recognition

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:15.573985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:02.195733Z digest=sha256:d4aa0bdcc73208a1f3592e0cdec542157c4b2bf7e309db066a99a50beee27a94

Observation 5ed91843-9620-4b79-b3ee-540ebacddac3 · outbound

This paper cites Visual prompt tuning.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Visual prompt tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:02.350655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:02.350655Z digest=sha256:e8cb5b52034442c71af3483cef857e65d38776654e819517b64419ab8a9f63eb

Observation b0e658ee-898a-427e-9772-cbf0b22680da · outbound

This paper cites Exploring sparse visual prompt for domain adaptive dense prediction.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Exploring sparse visual prompt for domain adaptive dense prediction

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:15.280623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:02.452453Z digest=sha256:171267936d00bc78df64d0c78490bf55500b25689b581f3ebac4ed6771e6e75d

Observation 1baac4a4-0667-4996-be69-6955ba6213c8 · outbound

This paper cites Point-peft: Parameter-efficient fine-tuning for 3d pre-trained models.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Point-peft: Parameter-efficient fine-tuning for 3d pre-trained models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:15.111658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:02.633769Z digest=sha256:9508eeffdc4fdb69de1a7172481ff436c7740d11ca903667a36c3dcb579c07cb

Observation 68f3f28d-e32f-4823-9c75-1f391001e74a · outbound

This paper cites Gaprompt: Geometry-aware point cloud prompt for 3d vision model.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Gaprompt: Geometry-aware point cloud prompt for 3d vision model

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:14.988583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:02.774666Z digest=sha256:fd5917a428649ae667bb7a3f1f91ba00bdfaeee1140ef13f8fa246291fae1d5a

Observation 9fdb3093-5059-4acf-adb3-4568ac43c35c · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:02.918450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:02.918450Z digest=sha256:bc501f343383b317b9218e706d1959e396ea917566f604d6c2cb5ccee05e560f

Observation 789bd28e-c386-4c5d-ae78-3aeea2d6c50d · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:03.077802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:03.077802Z digest=sha256:f41578efabd8b9c59d79a7c1ece63e75fee51074ce0d77a62eb8a4b45a0ed653

Observation a0730824-d762-4294-a85b-1679243c5df9 · outbound

This paper cites Language models are unsupervised multitask learners.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Language models are unsupervised multitask learners

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:03.206959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:03.206959Z digest=sha256:7dd7078b05520f4ea5b52574d583ec45d1f17fbb137485a13b2c00a72523ae88

Observation cfdfcd6e-3106-4a00-b803-273515f27447 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Training data-efficient image transformers & distillation through attention

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:14.821934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:03.366466Z digest=sha256:c91e2dd8402cd90de26a205ca76f03cce6b03db62e5e8a50952c1a595ed21952

Observation a0a6ad1d-e84c-4249-8d2c-30989aefec82 · outbound

This paper cites AST: Audio Spectrogram Transformer.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding AST: Audio Spectrogram Transformer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:03.520260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:03.520260Z digest=sha256:a8ba75a0fcbda7124add623e49a342a0d8a2eb545e7380ec3953fd9c37b08f3c

Observation 7cea26ac-4cf3-477a-9764-7611f6df012b · outbound

This paper cites Ssast: Self-supervised audio spectrogram transformer.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Ssast: Self-supervised audio spectrogram transformer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:14.674060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:03.693827Z digest=sha256:e848fad7ed03c3ee66a29e8b738a26cdfc0824717817b169f02954fd74dff35e

Observation 87198369-ad0e-4b68-9f9b-8dfc69140e0a · outbound

This paper cites Imagebind: One embedding space to bind them all.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Imagebind: One embedding space to bind them all

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:14.528061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:03.840165Z digest=sha256:9ef5f5b5a60f1876824253f31b12de4891b1a6903961587ba37b063beeea2b09

Observation 6e690621-f026-44ed-a64d-e3820e529b0c · outbound

This paper cites Pointclip: Point cloud understanding by clip.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pointclip: Point cloud understanding by clip

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:14.414825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:03.988866Z digest=sha256:9e5c0593534e1fffa32e3d2472f544717e63e2dbc59ab3c83561ac0b42e65974

Observation d4f75a62-210c-402f-a5ea-b8c6681a3d01 · outbound

This paper cites Pointclip v2: Prompting clip and gpt for powerful 3d open-world learning.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pointclip v2: Prompting clip and gpt for powerful 3d open-world learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:14.314200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:04.159282Z digest=sha256:fbc12a3df6988a1f792b072af3fb4ae10e8a38e71ecfd9b1f108145779cf189c

Observation 94e6ee63-7215-4850-bba1-4f3aea878fc0 · outbound

This paper cites P2p: Tuning pre-trained image models for point cloud analysis with point-to-pixel prompting.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding P2p: Tuning pre-trained image models for point cloud analysis with point-to-pixel prompting

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:14.168323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:04.274153Z digest=sha256:9636d3e873c53df93a2fcd8babbf040a7d8afaa9c4f93eb29ae8f79df417890f

Observation 2ea0b34d-2abd-473e-8828-b8a791275990 · outbound

This paper cites Image2point: 3d point-cloud understand- ing with 2d image pretrained models.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Image2point: 3d point-cloud understand- ing with 2d image pretrained models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:14.060148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:04.419084Z digest=sha256:5fbe780e637d0fa786204b27354d4c46dd324e4197640603035c9ce95c40508d

Observation 75a276ea-305f-434d-bb40-67d08cba74c7 · outbound

This paper cites Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:04.494724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:04.494724Z digest=sha256:ce1c70d5d27626c792ef05909470fd885bda9030b86ec334a04f4dad4e48fa09

Observation e9ee25db-96c0-4245-b971-7038ee1c698e · outbound

This paper cites Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Ulip: Learning a unified representation of language, images, and point clouds for 3d understanding

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:13.955946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:04.615616Z digest=sha256:9aa37b9e83889ac415086b46b06eb2f71abde2809cfde2ab2b61cd9f9aec4e8d

Observation a9748a4e-b614-42ea-a503-5354ddc573fc · outbound

This paper cites Point Cloud GAN.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Point Cloud GAN

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:04.711501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:04.711501Z digest=sha256:615a8894819e87a2812cce2d298df1b725aaa7b41350179bcf0b6b74a2e9dc26

Observation 14d1b597-b42e-4c36-914c-b579972c597f · outbound

This paper cites Large-scale point cloud semantic segmentation with superpoint graphs.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Large-scale point cloud semantic segmentation with superpoint graphs

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:13.858926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:04.815158Z digest=sha256:06043966fa0c5b05bb1fe8255b004b8394b7da25e502995adb44ca5094e96c72

Observation 8ac327fa-b1bb-44f3-9f0f-cdec396b5fc0 · outbound

This paper cites Open3D: A Modern Library for 3D Data Processing.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Open3D: A Modern Library for 3D Data Processing

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:04.896833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:04.896833Z digest=sha256:110f6f92dfef49b77039847341f98b59bd2da8817d5dbd1806a1bed1658fa80d

Observation 7d394341-382f-44fa-8a90-8bee41170ccf · outbound

This paper cites Pointnext: Revisiting pointnet++ with improved training and scaling strategies.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pointnext: Revisiting pointnet++ with improved training and scaling strategies

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:13.752251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.006880Z digest=sha256:8ac9d7954fc06b876158b5539269643c26b9222ff6fdb49f329e0f23fc513568

Observation 99825e7d-ad8f-452c-8fad-ea6bd7389744 · outbound

This paper cites Investigating self-supervised methods for label-efficient learning.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Investigating self-supervised methods for label-efficient learning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:13.611873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.084772Z digest=sha256:733ec4d978b9fba36750c011f085d3ccdd90ab199efc9cd9ec8196786c373797

Observation 9e4f9fd3-2583-46af-8722-16a71c0d5c23 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Sinkhorn distances: Lightspeed computation of optimal transport

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:13.491294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.189212Z digest=sha256:3f2bcf033cb207e4e9a8a6a426c58fb64a7734a2902d4ea828f9e64b3083c211

Observation 28138aac-6cc3-4975-814e-59fabdbbc063 · outbound

This paper cites Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:13.315313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.290663Z digest=sha256:4a2c93e4ce469e1e8ee96dbc1c3b4d2051cab297599d35cef63d21dfd3339fc3

Observation 7faa0cba-343f-49d5-a972-ab9d242d10fc · outbound

This paper cites Pointclip v2: Adapting clip for powerful 3d open-world learning.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Pointclip v2: Adapting clip for powerful 3d open-world learning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:13.160276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.352919Z digest=sha256:51bf7ad1b7225a79b6bd3cd1e009314af9b531a2a76e01895571ff07b8ea8a8f

Observation e4a377a2-62d1-4f50-9655-6c70395cd150 · outbound

This paper cites Adamw and super-convergence is now the fastest way to train neural nets.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Adamw and super-convergence is now the fastest way to train neural nets

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:12.893549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.431811Z digest=sha256:54179b83aef7a6cbc5b0ed5603ecee97edf5e47739cedf26e9e03732f8e8cd7a

Observation bfe268e0-70b4-4f9d-b6d4-09c47cb52d63 · outbound

This paper cites Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:12.709788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.500714Z digest=sha256:715077b7a4391b3413c9c9b1ec77c3d4eec122af6c4aa42548c69d2d56933a37

Observation 01e67b11-7a48-445a-a1a5-ce1cc6aa8d6b · outbound

This paper cites Partdistill: 3d shape part segmentation by vision-language model distillation.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Partdistill: 3d shape part segmentation by vision-language model distillation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:12.481481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.625224Z digest=sha256:97dad924fb3f0d33de0d0ef8b6a783da5d901932b9c15d706f265dbaac767e35

Observation a810e214-03c8-4547-905e-8487d685716d · outbound

This paper cites Maskclip: Masked self-distillation advances contrastive language-image pretraining.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Maskclip: Masked self-distillation advances contrastive language-image pretraining

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:12.208669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.735561Z digest=sha256:814803936d862cceb9275e496bc23fa49d3a5c52faa825b6582d2a6ea69f6b1c

Observation 53680c93-4084-4a8f-8f03-b729c09e8532 · outbound

This paper cites Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:11.976053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.815290Z digest=sha256:a6d247fca5874e73a514964d2c28f60f0916c841b76d5aef0118cd3770cd5c29

Observation c8142e72-550d-4625-a425-b0da7fac18d7 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Openscene: 3d scene understanding with open vocabularies

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:11.758045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:05.922596Z digest=sha256:a6b78fa4009e91309dee1b7f283ffe8cc645a197255f8c9b2898191b8815db02

Observation 0faf4c34-4914-40a5-9204-4430f0d37b13 · outbound

This paper cites Geometrically-driven aggregation for zero-shot 3d point cloud understanding.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Geometrically-driven aggregation for zero-shot 3d point cloud understanding

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:11.556463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:06.050217Z digest=sha256:6c6ab51adcd1da968d34b8877415a35d0629c344f492ebf5f9d144968c7abfbb

Observation b8a750a7-0b91-4765-aeb2-b2dc3197e3a5 · outbound

This paper cites Cus3d: Clip-based unsupervised 3d segmentation via object-level denoise.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Cus3d: Clip-based unsupervised 3d segmentation via object-level denoise

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:11.244506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:06.168100Z digest=sha256:7463f4d425b727e6cfced14d09fc2a28ef69304bf3b218e50af2fe782d347f96

Observation 3dfdcbdc-e185-4dc7-a04a-2d9f166a272e · outbound

This paper cites 3d shapenets: A deep representation for volumetric shapes.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding 3d shapenets: A deep representation for volumetric shapes

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:11.001973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:06.311658Z digest=sha256:91f3f7b5df9c3c56b20a863823fce42105d9aed97e720ca1c5ef1e0bc7dca986

Observation 73a06799-2662-45ac-b295-20f60a72e594 · outbound

This paper cites Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:10.740107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:06.397888Z digest=sha256:8f45bccc728d621debabba6dd48cfdc195dd832ec9e5c485b65f88ba45958094

Observation 87b5d476-d0cd-4e47-a5e8-e32425b6345c · outbound

This paper cites Vconv-dae: Deep volumetric shape learning without object labels.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Vconv-dae: Deep volumetric shape learning without object labels

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:10.534535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:06.472466Z digest=sha256:927cdf8863e1887a9b84487f7196c5cc0fe45663aaea10d36900d6aa1897b7a7

Observation 7bba723b-3816-4e40-b358-f49efd7c79fe · outbound

This paper cites PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:06.584313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:06.584313Z digest=sha256:3e74aebce89cb9f3609c2cbb3ada80fe1180a9e1ce79713fe6b19e42a0009eaa

Observation c75da201-77e5-4e02-9ad6-d3f7c5f9dbbc · outbound

This paper cites Point-bert: Pre- training 3d point cloud transformers with masked point modeling.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Point-bert: Pre- training 3d point cloud transformers with masked point modeling

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:20.714521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:06.698938Z digest=sha256:00231c94fbb5597f200bb2d0f500057f8da869c0854cdae4fa002b73ca49dbdb

Observation 13a1e11d-f80a-4ab3-a6ad-183bcfe33d6d · outbound

This paper cites Masked discrimination for self-supervised learning on point clouds.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Masked discrimination for self-supervised learning on point clouds

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:10.262890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:06.793253Z digest=sha256:cbb24cfcf37aff3c5bc1cf942c12ac92c6c49537cb7301c721a8f9a353d55e52

Observation 7c03a7de-fa71-42c5-b8ec-128dd0eba52f · outbound

This paper cites Masked Surfel Prediction for Self-Supervised Point Cloud Learning.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Masked Surfel Prediction for Self-Supervised Point Cloud Learning

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:31:07.774994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:06.867336Z digest=sha256:62f6abfc9f70491ee0f650b41287dbb944892ceab1f2d4517856600207d05645

Observation 2b7acd45-e3fd-4e4c-b0e3-6aaf3fe37ff0 · outbound

This paper cites Towards compact 3d representations via point feature enhancement masked autoencoders.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Towards compact 3d representations via point feature enhancement masked autoencoders

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:10.038456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:06.942568Z digest=sha256:53bd5dc29aedad83d07bfd480d3660e47ccb44311ad6a714b31cebf08252c0bc

Observation 592ca33c-53f7-4c3c-ac42-0fb60a4d7661 · outbound

This paper cites Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:09.782559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:07.062231Z digest=sha256:66ebe7b29827e37361f7168db3e196ea9ffb6fe2a263a7a5291db652833b8246

Observation 0ed4aa4f-1760-4115-9eef-de4175e85b0b · outbound

This paper cites A scalable active framework for region annotation in 3d shape collections.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding A scalable active framework for region annotation in 3d shape collections

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:09.563356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:07.225619Z digest=sha256:1965fe8f4050d63a8479d7811076be09538de14dc614c73b51fef21e1db1639f

Observation f8750294-1275-4aed-a0c6-3b8c9be49ec0 · outbound

This paper cites Unsupervised point cloud pre-training via occlusion completion.

Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding Unsupervised point cloud pre-training via occlusion completion

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:31:09.306619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:31:07.374604Z digest=sha256:f3eccff77e127c7ede5b086af584ffa0c597bd5192015b6cbe8854f992770c5a

Pith citing papers

Observation 5abb7620-5d9e-420d-a649-a1b09772782e · inbound

Ultra Ethernet's Design Principles and Architectural Innovations cites this paper.

Ultra Ethernet's Design Principles and Architectural Innovations Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:34:15.124646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T17:34:14.621355Z digest=sha256:3039d2b38dec8c8cdd80e82fe4ae21fbf891f94fc5f00b1134cfcd59e7d9d267