Pith. sign in

Paper Citation Record · LEDGER

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP

As of 14 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2603.05962.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.05962 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-15T14:07:39.892672Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b75b2e5-1c3a-41ef-b654-34161a5a6193 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:d1c15962e81306c04a628643a93a08a631cb874de3f4adaf3c6c45f0441f6734

Observation 16f48e47-3aa3-4e18-b672-e4c201c07bc8 · outbound

This paper cites Masked autoencoders are scalable vision learners,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Masked autoencoders are scalable vision learners,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:41afec0f5ff8b088348864caa8e7f9a2f1dcc4efd0630a5b1d439a362e1a3251

Observation 92c9d89a-73d8-4c1b-ac04-65fb376299ef · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:d4254743aa9e99b3daca54319c275799be78400c3d4291004f926fd0cc230df2

Observation 611b93f7-047a-4f4e-93b2-bb6f6c8d3918 · outbound

This paper cites Flava: A foundational language and vision alignment model,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Flava: A foundational language and vision alignment model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:86a6f619d701a136172c8fcdc3c88b09b858a7aabe6807a3430663208a530f60

Observation 0fa5598e-b729-4188-9b37-d89559577fa9 · outbound

This paper cites Large Language Models Can Understanding Depth from Monocular Images.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Large Language Models Can Understanding Depth from Monocular Images

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:cb723742cd88953b2a1c997740b641e10b25af0d774a7c19324172833de9b1bf

Observation b723cfb6-611a-40cd-8b77-87c161635084 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Depth anything: Unleashing the power of large-scale unlabeled data,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:916161d66d223035bb45488fcc3e00ff076daead8849142756536aef5c57227b

Observation 111a5061-44b3-4a7b-b6f5-f05989653757 · outbound

This paper cites Pad: Self-supervised pre-training with patchwise-scale adapter for infrared images,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Pad: Self-supervised pre-training with patchwise-scale adapter for infrared images,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:68b01d55795d43d42cfc55c9d35dfa5e815a473fe3294d5a8eebadb67887f2f5

Observation 4d44a9db-978a-4e93-a2e7-8b5f2b33831d · outbound

This paper cites F-ViTA: Foundation Model Guided Visible to Thermal Translation.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP F-ViTA: Foundation Model Guided Visible to Thermal Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:167668d09047a4eb4c42ddbbfb29f69ea03e5d4a33488b9b54c72d7a1f367309

Observation cb6d7a1b-7ea7-4cf0-8a8e-490e188954be · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:f611bc855a1e72ea8364f5f0a9fc67aa2f03f7b0b4442c2bf87d2fe83178c382

Observation a1a9f6db-c3d4-4f08-a9f3-795658efce97 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Videomae v2: Scaling video masked autoencoders with dual masking,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:e33f7e4a37cc7539b898d09e0c699f06e7d7a9335233720d6a71adbb27d0ae1a

Observation 5cd75477-a027-4bc3-a3bb-406a26e6666f · outbound

This paper cites Omnivl: One foundation model for image-language and video- language tasks,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Omnivl: One foundation model for image-language and video- language tasks,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:e21df94167fb2ff50b2f696059577f6dce9478ed1e283000db524dd3171df0f8

Observation 2488037c-a727-417a-86f3-ba826fa6122d · outbound

This paper cites P2p: Tuning pre-trained image models for point cloud analysis with point-to-pixel prompting,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP P2p: Tuning pre-trained image models for point cloud analysis with point-to-pixel prompting,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:525423ebd6e725870cb38ce224da526f0fd0b6de3dfba5040bc190dd96ef9697

Observation 8ff34b17-245a-4985-b3c0-d6a3df311583 · outbound

This paper cites Pointclip: Point cloud understanding by clip,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Pointclip: Point cloud understanding by clip,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:ac9edb106aa24f3fc3986b9e2d2e3f2dbf1ab70904966dc15c292daf8ea50f55

Observation e8ed976e-f73d-4c0e-a362-3c5a0ec23c39 · outbound

This paper cites Diffusion models as masked autoencoders,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Diffusion models as masked autoencoders,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:7036979bb025988c44b5cb83094cdbc2aaaca4985f98ee4cb9718c6537337f6b

Observation e2978196-0f8b-4c2b-8286-b08a02d76dab · outbound

This paper cites Hierarchical recurrent neural network for skeleton based action recogni- tion,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Hierarchical recurrent neural network for skeleton based action recogni- tion,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:e7c6be05c1d62abb6250a302583077c8e2fad8a5b070357d66d8bcc04437172b

Observation 5523ccd9-7c6d-4732-8217-9691868bda82 · outbound

This paper cites Skeleton-based action recognition using spatio-temporal lstm network with trust gates,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton-based action recognition using spatio-temporal lstm network with trust gates,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:a5280f4529165c2d76e15d4c22ac45698109f168cecefa189db24cdbfd37c864

Observation 796176f6-e6ad-4eae-a3b5-6aef00f31692 · outbound

This paper cites View adaptive recurrent neural networks for high performance human action recognition from skeleton data,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP View adaptive recurrent neural networks for high performance human action recognition from skeleton data,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:95be2b30effb50bb6e2f4347f639d4bc05c723f1b53f502b46029efbeb40d438

Observation 20f3bb99-eea7-42ce-890b-d1798d643c43 · outbound

This paper cites Skeleton based action recognition with convolutional neural network,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton based action recognition with convolutional neural network,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:4325be059e9c82869cc5e3031da9532bed0df68b798cb25dcce2eea65a90ae39

Observation 15bc7c7d-9e61-43b4-955d-1681751aecb2 · outbound

This paper cites A new representation of skeleton sequences for 3d action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP A new representation of skeleton sequences for 3d action recognition,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:eed3376539b2f77709b0bee437a7def27d45f4678930edcbb71ea77cde8ec454

Observation 2e481554-1aec-4c89-a9af-94710de46594 · outbound

This paper cites Co-occurrence feature learning from skeleton data for action recog- nition and detection with hierarchical aggregation,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Co-occurrence feature learning from skeleton data for action recog- nition and detection with hierarchical aggregation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:9d9ad687db86b669fd252a354da8e99e927ca5721eb836d805b919fb47ea26a9

Observation 35da2927-a66f-4c57-8013-ed70c2d900e5 · outbound

This paper cites Spatial temporal graph convolutional networks for skeleton-based ac- tion recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Spatial temporal graph convolutional networks for skeleton-based ac- tion recognition,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:59cedeccf557b038121de855cb8d6aefa695037826343524a6fd4d4db6a9769a

Observation 3b0ed4d5-9bcd-4331-8754-fc6dbfa591b0 · outbound

This paper cites Two- stream adaptive graph convolutional networks for skeleton-based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Two- stream adaptive graph convolutional networks for skeleton-based action recognition,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:9d1658ab350e02a6d1a25b33a84257c794f542f9e1b9a869a4935ea9e66fcc28

Observation 685c5728-0bbd-42f3-a779-efdebabb2314 · outbound

This paper cites Channel-wise topology refinement graph convolution for skeleton-based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Channel-wise topology refinement graph convolution for skeleton-based action recognition,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:beedee68c9d4e068c25fd1aca3b48c02e8e07db25a68d9378554285cf23d509a

Observation 0dd03cef-bd49-4e90-8736-0dc1cf81cf84 · outbound

This paper cites Stst: Spatial-temporal specialized transformer for skeleton- based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Stst: Spatial-temporal specialized transformer for skeleton- based action recognition,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:545d992d8e1a4122aebc82a643ae918ce0285d76f8a40f4ffddddd502f088147

Observation 40655a0a-789d-48f3-841e-4af85e150dc8 · outbound

This paper cites Hypergraph transformer for skeleton-based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Hypergraph transformer for skeleton-based action recognition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:3af9c45f52d6bed284f8b681661144936eb95cf9cb25f8b23b8e6e567ab0f750

Observation 0e4e9a83-972a-4ea9-bee0-aafbbdb67098 · outbound

This paper cites 3d human action representation learning via cross-view consistency pursuit,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP 3d human action representation learning via cross-view consistency pursuit,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:ed60d25dcab33190d4526b3ba01484a3393ac6ee5ff255389eeaec7f19f25109

Observation ff3fe65c-ac7b-4946-866c-0cf3d22890f4 · outbound

This paper cites Contrastive learning from extremely aug- mented skeleton sequences for self-supervised action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Contrastive learning from extremely aug- mented skeleton sequences for self-supervised action recognition,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:7501437ea16485be41a111b692cacd091f7a29ea0bf3f57e541fac92247e7df2

Observation 36cbc4a1-ebc3-4c8d-9e15-b43af9bbe8e0 · outbound

This paper cites Contrastive positive mining for unsupervised 3d action represen- tation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Contrastive positive mining for unsupervised 3d action represen- tation learning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:d2b1183cff04cb1069b3e337efdc4a2b0e02fe78347196866bf6ce78e337e7b2

Observation d7610019-3dee-4acd-ae89-32d6669e20f4 · outbound

This paper cites Masked motion predictors are strong 3d action representation learners,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Masked motion predictors are strong 3d action representation learners,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:86eca7db5044cf05fc9155c4376e5ba1ae9e0579df8d3714dd5607def8ecd4d6

Observation e1b9ff96-6b20-470d-83d8-866757908ec9 · outbound

This paper cites Macdiff: Unified skeleton modeling with masked conditional diffusion,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Macdiff: Unified skeleton modeling with masked conditional diffusion,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:96c8dae1df3c1c3daf8deb36508cacba5bd2b89a91495f208dcf8b10be149aaf

Observation 8e3d93a3-9fd4-44d4-ad14-3bc87e06cc2c · outbound

This paper cites Skeletonmae: Spatial-temporal masked au- toencoders for self-supervised skeleton action recog- nition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeletonmae: Spatial-temporal masked au- toencoders for self-supervised skeleton action recog- nition,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:fa9643ce30ee1f6eabe7d46cfa63ef04d3b75623412853ef873f749d699cd1f3

Observation c1d1644d-adc5-4f5d-9dae-0c4a96b9cc1a · outbound

This paper cites Momen- tum contrast for unsupervised visual representation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Momen- tum contrast for unsupervised visual representation learning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:672b588bd73d434c1c94563cd08e006c8cb200afce8930e452ba5dcf038d520d

Observation bc34bf90-a585-42aa-b5e5-4d1c071e875b · outbound

This paper cites A simple framework for contrastive learning of visual representations,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP A simple framework for contrastive learning of visual representations,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:8e97363464e9a0dee72b06c5677e9bf3fd959a28ac7f5a17581376aa8985c783

Observation 2d1cac2a-ec23-459a-ac93-88679ac68309 · outbound

This paper cites Spatiotemporal contrastive video representation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Spatiotemporal contrastive video representation learning,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:155c7c197b396a323de42f00afe30d2f8891dcf52dd6c4d14387fd80316f1597

Observation a1e20b50-8770-416c-9a95-5c0ebfe960be · outbound

This paper cites BEit: BERT pre-training of image transformers,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP BEit: BERT pre-training of image transformers,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:e6b7ac018b792dc0fffadf73a31a711c1e6aeb3ac0d58a7452affd31c5711d30

Observation 963da4a3-adcc-4625-9d6d-06b317d49c95 · outbound

This paper cites Masked feature prediction for self- supervised visual pre-training,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Masked feature prediction for self- supervised visual pre-training,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:9594fbbd8c45db8cc6734042edc6fc623055f3b28c79b806b603c784b7266f2f

Observation d72f2d0e-9b95-4758-a0ca-4fc926622d97 · outbound

This paper cites Dynamic multiscale graph neural networks for 3d skeleton based human motion prediction,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Dynamic multiscale graph neural networks for 3d skeleton based human motion prediction,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:4f9eb2f0a3ae45431560b04dfa8533dcd896fd1638fcf944c1b6dd3c8452ab06

Observation 45bf7f52-8a73-4808-89d0-3ae3b0957aa4 · outbound

This paper cites Exploiting spatial-temporal relationships for 3d pose estimation via graph con- volutional networks,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Exploiting spatial-temporal relationships for 3d pose estimation via graph con- volutional networks,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:fa5aabbef7592873126e67d30e1a01b31caa0750bc15c780d98f7866f0093f04

Observation e7fc133c-520b-4a72-8b5e-26dd6b02075d · outbound

This paper cites Multi-Scale Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Multi-Scale Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:b7b12924559ab76981ba41ab193b8e449c0bb0442b1deb2998bad15ad13c739b

Observation d86fc034-19a8-4dd6-919a-b288fbf868cd · outbound

This paper cites Skele- ton cloud colorization for unsupervised 3d action rep- resentation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skele- ton cloud colorization for unsupervised 3d action rep- resentation learning,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:3c4d97616439632e33ff8ddd212b4a36db0804451e02bb3f59faeda09832136d

Observation 954ff6b5-b7a8-4f83-9fd0-fb72b4f4c071 · outbound

This paper cites Collaborating domain-shared and target-specific feature clustering for cross-domain 3d action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Collaborating domain-shared and target-specific feature clustering for cross-domain 3d action recognition,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:24a050a542a25064e4148284c3f50a873bafb17395e91a4ee4e998e90d3c1c4a

Observation ad617174-7455-431d-a686-caf1be3f010a · outbound

This paper cites Ntu rgb+d: A large scale dataset for 3d human activity analysis,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Ntu rgb+d: A large scale dataset for 3d human activity analysis,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:6f9defe80f512c2e1afb4212bb39e570819b1655c216967dfe2a0371d23dfccb

Observation 740dac90-8092-4db9-8de0-d5d43b59df45 · outbound

This paper cites Ntu rgb+d 120: A large-scale bench- mark for 3d human activity understanding,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Ntu rgb+d 120: A large-scale bench- mark for 3d human activity understanding,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:6f6f41498ed93c410a0bba2bfc55d67824e727bbb9cbb4c5c4719f9ca18bce51

Observation ff22536a-2572-415b-b8c0-3906f6120615 · outbound

This paper cites A bench- mark dataset and comparison study for multi-modal human action analytics,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP A bench- mark dataset and comparison study for multi-modal human action analytics,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:f58cadd1e0fe76690bb3f7193e61588ac93437d77254d3ab0060308ed6a51fc4

Observation f5abd89c-e6a2-404d-9f42-56e41e5d601f · outbound

This paper cites Cross- view action modeling, learning and recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Cross- view action modeling, learning and recognition,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:90074ecf5ce8f9b5623d5e0a2d23582ef926f965b749c3409be02a36625fb7af

Observation d2685dd8-9338-4a53-9c5f-5a8340194899 · outbound

This paper cites Toyota smarthome: Real-world activities of daily living,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Toyota smarthome: Real-world activities of daily living,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:3e825847da0499a366bc28d8e49ec785bad750b0ad29bcb07f83e79d58402bc6

Observation e3c34f84-7f86-4f64-bbe4-10bc643a5541 · outbound

This paper cites Semantics-guided neural networks for efficient skeleton-based human action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Semantics-guided neural networks for efficient skeleton-based human action recognition,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:6b6ffde95d4b4000f5086605cbcb064f47d89804d240c91387b252d43c29c411

Observation 07840b59-f146-4267-82a8-24f6ea7549ab · outbound

This paper cites Skeleton-based action recognition with shift graph convolutional network,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton-based action recognition with shift graph convolutional network,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:37ce7c1c5619a4ed39c167f59fd99653f17f6eb88a18ee82614ff2287c40ca0f

Observation b3831f3d-0de6-44a6-a7fa-d9684c4d738e · outbound

This paper cites Unsupervised representation learning with JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13 long-term dynamics for skeleton based action recogni- tion,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Unsupervised representation learning with JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13 long-term dynamics for skeleton based action recogni- tion,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:3d9f3467ba7a47de162abdc517e95914452765ce88e6ae602b778e290944e264

Observation a4dc1a56-1fcd-4595-93c2-7a8bcba78426 · outbound

This paper cites Predict & cluster: Unsupervised skeleton based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Predict & cluster: Unsupervised skeleton based action recognition,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:a9993292274d9731ea3b0e7444a78b95748359586600bb45cabf83d626b0f0a2

Observation 366592a3-c611-4679-9308-2319924da831 · outbound

This paper cites Ms2l: Multi- task self-supervised learning for skeleton based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Ms2l: Multi- task self-supervised learning for skeleton based action recognition,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:abc822cd2fac47b85ec50087ab18a11f00874f4d988d359429c9fd7febff3646

Observation a19d53fe-c254-490d-9dd5-be4d7fee82b3 · outbound

This paper cites Skeleton- contrastive 3d action representation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton- contrastive 3d action representation learning,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:3814f9c1d4e579a3ba3f9fcc4cb49cdbfda41607535395739a38e6c31f325500

Observation 802284f4-345b-49f0-ac4c-5ccd63fcca5c · outbound

This paper cites Global- local motion transformer for unsupervised skeleton- based action learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Global- local motion transformer for unsupervised skeleton- based action learning,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:b5c4469b2685a7651e9bc5e8aef0011056a08bb366b1fbf85462af4ee9a274a3

Observation 7e277e30-1c62-4f1d-9927-ac0c2ce9a2f2 · outbound

This paper cites Cmd: Self-supervised 3d action representation learning with cross-modal mutual distillation,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Cmd: Self-supervised 3d action representation learning with cross-modal mutual distillation,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:0e0aa139a3e622e321a96267e1f85893cc320ae890f641f6e76e43dba6744152

Observation c8c519ac-761c-4d41-afd3-011be7da6b74 · outbound

This paper cites Actionlet-dependent contrastive learning for unsupervised skeleton-based action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Actionlet-dependent contrastive learning for unsupervised skeleton-based action recognition,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:07889a7f7425c0a69afde42e705c52b7e39c5be9844c3007095dc7deaa0061a7

Observation 90aecfd8-ccd5-491b-92e3-9ecb43fed72d · outbound

This paper cites Self-supervised 3d skeleton action representation learning with motion consis- tency and continuity,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Self-supervised 3d skeleton action representation learning with motion consis- tency and continuity,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:ad33e14132f1c82d893bb3a99cc4820217f055d755b364afa95650b5656de10d

Observation 6db2b46a-de0f-4d0c-ba3b-f61c51bc18c4 · outbound

This paper cites View-invariant skele- ton action representation learning via motion retarget- ing,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP View-invariant skele- ton action representation learning via motion retarget- ing,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:507c76e8f0f080080d7787ec883ee954f88cc9569bd94004ea608c354b66b661

Observation d3d46f76-0065-48ac-8b8a-6ff8312fb83e · outbound

This paper cites Hierarchically self- supervised transformer for human skeleton represen- tation learning,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Hierarchically self- supervised transformer for human skeleton represen- tation learning,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:4c9e8613253492e8ee494bbed25373800222be56d8ead85d9ffa8c66e5b8bb2e

Observation c66f9d00-dc82-41f9-aac8-616cf56e6707 · outbound

This paper cites Adversarial self-supervised learning for semi-supervised 3d action recognition,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Adversarial self-supervised learning for semi-supervised 3d action recognition,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:acaf1c0e1f42e63b79d017ecf30172c9ce1a6c4b4d51aa60fb30c89f4e88f952

Observation 3d8de597-38e6-4b43-a71f-51f897f56250 · outbound

This paper cites UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:ce3915646c77fcb23c8c5fdae3ed40b7911f6f1f59b5ddd1c4ac0e3ad6bfe46c

Observation 0d29d8a5-a399-450c-a9ac-72ec1b85f4ca · outbound

This paper cites Attention is all you need,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Attention is all you need,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:11653f5c2ed75fc7a8a3593a367a480cfc3297552a695d7147ea2413ff6dcf9b

Observation 074e8d82-25c1-415f-b0b8-120a18f04c28 · outbound

This paper cites Layer Normalization.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Layer Normalization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:222acede8af8e2407f54f3d55f5ceaa76fd8e34b9bbab5be4bda0964163a29c6

Observation 1649a5a0-0926-4997-bedf-1ebce901ada5 · outbound

This paper cites Denoising diffusion probabilistic models,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Denoising diffusion probabilistic models,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:f5e0f9ad202be7bd242c596d70fbc385b0396f3573505ba474bf9537c1c607fa

Observation 9cb9a11b-b15b-47ae-9fe5-27f1c93b65f1 · outbound

This paper cites Lcr-net++: Multi-person 2d and 3d pose detection in natural im- ages,.

A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Lcr-net++: Multi-person 2d and 3d pose detection in natural im- ages,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-15T14:07:39.892672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:07:39.892672Z digest=sha256:82fd768ed422352d0b7471da1c3663f8e2641f728e9907c3885ed9a9d041d5e8

Pith citing papers

No inbound Pith citation observations are available.