Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-15T14:07:39.892672Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2603.05962.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-15T14:07:39.892672Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0b75b2e5-1c3a-41ef-b654-34161a5a6193 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP An image is worth 16x16 words: Transformers for image recognition at scale,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16f48e47-3aa3-4e18-b672-e4c201c07bc8 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Masked autoencoders are scalable vision learners,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92c9d89a-73d8-4c1b-ac04-65fb376299ef · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Scaling up visual and vision-language representation learning with noisy text supervision,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 611b93f7-047a-4f4e-93b2-bb6f6c8d3918 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Flava: A foundational language and vision alignment model,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fa5598e-b729-4188-9b37-d89559577fa9 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Large Language Models Can Understanding Depth from Monocular Images
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b723cfb6-611a-40cd-8b77-87c161635084 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Depth anything: Unleashing the power of large-scale unlabeled data,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111a5061-44b3-4a7b-b6f5-f05989653757 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Pad: Self-supervised pre-training with patchwise-scale adapter for infrared images,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d44a9db-978a-4e93-a2e7-8b5f2b33831d · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP F-ViTA: Foundation Model Guided Visible to Thermal Translation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb6d7a1b-7ea7-4cf0-8a8e-490e188954be · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a9f6db-c3d4-4f08-a9f3-795658efce97 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Videomae v2: Scaling video masked autoencoders with dual masking,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cd75477-a027-4bc3-a3bb-406a26e6666f · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Omnivl: One foundation model for image-language and video- language tasks,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2488037c-a727-417a-86f3-ba826fa6122d · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP P2p: Tuning pre-trained image models for point cloud analysis with point-to-pixel prompting,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ff34b17-245a-4985-b3c0-d6a3df311583 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Pointclip: Point cloud understanding by clip,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ed976e-f73d-4c0e-a362-3c5a0ec23c39 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Diffusion models as masked autoencoders,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2978196-0f8b-4c2b-8286-b08a02d76dab · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Hierarchical recurrent neural network for skeleton based action recogni- tion,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5523ccd9-7c6d-4732-8217-9691868bda82 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton-based action recognition using spatio-temporal lstm network with trust gates,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 796176f6-e6ad-4eae-a3b5-6aef00f31692 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP View adaptive recurrent neural networks for high performance human action recognition from skeleton data,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20f3bb99-eea7-42ce-890b-d1798d643c43 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton based action recognition with convolutional neural network,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15bc7c7d-9e61-43b4-955d-1681751aecb2 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP A new representation of skeleton sequences for 3d action recognition,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e481554-1aec-4c89-a9af-94710de46594 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Co-occurrence feature learning from skeleton data for action recog- nition and detection with hierarchical aggregation,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35da2927-a66f-4c57-8013-ed70c2d900e5 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Spatial temporal graph convolutional networks for skeleton-based ac- tion recognition,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b0ed4d5-9bcd-4331-8754-fc6dbfa591b0 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Two- stream adaptive graph convolutional networks for skeleton-based action recognition,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 685c5728-0bbd-42f3-a779-efdebabb2314 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Channel-wise topology refinement graph convolution for skeleton-based action recognition,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd03cef-bd49-4e90-8736-0dc1cf81cf84 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Stst: Spatial-temporal specialized transformer for skeleton- based action recognition,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40655a0a-789d-48f3-841e-4af85e150dc8 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Hypergraph transformer for skeleton-based action recognition,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e4e9a83-972a-4ea9-bee0-aafbbdb67098 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP 3d human action representation learning via cross-view consistency pursuit,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff3fe65c-ac7b-4946-866c-0cf3d22890f4 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Contrastive learning from extremely aug- mented skeleton sequences for self-supervised action recognition,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36cbc4a1-ebc3-4c8d-9e15-b43af9bbe8e0 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Contrastive positive mining for unsupervised 3d action represen- tation learning,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7610019-3dee-4acd-ae89-32d6669e20f4 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Masked motion predictors are strong 3d action representation learners,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b9ff96-6b20-470d-83d8-866757908ec9 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Macdiff: Unified skeleton modeling with masked conditional diffusion,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e3d93a3-9fd4-44d4-ad14-3bc87e06cc2c · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeletonmae: Spatial-temporal masked au- toencoders for self-supervised skeleton action recog- nition,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d1644d-adc5-4f5d-9dae-0c4a96b9cc1a · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Momen- tum contrast for unsupervised visual representation learning,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc34bf90-a585-42aa-b5e5-4d1c071e875b · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP A simple framework for contrastive learning of visual representations,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d1cac2a-ec23-459a-ac93-88679ac68309 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Spatiotemporal contrastive video representation learning,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e20b50-8770-416c-9a95-5c0ebfe960be · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP BEit: BERT pre-training of image transformers,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 963da4a3-adcc-4625-9d6d-06b317d49c95 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Masked feature prediction for self- supervised visual pre-training,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72f2d0e-9b95-4758-a0ca-4fc926622d97 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Dynamic multiscale graph neural networks for 3d skeleton based human motion prediction,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45bf7f52-8a73-4808-89d0-3ae3b0957aa4 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Exploiting spatial-temporal relationships for 3d pose estimation via graph con- volutional networks,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7fc133c-520b-4a72-8b5e-26dd6b02075d · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Multi-Scale Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86fc034-19a8-4dd6-919a-b288fbf868cd · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skele- ton cloud colorization for unsupervised 3d action rep- resentation learning,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954ff6b5-b7a8-4f83-9fd0-fb72b4f4c071 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Collaborating domain-shared and target-specific feature clustering for cross-domain 3d action recognition,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad617174-7455-431d-a686-caf1be3f010a · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Ntu rgb+d: A large scale dataset for 3d human activity analysis,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740dac90-8092-4db9-8de0-d5d43b59df45 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Ntu rgb+d 120: A large-scale bench- mark for 3d human activity understanding,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff22536a-2572-415b-b8c0-3906f6120615 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP A bench- mark dataset and comparison study for multi-modal human action analytics,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5abd89c-e6a2-404d-9f42-56e41e5d601f · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Cross- view action modeling, learning and recognition,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2685dd8-9338-4a53-9c5f-5a8340194899 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Toyota smarthome: Real-world activities of daily living,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c34f84-7f86-4f64-bbe4-10bc643a5541 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Semantics-guided neural networks for efficient skeleton-based human action recognition,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07840b59-f146-4267-82a8-24f6ea7549ab · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton-based action recognition with shift graph convolutional network,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3831f3d-0de6-44a6-a7fa-d9684c4d738e · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Unsupervised representation learning with JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13 long-term dynamics for skeleton based action recogni- tion,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4dc1a56-1fcd-4595-93c2-7a8bcba78426 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Predict & cluster: Unsupervised skeleton based action recognition,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 366592a3-c611-4679-9308-2319924da831 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Ms2l: Multi- task self-supervised learning for skeleton based action recognition,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19d53fe-c254-490d-9dd5-be4d7fee82b3 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Skeleton- contrastive 3d action representation learning,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 802284f4-345b-49f0-ac4c-5ccd63fcca5c · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Global- local motion transformer for unsupervised skeleton- based action learning,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e277e30-1c62-4f1d-9927-ac0c2ce9a2f2 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Cmd: Self-supervised 3d action representation learning with cross-modal mutual distillation,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8c519ac-761c-4d41-afd3-011be7da6b74 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Actionlet-dependent contrastive learning for unsupervised skeleton-based action recognition,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90aecfd8-ccd5-491b-92e3-9ecb43fed72d · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Self-supervised 3d skeleton action representation learning with motion consis- tency and continuity,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db2b46a-de0f-4d0c-ba3b-f61c51bc18c4 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP View-invariant skele- ton action representation learning via motion retarget- ing,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d46f76-0065-48ac-8b8a-6ff8312fb83e · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Hierarchically self- supervised transformer for human skeleton represen- tation learning,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c66f9d00-dc82-41f9-aac8-616cf56e6707 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Adversarial self-supervised learning for semi-supervised 3d action recognition,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d8de597-38e6-4b43-a71f-51f897f56250 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d29d8a5-a399-450c-a9ac-72ec1b85f4ca · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Attention is all you need,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 074e8d82-25c1-415f-b0b8-120a18f04c28 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Layer Normalization
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1649a5a0-0926-4997-bedf-1ebce901ada5 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Denoising diffusion probabilistic models,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb9a11b-b15b-47ae-9fe5-27f1c93b65f1 · outbound
A novel Framework for Open-Vocabulary Multi-Object Recognition using CLIP Lcr-net++: Multi-person 2d and 3d pose detection in natural im- ages,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.