Pith. sign in

Paper Citation Record · LEDGER

Long-CLIP: Unlocking the Long-Text Capability of CLIP

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2403.15378.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.15378 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:56:06.568891Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:29:57.890361Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03a6e17e-2b29-4780-b90a-e008cf09c589 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.852840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:8fdd895fdddf0559a6afebfbf9fe6db6f02d305ca25c9dab91a4b96e83edb09f

Observation e14f8811-1c17-4665-b344-4424d20d57a6 · inbound

E5-V: Universal Embeddings with Multimodal Large Language Models cites this paper.

E5-V: Universal Embeddings with Multimodal Large Language Models Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:52:20.952536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T22:52:20.935555Z digest=sha256:78028dcd22b648dd7e2a5c6fcd81ff512024e006c82fa5762f534e61c9f0da32

Observation 2d24c723-707f-469f-b490-3bc80d67e801 · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 141

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.568891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.568891Z digest=sha256:e76bd12ccacd74cf8a263173e2ceac1996b0aa943f6f34c9fd6ed0ee2baa48f0

Observation 2a64a548-2eac-47b4-bf76-44d2785771f1 · inbound

ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries cites this paper.

ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:48.404331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:48.404331Z digest=sha256:70b802bc593d85f8b59f5756a149da4ed6276576e086e7d2979e116400425dbe

Observation 34575fdb-50b9-4bf4-a42f-09ffe95cc596 · inbound

I0T: Embedding Standardization Method Towards Zero Modality Gap cites this paper.

I0T: Embedding Standardization Method Towards Zero Modality Gap Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:20:33.175525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:20:33.175525Z digest=sha256:d4dcc57cf671a01aa0e32df56898651896713386bf23809a1471e31c2b257e95

Observation d0f24e1a-4b21-496e-8490-527d6d8d12d8 · inbound

Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection cites this paper.

Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:51.947862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:12:51.947862Z digest=sha256:77d263a0d152f1206bb9822d0f4389c65bb716bba070030f36d9ef91b185c1e4

Observation 1d638a5f-9c0e-4fc3-bba7-784355d7cbdb · inbound

Seeing with Partial Certainty: Conformal Prediction for Robotic Scene Recognition in Built Environments cites this paper.

Seeing with Partial Certainty: Conformal Prediction for Robotic Scene Recognition in Built Environments Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:26:07.389266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:26:07.389266Z digest=sha256:a1593eb8418b23d7db37f1099bab1eb8883a66fe9cb093857dcb59d97c38f51c

Observation 6cb0ee77-9ec0-4717-8e75-9b406f2599c7 · inbound

Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency cites this paper.

Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T23:40:21.599948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:40:21.599948Z digest=sha256:285c3f8d935c058369bf69ce16fa80f94c464cf4d475f4d7a888d55e632b6f46

Observation 5b51070d-d63d-4fc0-976a-77c714bd43a5 · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.217443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.217443Z digest=sha256:1fc9fc44501d643a7eba8745a42f281e6383689cc0a6c1984171390e9fb00053

Observation 1d6f4e5d-57fc-41b3-9b1f-fd9d40e586a1 · inbound

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation cites this paper.

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:55.385363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:55.385363Z digest=sha256:45eb7b4ebdb50fe5f4799b6387655d6bf6b7a501113e61bb787a377c558306b1

Observation 766675df-dbfd-4115-975d-536f1a31855b · inbound

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations cites this paper.

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.089363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.089363Z digest=sha256:d32333662b168e993a07fdcf1723c039844f66c90096f4a41db3a29538309e3e

Observation 5a9871a1-57b7-40e4-bd16-eda9ed779b15 · inbound

ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model cites this paper.

ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:47.273964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:47.273964Z digest=sha256:c9bb6cfc60383565a39bcc40ecaab6a91fc6dcfdca8461a871ff7568ac1b6e31

Observation 38f709d8-3a45-402a-a049-1a51bc49f3f1 · inbound

GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles cites this paper.

GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.804567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.804567Z digest=sha256:0fddda46d45cbbf2c4ee4195c67ae49d254320f99387b214bc2b4a8b9c494ad5

Observation d91d90a1-acd1-4d8e-ad0a-ea4de23d1568 · inbound

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text cites this paper.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.501373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.501373Z digest=sha256:2c6d9579af49d7bd8bfc39e171943c82801607d66e161c1d7537752e3de66925

Observation 7fecc64b-1d59-4c7c-9aa8-12589220b431 · inbound

Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation cites this paper.

Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:57.092222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:57.092222Z digest=sha256:e8b6ac667e4b8863c29bf37848b48beb705424f61442ddd8734f598a671a2a41

Observation 4160fdf5-77db-4682-8f0b-cf6b232dfb42 · inbound

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement cites this paper.

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:22:49.772744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:22:49.772744Z digest=sha256:c6bfaa9e14602dc375bab557e359bcf4ef702fa36f82a06d6de42be738ad3232

Observation c4c08ade-a215-4f17-8092-44de9aca2019 · inbound

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach cites this paper.

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.671528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T12:39:56.396231Z digest=sha256:ceca1f392f28fd4186967d4be11e4961956296a27cd4f16455a96ebb94ea2879

Observation 21865d75-2172-4a81-abe7-a9c3a11ce13c · inbound

LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning cites this paper.

LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:22:52.400078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T12:21:19.606631Z digest=sha256:c9306f154f9a618fde9fd7e78469d5352f1337ba69e6eeca8aa0515cfb4a91c8

Observation b9a0f02a-d2d4-4809-99f8-99d96cc9c187 · inbound

AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian Splatting cites this paper.

AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian Splatting Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:50:56.206901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:50:56.206901Z digest=sha256:57523496739b1d99ad3ca6bd582f484849212212344c25cf316cfead213c72f6

Observation 5b837b95-8915-4b84-aca9-6f6f0cc107e9 · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.288367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:e3ccc6ff9ea51f6517f199f41941f5cb5167e9dd9649b11dafb4b68d9aeec6d0

Observation 84f41b62-ca6e-407d-8f7b-90a806f33b21 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.725478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T01:52:14.874049Z digest=sha256:665b36ae76bb6b56e91013fed015ccbd904ac832bdb2b93e63048ebfff601d0a

Observation 62d63b4f-850f-4b2b-8183-983e4e591e26 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:59:06.461323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T21:54:33.902256Z digest=sha256:7557ab5ad05b95fdafe85ece6533a417422b8c9f48e655218b5ada5d3ca78627

Observation 873a7819-028c-47e2-a0c4-a30081ab958e · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.624844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T09:12:35.777810Z digest=sha256:cf9e1fd0aa7486ba5b8ddd7f3f4228b1a6cd6b8690be046594f65a8d789f0f42

Observation 65d9bc3f-5700-4c3f-847a-0ce1a24e5f20 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:45:05.533477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T21:44:21.427570Z digest=sha256:9cceefb77b230f174f7112c1bc6c8c0afda03e757e4e26f1a462043ae2777dd4

Observation ac12bc4f-551e-49c0-9619-29fecc202c2d · inbound

Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation cites this paper.

Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:43:38.647922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T18:41:12.533556Z digest=sha256:d92a6495789ea0cdb4718ce07cbf4c60dbddae8adfd99db6e9a7c846cbfdccf6

Observation 29d9322a-d926-4c8b-a69c-81a353b8bb01 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.630748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:0116f118366bafd852116d476b35b44563b3807917d975e64c0bef1b2a6cf821

Observation ab0c928c-5f9b-456d-95c7-e6a77702e2f0 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.891855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T00:29:41.291832Z digest=sha256:b5c85d71528b745ea6555d958faf8f7ebedcf2a185bcb8d205550a9331807904

Observation f8971470-a521-43ab-9a07-60b75aa612e5 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:03:32.224194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:a307086bf07249d474d3748881c22aea592250ba046880304b7b2bcdb03ef082

Observation 016f124c-dda2-4a77-b01e-e8e688f0ee17 · inbound

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators cites this paper.

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:48.757577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T01:02:20.024793Z digest=sha256:d1a452c3b3f361f87cde5c42ac641ea0cdc300c16424e2b0086e66db1bcd8b94

Observation c9eee4a0-f03e-48e6-b829-770ae20f3b26 · inbound

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators cites this paper.

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T11:46:50.520083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:46:50.520083Z digest=sha256:1dc52cc46b5ff3c10c21e4d42eac44b0d9e76264ecf47d5f6ae21af48bd5f6fa