Pith. sign in

Paper Citation Record · LEDGER

Long-CLIP: Unlocking the Long-Text Capability of CLIP

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2403.15378.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.15378 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:40:21.599948Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:29:57.890361Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03a6e17e-2b29-4780-b90a-e008cf09c589 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.852840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:c0e7ea85530f742f223305e7f8752c1f86f42ca7cc2aa3a1e7b7426f6e881024

Observation e14f8811-1c17-4665-b344-4424d20d57a6 · inbound

E5-V: Universal Embeddings with Multimodal Large Language Models cites this paper.

E5-V: Universal Embeddings with Multimodal Large Language Models Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:52:20.952536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:52:20.935555Z digest=sha256:a6ed4f52a4f7bd065354d01030f59c2a6baf99bb904632cc51e13e8deadf65f0

Observation 6cb0ee77-9ec0-4717-8e75-9b406f2599c7 · inbound

Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency cites this paper.

Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T23:40:21.599948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:40:21.599948Z digest=sha256:4f40ba1c6cf501f3ccfaff9be33a7dae3d24ecfd26829e4e82d51f3215a7a236

Observation 5b51070d-d63d-4fc0-976a-77c714bd43a5 · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.217443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.217443Z digest=sha256:ea9680ee2d75ba7393334c81a7dc1ad1a3c5b0f5a252a8297f64a02d14b2bf85

Observation 1d6f4e5d-57fc-41b3-9b1f-fd9d40e586a1 · inbound

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation cites this paper.

Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:55.385363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:55.385363Z digest=sha256:45eb7b4ebdb50fe5f4799b6387655d6bf6b7a501113e61bb787a377c558306b1

Observation 766675df-dbfd-4115-975d-536f1a31855b · inbound

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations cites this paper.

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.089363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.089363Z digest=sha256:d32333662b168e993a07fdcf1723c039844f66c90096f4a41db3a29538309e3e

Observation 5a9871a1-57b7-40e4-bd16-eda9ed779b15 · inbound

ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model cites this paper.

ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:47.273964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:47.273964Z digest=sha256:6a1bb21a662750eff560062f5fd77158680314fce871d4fa4fb15dd8439e4876

Observation 38f709d8-3a45-402a-a049-1a51bc49f3f1 · inbound

GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles cites this paper.

GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.804567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.804567Z digest=sha256:0fddda46d45cbbf2c4ee4195c67ae49d254320f99387b214bc2b4a8b9c494ad5

Observation d91d90a1-acd1-4d8e-ad0a-ea4de23d1568 · inbound

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text cites this paper.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.501373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.501373Z digest=sha256:a3ccccdb7d816d34332a447f697a52e547d96905381acaca08e01928912a8320

Observation 7fecc64b-1d59-4c7c-9aa8-12589220b431 · inbound

Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation cites this paper.

Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:57.092222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:57.092222Z digest=sha256:e8b6ac667e4b8863c29bf37848b48beb705424f61442ddd8734f598a671a2a41

Observation 4160fdf5-77db-4682-8f0b-cf6b232dfb42 · inbound

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement cites this paper.

PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:22:49.772744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:22:49.772744Z digest=sha256:c6bfaa9e14602dc375bab557e359bcf4ef702fa36f82a06d6de42be738ad3232

Observation c4c08ade-a215-4f17-8092-44de9aca2019 · inbound

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach cites this paper.

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.671528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T12:39:56.396231Z digest=sha256:8352ba0f98fa636fec11833d80bf528e42faea82abbd987d5abbcf3daf15dfba

Observation 21865d75-2172-4a81-abe7-a9c3a11ce13c · inbound

LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning cites this paper.

LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:22:52.400078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:21:19.606631Z digest=sha256:b6ba34160e3b131eeb8969788a090fd6da72ea4e82cad064429a59b3696a4886

Observation b9a0f02a-d2d4-4809-99f8-99d96cc9c187 · inbound

AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian Splatting cites this paper.

AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian Splatting Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:50:56.206901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:50:56.206901Z digest=sha256:bc2005080aefaf6ad0e14ed359a3b90ac1db02283c79e49ac9b78b68f7459b43

Observation 5b837b95-8915-4b84-aca9-6f6f0cc107e9 · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.288367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:414713c2faf37263ca3b7ce4d6c5ff4f2f121e6cd2dca347badc2345c0e371d7

Observation 84f41b62-ca6e-407d-8f7b-90a806f33b21 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:53:28.725478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T01:52:14.874049Z digest=sha256:b4ccd0340ae665435191365b122ae01e4dc4d19f2b65b95aa7bdb93f382d347b

Observation 62d63b4f-850f-4b2b-8183-983e4e591e26 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:59:06.461323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T21:54:33.902256Z digest=sha256:8986d9d0a764f09f28473e77077348764ae59e6d45c5e9af2a4da9ac329dcf47

Observation 873a7819-028c-47e2-a0c4-a30081ab958e · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.624844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T09:12:35.777810Z digest=sha256:1245e5a4cdc9201df05edddfd7913eb54e4c8a2a021c8308416bdd0783ae81ca

Observation 65d9bc3f-5700-4c3f-847a-0ce1a24e5f20 · inbound

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation cites this paper.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:45:05.533477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:44:21.427570Z digest=sha256:d70b2008308838c06eb4f1af5165432f2504f317478fc3ee686359e4f095424f

Observation ac12bc4f-551e-49c0-9619-29fecc202c2d · inbound

Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation cites this paper.

Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:43:38.647922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T18:41:12.533556Z digest=sha256:51dde1a49ff3342594d88db2fcb90296f7fa66d7e58d6a23f8d30b406c68e5b1

Observation 29d9322a-d926-4c8b-a69c-81a353b8bb01 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.630748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:5c83a77a83cd01959f2b79a9c3cd1d0d58c64dc2059c7f80d3c5a63d2b63a721

Observation ab0c928c-5f9b-456d-95c7-e6a77702e2f0 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.891855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:29:41.291832Z digest=sha256:957dfdd128273c295130ee8a9d9b8111a855ad31ae9b39ef97ad7ed4cec20952

Observation f8971470-a521-43ab-9a07-60b75aa612e5 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:03:32.224194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:d184c68d7658cfe779ef08bdd8f71fe01ea83893306180c6a99fae505a880441

Observation 016f124c-dda2-4a77-b01e-e8e688f0ee17 · inbound

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators cites this paper.

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:48.757577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T01:02:20.024793Z digest=sha256:dde52d15f400a80ced1fff6cd021f09f4036b9a2faa3b80f81c49f4416f9e477

Observation c9eee4a0-f03e-48e6-b829-770ae20f3b26 · inbound

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators cites this paper.

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T11:46:50.520083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:46:50.520083Z digest=sha256:1dc52cc46b5ff3c10c21e4d42eac44b0d9e76264ecf47d5f6ae21af48bd5f6fa