Pith. sign in

Paper Citation Record · LEDGER

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2402.04252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04252 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:41.212978Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 761d635e-cc0e-4563-abd0-2d153a7b3767 · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.804400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:8049d8decdebc0670fab94a1190a3766b7021dac6b95e8c293bdf1ac7a87ffff

Observation 33163676-9942-43da-96e7-7883f6a5b377 · inbound

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM cites this paper.

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:41.212978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:41.212978Z digest=sha256:516e9bd680279779e7451aa5bfcf478d524e1438b1a843a46bdcfdeba39c230d

Observation 6141965c-c4e6-409f-9782-80e21e49fed8 · inbound

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation cites this paper.

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:26.469919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:42:26.469919Z digest=sha256:535a66afce171fb94106b2f155d22931aff994f41c43e204021483662a82322a

Observation 4ac52ac6-a3b7-487b-a43c-42578b84fa9e · inbound

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying cites this paper.

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:18:47.669708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:18:47.669708Z digest=sha256:4062585af9b6418ef755bfb7860b0374e71ac038083ebce67f1ede4e10dccbdc

Observation b287334b-6846-47e4-9a02-02e8a143c3d6 · inbound

Can Argus Judge Them All? Comparing VLMs Across Domains cites this paper.

Can Argus Judge Them All? Comparing VLMs Across Domains EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:52.199619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:52.199619Z digest=sha256:222dc7d7a30c9e6de2b8fad9b920998aeb9a534e0517a1957e0e9f4406e57ed7

Observation 4d1c2b4e-a74b-4f70-8c1c-f0b2b2932a61 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.749785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.749785Z digest=sha256:de8ef6223674f4fb70497477cf3ce96238767d5cc01d147f3b853d1be82a88d5

Observation 1e092c2f-2333-4d1d-849b-1b18121ed4e0 · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.406624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.406624Z digest=sha256:6df54a6e613f4796c06d9ed55f67e08e8591302b3e1db3a60bb168a79afd3041

Observation 5006196a-e851-43d4-b838-bb79c4440a6c · inbound

Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders cites this paper.

Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T13:59:44.047838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:59:44.047838Z digest=sha256:867425b6540b2c89ae03417080cb3a0fc9be8f836f1cceec0a643cb6d9977db9

Observation a5a8c8bd-7223-4a3f-a987-fd10e00d65c9 · inbound

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents cites this paper.

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:41:22.579211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T12:40:19.544260Z digest=sha256:4fb15d7e24ddb00e2c61488136d224b0052d9430bfb02202e1aed1af101aee0d

Observation c74decde-c081-47ff-baf8-5c04afc10656 · inbound

QKVQA: Question-Focused Filtering for Knowledge-based VQA cites this paper.

QKVQA: Question-Focused Filtering for Knowledge-based VQA EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:54.033143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:53:59.668663Z digest=sha256:72a8c0eb9267719011f3cbd5baae93f64d6a4101c30735c36e8dfb74b3cdea9c

Observation 24b1a7b5-fabe-48c9-99c0-5ccc3d4e1fe9 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.341823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.341823Z digest=sha256:0544939ef203ef00e2a686c8adfa5daee029201984c031d5e22087d987f0fe91

Observation 210dbf3f-69e6-41c3-ba71-1ed8a84f9b3b · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:02.006151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:09:51.816554Z digest=sha256:2daf2dbc5b570321c4a6e5f89a55ae49953c4baefc0dc2e7c6d5590f35052922

Observation a43f5ecc-af12-4c0e-b663-17de2f84e7cd · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T22:43:59.727378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:43:59.727378Z digest=sha256:34b64e74aea73be71febc085a5aebeb6a416ea8e2d2969d1d7a376f63d97ac58

Observation ab31082e-b61d-4101-a65d-f26a4078c97f · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:14.532847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:c180a5228c5a25421e3c3ba7389b918eb18302689a865e212ce80266683da0b5

Observation e9dd7262-0e5a-42be-91d3-07b0c2fb289d · inbound

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation cites this paper.

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:00.915838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T17:12:38.618233Z digest=sha256:92126b8f9cfb10a01e69b3d346911d8c7cb294491f7bb091ee239203c46f039e

Observation 58593182-a8bb-40e1-84b2-f315e94771f7 · inbound

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation cites this paper.

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T19:54:25.294067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:54:25.294067Z digest=sha256:8f3782138e794b0b2bb77676c1f3baeb214c0729e9f83ab933ebf162666fa1db

Observation 3f1319da-4fef-4b53-9cca-409ce23938c0 · inbound

Benchmarking Deflection and Hallucination in Large Vision-Language Models cites this paper.

Benchmarking Deflection and Hallucination in Large Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T10:31:00.678352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:28:26.268341Z digest=sha256:17f910f3de2af617619f8602d3130d0f44c2947e3ef03a47823a45809637b907

Observation 1e3b25e8-7e0a-4d90-8e9a-b9b75bd515fe · inbound

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models cites this paper.

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.682843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:16:51.109113Z digest=sha256:0d58c54c473f7d48226c051d408c0ec3bb429edf53c6bb83c414b613199739cd

Observation f31a42a9-3475-4e02-aa22-a9e3b35373a2 · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:49.461872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:7bf2df0809831b6e0dcc51d96bea162ddcfb5f7c2825d4ab6ff60afeef1085d7

Observation 5ec127f8-c038-4205-9379-8c004079cb78 · inbound

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration cites this paper.

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:26:02.222190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T21:55:35.699057Z digest=sha256:914a7559653b166494d0c906bea022e672cffefdff180cbe304f28ec14460958

Observation 1219347e-580a-4b29-802e-2749de453990 · inbound

Sensorimotor World Models: Perception for Action via Inverse Dynamics cites this paper.

Sensorimotor World Models: Perception for Action via Inverse Dynamics EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:30.299865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T18:21:51.580903Z digest=sha256:38ca23874f32aece4b3ddb77a8d2c38b5f161218e4bc5874f212351a448595ff

Observation f7e01b75-078c-4ca5-8725-73c52626a467 · inbound

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval cites this paper.

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:49:36.655493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T15:34:55.062016Z digest=sha256:0fe2a0076a16c8688712d2df87e4d82955d803b0e153fc5c40b77e66650a3d9f

Observation 7bafcb59-91f1-40f5-bec6-abf94911e215 · inbound

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification cites this paper.

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:59:46.881141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:12:14.829556Z digest=sha256:89306f4c72d20bedd1f91be56c239e1cd3ce8a35ec131cd82dfa5d61e4ad6e59

Observation 3c2d5fa3-575a-4d5f-a188-1c1710b824b6 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.784830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:498bc604dea88134284b62e840df3d9406154e346a9ac112f5007fa4fc118f40

Observation 38118793-1777-41f9-b8f6-6a1acc7ce684 · inbound

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting cites this paper.

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:17:17.749685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T19:13:44.015782Z digest=sha256:6437a5ccb2cd5c9d3d688e07fc83764ca0383ee11512d6f376a0fd13da62eb1d

Observation 851cac4c-9e90-4435-b87d-30ba397117c2 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:17.047603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:17.047603Z digest=sha256:a26fd3da14ed72f21a2603f4c67c7edbd50b5238b0af0ddb1caa7ca5e2bb2f56

Observation 353c7071-a8c2-4dbc-9acf-c8a45a9ce76c · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:52.837353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:52.837353Z digest=sha256:d26640589b32b490c35b27cb9f1fedcd9c76bf61407a9746a1568adc6589fcb2