Pith. sign in

Paper Citation Record · LEDGER

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2402.04252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04252 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:18:21.824481Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 604dbd24-601d-4acc-8380-8d19cf97cc55 · inbound

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering cites this paper.

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T12:52:30.936377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:52:30.936377Z digest=sha256:f93be955fe4f6653d010993ef8351a44768a6ef27eac9df9dc5ae6dc04bbec35

Observation 7ddc42c3-c86a-4170-bcf8-4bac5e3c75cd · inbound

HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator cites this paper.

HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:25:18.540134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:25:18.540134Z digest=sha256:bc813092b12cea7021238b881a0f7e42b658af8b113f11ddf5b8f18331271a70

Observation fea92edb-8c75-4ce7-bf16-cfa6668d7746 · inbound

VladVA: Discriminative Fine-tuning of LVLMs cites this paper.

VladVA: Discriminative Fine-tuning of LVLMs EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:26.482104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:26.482104Z digest=sha256:430e3fac735ef7d3d723d553a08ebbc472824d741b2fea47104ebac85abd100f

Observation d6b354a4-655a-47d7-95f2-dbe1febb5d04 · inbound

VCA: Video Curious Agent for Long Video Understanding cites this paper.

VCA: Video Curious Agent for Long Video Understanding EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.266437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.266437Z digest=sha256:f5ac90ada80777a9c5991228b928575c0fd988bb22975122a8864ce5f5853551

Observation b522667c-c3bc-4f99-85bb-2221c3497082 · inbound

DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment cites this paper.

DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T10:45:04.412388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:45:04.412388Z digest=sha256:129887b1700dfc790db26628159f620245e51fd30e7e691f465a27cd74403974

Observation 05160980-7f19-460b-92ea-469bcdb45e34 · inbound

Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation cites this paper.

Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:25:41.992024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:25:41.992024Z digest=sha256:d70bda7f0cea61c9bbfe4d3759c7b9c76009b66658c87845357d9fcf5e1f103e

Observation 0f91f3df-ccea-4e11-afc9-3cfe486b66ee · inbound

Towards Conscious Service Robots cites this paper.

Towards Conscious Service Robots EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-10T14:33:14.601690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:33:14.601690Z digest=sha256:caf216a63457f4f43c3dcca633d7933bac514a9d850937ca16c79f1b3709c45e

Observation 6c137259-cc76-4756-84a7-bcbf761cb4c4 · inbound

Understanding Long Videos via LLM-Powered Entity Relation Graphs cites this paper.

Understanding Long Videos via LLM-Powered Entity Relation Graphs EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T13:53:26.637867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:53:26.637867Z digest=sha256:4db1ed7ac55e8e59297846292dece5ebd5fb78d500ddd522c458d55e493ced95

Observation 95ad16bf-1e74-47a6-9bd8-d43c48855210 · inbound

LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models cites this paper.

LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T00:12:36.489685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:12:36.489685Z digest=sha256:b8dd3f79c48457ae62e0e86de6c0f9835d69dd2d430bd303c188119b7cd09a30

Observation 6bf53bbe-ea34-459b-95a4-8b2c145abedf · inbound

Conformal Predictions for Human Action Recognition with Vision-Language Models cites this paper.

Conformal Predictions for Human Action Recognition with Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.636325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.636325Z digest=sha256:7b4557f3f99df59966dd874a178494bf8fa69c68cab6ba837a47aa9caa17452f

Observation 20477537-a323-4c31-956e-b3057dc27188 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.850157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.850157Z digest=sha256:305f6e6a586216f20cd99fef50a1bc50b93f7b53d48551745d7bb732630c6747

Observation 761d635e-cc0e-4563-abd0-2d153a7b3767 · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:21:15.804400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:fe25815332ebe878d342e776add5e65ab719ed1d79b690ea12f9c7b9ef38ab22

Observation 33163676-9942-43da-96e7-7883f6a5b377 · inbound

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM cites this paper.

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:41.212978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:41.212978Z digest=sha256:7f2530bd6b5255d6818a1cb2ca52192030c521b208c525241c4ecf41384c88eb

Observation 6141965c-c4e6-409f-9782-80e21e49fed8 · inbound

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation cites this paper.

mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:26.469919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:42:26.469919Z digest=sha256:35cf12325842a6eb913f8ac3c270490655127505e9841c484a094435601329d0

Observation 4ac52ac6-a3b7-487b-a43c-42578b84fa9e · inbound

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying cites this paper.

Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:18:47.669708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:18:47.669708Z digest=sha256:c85f9fa9b0a7652a36f1a2dd6892c39065e76713fad09141db16cfebe9acc276

Observation b287334b-6846-47e4-9a02-02e8a143c3d6 · inbound

Can Argus Judge Them All? Comparing VLMs Across Domains cites this paper.

Can Argus Judge Them All? Comparing VLMs Across Domains EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:52.199619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:52.199619Z digest=sha256:ec4a67995037c48b662a03ae5d620139c47863ce49af6a9bb7a02e1624ec5c3f

Observation 4d1c2b4e-a74b-4f70-8c1c-f0b2b2932a61 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.749785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.749785Z digest=sha256:1e929b0e04c91e5b815f8a85eae0a027f95b6395e4f0cfc5d2b457be837924db

Observation 1e092c2f-2333-4d1d-849b-1b18121ed4e0 · inbound

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments cites this paper.

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:55:45.406624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:55:45.406624Z digest=sha256:0311acdc1591029e581d4698434eb984cb3aa3ff6e1e1c4f7a936cccaf2eb6b6

Observation 5006196a-e851-43d4-b838-bb79c4440a6c · inbound

Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders cites this paper.

Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T13:59:44.047838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:59:44.047838Z digest=sha256:7df735443043c84fc034b6f536f858a4746c5cd6b5e441a73f54791c9970bdb7

Observation a5a8c8bd-7223-4a3f-a987-fd10e00d65c9 · inbound

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents cites this paper.

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:41:22.579211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T12:40:19.544260Z digest=sha256:6ae54ee0cdc91a3deb8b3b6a6866f1b84302105feace3c3b750e931ed37e57cc

Observation c74decde-c081-47ff-baf8-5c04afc10656 · inbound

QKVQA: Question-Focused Filtering for Knowledge-based VQA cites this paper.

QKVQA: Question-Focused Filtering for Knowledge-based VQA EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:54.033143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T12:53:59.668663Z digest=sha256:f3c370b8bf8e68b2adc5e6f738530064158afee7cab0cefb1998fc380f30dd7c

Observation 24b1a7b5-fabe-48c9-99c0-5ccc3d4e1fe9 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:57.341823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:57.341823Z digest=sha256:ebf96c25f4df2ba7059eb601ea26066be96a5af0a47f18e4f6996ca5b4a3ab02

Observation 210dbf3f-69e6-41c3-ba71-1ed8a84f9b3b · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:02.006151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T11:09:51.816554Z digest=sha256:7af9714dd89228ccb5feaee6ecdf14b42c36149f3646ac4ac77c77da1febd573

Observation a43f5ecc-af12-4c0e-b663-17de2f84e7cd · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T22:43:59.727378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:43:59.727378Z digest=sha256:d49e6696eacdd4cd1fb44fdaa36e448a9d1103d3f1da2c00e057fc84a3a7c5f6

Observation ab31082e-b61d-4101-a65d-f26a4078c97f · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:14.532847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:6b7c7bc16cf3a8d6680a0f4fc20e603f5da8eb877595e64804020c9e35a3c378

Observation e9dd7262-0e5a-42be-91d3-07b0c2fb289d · inbound

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation cites this paper.

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:00.915838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T17:12:38.618233Z digest=sha256:7d3e166c402f51c3db890c3d4aae84f2ae6d772c50857d740a0c725776431a40

Observation 58593182-a8bb-40e1-84b2-f315e94771f7 · inbound

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation cites this paper.

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-14T19:54:25.294067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:54:25.294067Z digest=sha256:cafa3ca9c732416ba9c3c02fc0c86d236907781e08b7b5285178c8b4fec5850e

Observation 3f1319da-4fef-4b53-9cca-409ce23938c0 · inbound

Benchmarking Deflection and Hallucination in Large Vision-Language Models cites this paper.

Benchmarking Deflection and Hallucination in Large Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T10:31:00.678352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:28:26.268341Z digest=sha256:f44b888a4b7ce6f714cc65345d8ed9e86121628d51ff8ab862cc2ecfd6e2ac93

Observation 1e3b25e8-7e0a-4d90-8e9a-b9b75bd515fe · inbound

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models cites this paper.

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.682843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T16:16:51.109113Z digest=sha256:2fe8a4447f73d5b90c9a2e266eefe0eeed6a6fc0d17b6e7a00e4100307ef1a62

Observation f31a42a9-3475-4e02-aa22-a9e3b35373a2 · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:49.461872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:d85f651df3c79ec3520dc830432efd97119cef465c2f314eb3106be8eb58de0e

Observation 5ec127f8-c038-4205-9379-8c004079cb78 · inbound

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration cites this paper.

HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:26:02.222190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T21:55:35.699057Z digest=sha256:764912066d6eb26d84cacd675e566e3cdd7201be672fc289ec5260fe22574405

Observation 1219347e-580a-4b29-802e-2749de453990 · inbound

Sensorimotor World Models: Perception for Action via Inverse Dynamics cites this paper.

Sensorimotor World Models: Perception for Action via Inverse Dynamics EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:30.299865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T18:21:51.580903Z digest=sha256:af14d7d2c320f912878abe0c9c3241aeaad158124b02d2683ca723c0593351ef

Observation f7e01b75-078c-4ca5-8725-73c52626a467 · inbound

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval cites this paper.

ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:49:36.655493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T15:34:55.062016Z digest=sha256:37c20e957d2ad1117979df1a0709d0c82fee5c361d222da06324881dc2e7c11f

Observation 7bafcb59-91f1-40f5-bec6-abf94911e215 · inbound

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification cites this paper.

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:59:46.881141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T08:12:14.829556Z digest=sha256:91b8ca71c6d74013ad80c008ed6f4b4ce141d4471b943f67e8383a16ee0a23c0

Observation 3c2d5fa3-575a-4d5f-a188-1c1710b824b6 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.784830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:4455c2771ff6398fbb090dd45a625eca5f0693ebcfa107eb5667f1629cac525a

Observation 38118793-1777-41f9-b8f6-6a1acc7ce684 · inbound

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting cites this paper.

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:17:17.749685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T19:13:44.015782Z digest=sha256:c5bcc9c3924132d3eb09ef989e7de2b60343ab382b18750175a78ad1c6fbafc5

Observation 851cac4c-9e90-4435-b87d-30ba397117c2 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:17.047603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:17.047603Z digest=sha256:f67d54f8d99d86327c11e5f7c7447988dde8b43669050bf9cfc84b713445b55d

Observation 353c7071-a8c2-4dbc-9acf-c8a45a9ce76c · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:52.837353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:52.837353Z digest=sha256:7e360771f2870e844eb48e1336ad15c58b0a28b385bd45099444a26de9b1b991

Observation 73156e81-089b-42ab-81a5-5f4185047780 · inbound

LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection cites this paper.

LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T13:46:27.714417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:46:27.714417Z digest=sha256:bcc639894c508717038b673c99328b9e8c72af010aaa718909f85877a8ce5dbf

Observation b9564f85-1ea9-44da-887d-e3491ad51c66 · inbound

LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection cites this paper.

LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:21.824481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:21.824481Z digest=sha256:1d14434a5054d475cf4f9fd14610765f607190dc65c303baa6f5a23346b480c8