Pith. sign in

Paper Citation Record · LEDGER

EVA-CLIP: Improved Training Techniques for CLIP at Scale

As of 5 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 100 inbound Pith citation observations for arXiv:2303.15389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.15389 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T01:54:21.943160Z

measured 153 of 153 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 140 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:01:21.903186Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact14
  • verified fuzzy38
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

80
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5d4ac27b-e57c-4063-999b-dfd616550206 · outbound

This paper cites https://laion.ai/blog/giant-openclip/.

EVA-CLIP: Improved Training Techniques for CLIP at Scale https://laion.ai/blog/giant-openclip/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.129995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:9b4a757cff618396c10b4842429bc24f57f6d3e48a6f3385f3901245996286f2

Observation 905cb502-ae4f-4887-ac92-39bcffaf06ff · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

EVA-CLIP: Improved Training Techniques for CLIP at Scale BEiT: BERT Pre-Training of Image Transformers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:50:11.580126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:20a5960b664bbdae4c257f26d522d47bbc16b7a79d999442d4e3ea5afdcaa617

Observation 47538e7b-3219-4811-a63f-f13c3d2dd47c · outbound

This paper cites Ob- jectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Ob- jectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.102541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:e95d14690feb8b94491a51a7629c6f8a1df5e3b62698c787542eee8156f717cc

Observation 84c86ebd-289e-49b5-a5f1-2bc66792c690 · outbound

This paper cites Birdsnap: Large- 5 config EV A-01-CLIP-g / EV A-02-CLIP-g+ image enc.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Birdsnap: Large- 5 config EV A-01-CLIP-g / EV A-02-CLIP-g+ image enc

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.105212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:d83e71c91022581ef0e7f516d9a0e0bb40cfecc340f6e84cba9bb2bd8302d079

Observation 3be0339e-3f28-4a58-95f6-c2d4a21eb001 · outbound

This paper cites Food- 101–mining discriminative components with random forests.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Food- 101–mining discriminative components with random forests

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.108178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:eda178f7d5f8ee1fb2d46edf9a5ff937ae2a274b8a5ab5f4c82a96b38982d743

Observation 3c713905-633c-42f9-9547-24cb2b680173 · outbound

This paper cites Coyo-700m: Image- text pair dataset.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Coyo-700m: Image- text pair dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.110684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:4106c4b0b24c8f1d7f2dd06ebd3e38fb8ecb0f29140453b4e4b26d29e63f60f5

Observation 9227a8ef-d6a0-4d3c-9757-84e1b04262e7 · outbound

This paper cites A Short Note about Kinetics-600.

EVA-CLIP: Improved Training Techniques for CLIP at Scale A Short Note about Kinetics-600

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:21.979398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:baab1554b6fd7e21a7fa6c97d7623bf18b6bc491ad90e7dd52e31b784345cb6f

Observation 40df6e5f-c463-4583-b6c4-1466c3d0c92a · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

EVA-CLIP: Improved Training Techniques for CLIP at Scale A Short Note on the Kinetics-700 Human Action Dataset

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:21.983170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:27e112e115d7d15c89bfc59e8588614b673603c9c4cd3019c4dae5852bedbb75

Observation 3847e30a-d5d1-4bb2-803f-4f612dcd4e8a · outbound

This paper cites Quo vadis, action recogni- tion? a new model and the kinetics dataset.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Quo vadis, action recogni- tion? a new model and the kinetics dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.119266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:659d9a01fc6436f84c518539003eb773ea38cb05505fb9f3593cfdecfcc2928c

Observation c06272bf-e3a1-4305-bfd0-89ea29a742b7 · outbound

This paper cites Train- ing deep nets with sublinear memory cost.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Train- ing deep nets with sublinear memory cost

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.121938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:c0bf787fabd7626b62ca80f78e40dcc0fb776cce4504390ef17d442971a2b513

Observation 3e0c82d9-1ca8-4a33-a341-5cf5fb20c7b0 · outbound

This paper cites Remote sensing im- age scene classification: Benchmark and state of the art.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Remote sensing im- age scene classification: Benchmark and state of the art

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.124613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:c09c6645ec10663c4419b04b80d91d35e90979ef214c7a4fb2b92977158e9595

Observation 2f2bd24a-7cf9-41b1-84e6-3b6d0e2f7775 · outbound

This paper cites Cimpoi, S.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Cimpoi, S

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.127191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:212f2bdbba69b7606e2416d2329c92ec063acf3f3f9e25e5ce7a5823f0532a01

Observation 8929ef37-d1b6-4a04-8b9b-f0402469abd1 · outbound

This paper cites ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators.

EVA-CLIP: Improved Training Techniques for CLIP at Scale ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:26:47.677366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:b8613e0b60a7b0b9f9bc2988ef8feb8fb55d7f4d1cd3418014ca09375b06e538

Observation 717a9c38-ba92-48d0-86f0-8dc0f1e20d74 · outbound

This paper cites An analysis of single- layer networks in unsupervised feature learning.

EVA-CLIP: Improved Training Techniques for CLIP at Scale An analysis of single- layer networks in unsupervised feature learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.132329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:5b1b9ad6257b63ad12e0603baa3cfc8ab6f5ebcc350327bbff03f9d5a42149a7

Observation d8d4586e-eca9-4adb-857f-349b281b601d · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.028248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:a7616d2888e4aec0336695ac14a87340d280c18837d7f850d4712d5dc48071a9

Observation 6cfbdbf5-7aa9-4573-8694-628b8893427a · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Imagenet: A large-scale hierarchical image database

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.030998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:b298d057ea4135a86347fb2ff5de76f14d49b0d0e1f6210baaaa4ce01e043d16

Observation 72644616-38fc-45fd-bedc-7a895f2b1e88 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

EVA-CLIP: Improved Training Techniques for CLIP at Scale An image is worth 16x16 words: Transformers for image recognition at scale

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.033451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:50efc4572cbbccc47dde106d4d294c1a2de6899b4a555f5ec0d0533a0b686003

Observation c77d9479-4b93-49f0-8659-f4903cc3dc19 · outbound

This paper cites The pascal visual object classes challenge: A retrospective.

EVA-CLIP: Improved Training Techniques for CLIP at Scale The pascal visual object classes challenge: A retrospective

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.035899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:df7b4de5e778888ca60a208cbe9db20440be05609d49c3ababe8b1f1a12e15d3

Observation e6073abb-c3e6-41e7-9142-55638d434202 · outbound

This paper cites EVA-02: A Visual Representation for Neon Genesis.

EVA-CLIP: Improved Training Techniques for CLIP at Scale EVA-02: A Visual Representation for Neon Genesis

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:21.987206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:df67205f15e6d3de11bf1e3c1ff1bff05008e7ac11118d2cce6fd1e74f14c1f0

Observation 4b240e77-6838-4723-a1a2-7770350f0f94 · outbound

This paper cites EVA: Exploring the Limits of Masked Visual Representation Learning at Scale.

EVA-CLIP: Improved Training Techniques for CLIP at Scale EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:21.991369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:182658f7a766532c212778be75f4f4f3e423811fe56d89ac6e98c02d58d1dc51

Observation e20181b7-f29b-458d-8159-f674dc8cb0a1 · outbound

This paper cites Learning generative vi- sual models from few training examples: An incremental bayesian approach tested on 101 object categories.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Learning generative vi- sual models from few training examples: An incremental bayesian approach tested on 101 object categories

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.044325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:92090acd114f64ba1e642bf9ea62183039d1cf4a1b7a3000c08b4218239d6c19

Observation 0ce706aa-5ae4-4af5-8c40-090708f9b9a4 · outbound

This paper cites Challenges in repre- sentation learning: A report on three machine learning contests.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Challenges in repre- sentation learning: A report on three machine learning contests

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.047295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:92d44db0b1879dc778deafccb4c1ef4a95c8221d715075bbf0cc8fbc8446d554

Observation 7058e54c-3bd3-4e16-95ea-db0cbdf8c405 · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.IEEE J.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.IEEE J

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.050108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:71fe46499e03d8975d4d459b7ff3726b12e79c462e1ff092d9a020e3add9a210

Observation 4ed2b2a1-72cb-4008-bdaa-7c97f5bb18d2 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.

EVA-CLIP: Improved Training Techniques for CLIP at Scale The many faces of robustness: A critical analysis of out-of-distribution generalization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.052637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:cbe25d4af5cf0795d575152c0095aea978c4dff51156d1ba5fde0fc19c3bb7d9

Observation cc7ce154-af0c-4833-b404-32e89e3ff53c · outbound

This paper cites Natural adversarial examples.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Natural adversarial examples

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.055045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:fe97ea87b5d85a560e910c43f072b1a7e963dbeadd101c29ac60bfd41d98b036

Observation 68a7f5ce-8d31-450d-9300-2118792fe1cd · outbound

This paper cites Deep networks with stochastic depth.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Deep networks with stochastic depth

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.057627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:6a7d530ca5f3ae3ac66040b67017869179a8a3ad6385f48901a3012729b658b3

Observation fc3cdfbf-a4a1-43fc-b352-bc6849e14b7c · outbound

This paper cites Openclip.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Openclip

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.059793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:fdd2876ceb8016f9469e533e4a00cc5fa6e4ab55bc92cc13c195bc011867747b

Observation b3fc2eed-9b09-4827-854a-056c7e6415a1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Adam: A Method for Stochastic Optimization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:54:21.994898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:e55db949808f5bd6b5dfe051c12587dbbbc3a0033d7954d3a4b1c1982a5aabe7

Observation 4429f131-efeb-4bf8-8c68-07cd4f61083c · outbound

This paper cites 3d object representations for fine-grained categorization.

EVA-CLIP: Improved Training Techniques for CLIP at Scale 3d object representations for fine-grained categorization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.065627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:3fee8ba57ae917f567138616dd558b1aaf50712affbfa99dd5bc7964f94e7a67

Observation 9a2d9850-3d62-46b4-8589-03158af307fd · outbound

This paper cites Learning multiple layers of features from tiny images.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Learning multiple layers of features from tiny images

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.068005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:a401979d2cf8bf7b7c3b4d1885c2eee827e37904937f3eb1d6fd528f84e99f57

Observation 86370879-fa09-42fe-9550-642faf874be0 · outbound

This paper cites Gradient-based learning applied to document recognition.Proceed- ings of the IEEE.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Gradient-based learning applied to document recognition.Proceed- ings of the IEEE

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.070552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:03c16822aadbe96b5629c5aaa07bd0fd024727bd6f52d989064f95f6fa8c03f6

Observation 4157f1f2-ad2d-4c38-8bcc-eda18a0fd7fd · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

EVA-CLIP: Improved Training Techniques for CLIP at Scale BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:54:21.999367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:64b31aa56dbeb345b66da6701e11f5aa6a1016b62660572445d3ae865ba0a4c0

Observation 7fa97d15-8688-401f-b5b5-20dcf49962f8 · outbound

This paper cites Scaling language-image pre-training via masking.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Scaling language-image pre-training via masking

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.075790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:e8ca33dced63d3eb11c3f8e8f942442d69db41d77e62178e61f66a105295d7c4

Observation dc571e50-a4b3-46d6-8ac1-2cc18354ea40 · outbound

This paper cites Mi- crosoft coco: Common objects in context.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Mi- crosoft coco: Common objects in context

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.078251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:328e166c3db3d320d33639f446b1813fab3055f88fd8f125b8c0d1a02e9354d6

Observation 4dfa8e20-811f-45cd-92d3-6e833563521d · outbound

This paper cites Decoupled weight decay regu- larization.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Decoupled weight decay regu- larization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.081332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:37787325206631080ce71962223c0cb3188fd2b8da9cc57beaccf294028b52f2

Observation 585dcffc-9241-46b2-b18f-85bb2e2024be · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Fine-Grained Visual Classification of Aircraft

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:54:22.003522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:735833271dafb3071a58fd8ca96ebb9910a984cdabdb7bc1ed7c2c11e9410ce3

Observation 1717ca60-b062-4b7f-8ee2-58923a460f94 · outbound

This paper cites Automated flower classification over a large number of classes.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Automated flower classification over a large number of classes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.086632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:ef400d6b9350e0071b0634857b33340cb079c6009cddcd211bd11bbc41aec6e8

Observation 4f03c5b8-54cc-4747-acdf-1c9d3e18005b · outbound

This paper cites Parkhi, Andrea Vedaldi, Andrew Zisserman, and C.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Parkhi, Andrea Vedaldi, Andrew Zisserman, and C

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.089019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:9eea21996987d6028ee62a26a98f90edc565de2062c2fa6099c0bbee4164d617

Observation e8d43270-fa26-4155-9939-2e343b7b816b · outbound

This paper cites Learning transferable visual models from natural language supervision.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Learning transferable visual models from natural language supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.093328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:35536aa08d8a3e763e72475b925c636d867c8fc1a73870afbd5822f4bb4a2564

Observation 5e2d96f6-9869-40b7-9a61-86d87ea4d916 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Zero: Memory optimizations toward training trillion parameter models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.096508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:12eda36766c0860b3b1b8681661c3b7665807837e58a713dded4a9c3229b9493

Observation 38a04513-429b-44bd-8d97-2e2d7bda9b58 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:54:22.007249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:262d798afc287f83e01d4931bc1bc0c0639dc6b8afec0f6c48ef21be69b6e9e9

Observation 8197d098-652e-44a6-945c-8404abe851bb · outbound

This paper cites Zero-Shot Text-to-Image Generation.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Zero-Shot Text-to-Image Generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:26:09.277502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:f2af47fe6d439964156bcbb75baa61a9ff77dc37d532d065f5b2dde714e7c1cc

Observation e833415e-e83a-40c0-9403-c017d844f0b1 · outbound

This paper cites Deepspeed: System optimizations enable training deep learn- ing models with over 100 billion parameters.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Deepspeed: System optimizations enable training deep learn- ing models with over 100 billion parameters

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.116495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:d632b675c1f8a20d1f8b00a88287389c2b1892dee739f0b6a1981add9d1282b6

Observation 24cbce26-2a3c-4a63-8a39-b3c00860d10f · outbound

This paper cites Do imagenet classifiers generalize to imagenet?.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Do imagenet classifiers generalize to imagenet?

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.038407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:c87b77d3296d4693ebbc309f43134c1e6e0b7fbefc97d1f7df501511c7cacff4

Observation a5ecfb2e-8233-4f49-be7d-837cd80cfe14 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

EVA-CLIP: Improved Training Techniques for CLIP at Scale LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:22:17.864689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:da12e373b63695980f00ab0ac7a5d0f9df71fd590b5fe4d7683eaa7423558a8d

Observation beddd684-84b6-4f4c-a0b7-0a12b308952f · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

EVA-CLIP: Improved Training Techniques for CLIP at Scale LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:54:22.020140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:e1719106b0a75bb4c436131dc223235fb71d91fa80bad14d29cbee15e1669c2b

Observation 84b45e81-4f67-413c-9e81-f7a43c26b1b4 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

EVA-CLIP: Improved Training Techniques for CLIP at Scale UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:54:22.024159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:f3fb4af79064df6a44c3c673d7c8fc83a84aaccea56e4f07324427f1d23b7a9f

Observation 0aac5f7d-690a-4287-87f5-ec929279514f · outbound

This paper cites an unresolved cited work.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-13T01:54:22.083782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:b9cb32950780bebe73654b6434bdebf56d6d76ef2e6ba0b7a85a841225e553c0

Observation 7d9c10ab-a031-40cc-8064-8973b00ecbac · outbound

This paper cites Rotation equivariant cnns for digital pathology.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Rotation equivariant cnns for digital pathology

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.099465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:287daeb87b9c083c71e4d435c85874f97de24f49a5cbd81be6b0a73737e03066

Observation 93db47bd-4d8a-4d1c-a88a-11671c53c8c5 · outbound

This paper cites Learning robust global representations by penalizing local predic- tive power.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Learning robust global representations by penalizing local predic- tive power

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.113561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:816eaa4f2adf9874318b7d09d5d99e093b99a68831336986686bddeb9b3eb8a4

Observation 3e5eb02c-88a3-4689-bc21-19da9493582b · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Sun database: Large-scale scene recognition from abbey to zoo

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.041452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:aabeb8117fea7fc740718e93f48f56a6744415c1cf36ad28379cca2b281e1707

Observation fbc6978b-0bef-4c9e-8c7e-b508f0ff55e1 · outbound

This paper cites Large batch optimization for deep learning: Training bert in 76 minutes.

EVA-CLIP: Improved Training Techniques for CLIP at Scale Large batch optimization for deep learning: Training bert in 76 minutes

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.062575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:b0299b04040871f9f79cd4c3941abe77d17085a7b03fb556af797c9812ee9228

Observation 094735e7-f008-43ea-90a5-b0ef0af9bcb5 · outbound

This paper cites From image descriptions to visual denotations: New similarity met- rics for semantic inference over event descriptions.

EVA-CLIP: Improved Training Techniques for CLIP at Scale From image descriptions to visual denotations: New similarity met- rics for semantic inference over event descriptions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:54:22.073193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:54:21.943160Z digest=sha256:651fd78d567487d1fd3da56d971d02a1b4ef33bd02c5fd9e111a7a1f782439a1

Pith citing papers

Observation ae2f0075-525f-4e22-9d3a-765eb44e0ea0 · inbound

Sigmoid Loss for Language Image Pre-Training cites this paper.

Sigmoid Loss for Language Image Pre-Training EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:05:36.495674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T13:05:36.460932Z digest=sha256:89db6d7cd67ef86485418034c0d184239b4626a04d0794e3d6827120c9fa4cda

Observation e12117b1-6c89-4e1c-a853-cc27b8436508 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.627028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:7fbd502f7877a3109dee72a52999056edc68b06d560d204522fa7fffa392ab12

Observation b02c8501-d249-4816-abf6-0acedb722186 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:56:42.459329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:57997fac78cb0cb90950e81f10238d9e1d5f886dc8e38342faf1b08b10237cb2

Observation d75a59a5-7492-4341-8443-e0ed414b8f02 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:30:22.717922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:c898fcdb44f1f0894291047da396e66161307578fc43fb9d9ee8a3f80e00ec6c

Observation 0c9fbff2-1c1d-46ed-8d07-0a09f3a29daf · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:35.139025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:8e7dd1985bb098546dfe9a24559ed626c1a1b691016b22dd60716e1f68a1ebd3

Observation 9f169a1d-6986-4a9a-849f-96dbd15aee9b · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 131

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:46:10.115832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:7bdaf30e01cd6d41625ddfdf4c8d3de1f60f3eca5faa3d8a26aa5fcf75c56454

Observation 0dafeb92-717b-4721-97ee-d5280200245b · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.118073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:4891cca0d845e261d4d56561dd7bcbf64906c31b7813e990c273d3772f2ec188

Observation 40e29025-82a3-4831-9d38-486880913f7b · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 103

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T14:25:59.311906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:b78c114e5ce5aec55e842739bf6fba0bec0318601caff54cd048614226aea343

Observation 8e50dc18-bbd7-4efc-bb83-1dbe5e3c12a5 · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:55:20.518354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:8801943d42b622dd5f0750e630ee1e4b2ab155990556f9a968bffb4bfd475c4b

Observation 221204e8-98ad-4dc9-9854-48309924cf39 · inbound

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition cites this paper.

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T03:33:50.836780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T03:31:58.848180Z digest=sha256:d8b2fa165010fbc0aff3e9370ae9e4abd70c7a6ee52a35a55790e0929000cae4

Observation b8d17ed1-6576-4d50-ad15-e05a10d25bd1 · inbound

BLINK: Multimodal Large Language Models Can See but Not Perceive cites this paper.

BLINK: Multimodal Large Language Models Can See but Not Perceive EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.564212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:440afdcfd884fce3feb02ae72c8dc09c53bf8ecb1233d33684c793923413a5c0

Observation 4c174f23-7ae4-45ae-9805-5c7e612debb2 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:55:30.107505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:5a3fe61c2c20bf3018f4ee11a12369e6b6b4f9011c0de08166f1b07c6a5526d3

Observation 49d8516d-9262-4c1c-a24f-9aa1bdd97fa8 · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:09:30.416809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:37f4ea7a426b9cd14bc7af7a912004b711a13302b08119cd2a8332452e480f2f

Observation dc55ca00-0a0b-4316-826b-277fe5ea1a14 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 123

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:05:03.704852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:6ecbea57db9c5dadd6f34cc2d79d9db6c4029534fe52400a915452d1a5addda3

Observation 01cd2b6d-5285-4c09-9eb0-122aa8847b25 · inbound

E5-V: Universal Embeddings with Multimodal Large Language Models cites this paper.

E5-V: Universal Embeddings with Multimodal Large Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:52:20.993088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T22:52:20.935555Z digest=sha256:64a7cdf885e0dc66a61e8e7c11b133f6123832bfea7d8c5936d2d1b13ad2a037

Observation a75f5f2b-8ca2-4d28-8161-68348ad4b46a · inbound

CogVLM2: Visual Language Models for Image and Video Understanding cites this paper.

CogVLM2: Visual Language Models for Image and Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:10:27.731179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T20:10:27.633010Z digest=sha256:da5060f4f80bb810b46cbf856a12f903d70953936aa1085e690a170e1ce75bde

Observation b9d723cf-4dce-4df9-ae41-a22f7e88f4af · inbound

VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks cites this paper.

VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:19:44.022275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:19:43.882232Z digest=sha256:bd2e1250e4544d784b4dca5508c5d159f2cae01ad3cc8087c9bb1522c10f3340

Observation f3c40a80-aaae-4644-a7e6-9cbeaecad50a · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:09:16.267045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:42ac75cd2c8d72c267ae9208f587b730245a5b3eefd124d0eac7ee90ac011cce

Observation c158e121-1f3b-4453-ba3a-598fe83fb317 · inbound

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos cites this paper.

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:12:43.893590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T08:08:01.889675Z digest=sha256:993424424c471553384b33c401cdff2048355cd8c4fd48af630db2f6d51a6139

Observation c211992d-0b7a-4c40-83ca-94bf2190d47a · inbound

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks cites this paper.

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:51:36.369148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:51:36.137985Z digest=sha256:34b54beb9b52abdc88c3a4fb09116b20686288b5d6e8c6bfecff952e9052c4b6

Observation d4b1ba38-dc02-4bf6-9bf2-3ea499ab52d6 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 187

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T07:51:13.391038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:b1d55858176e43405d3969aaf8aec116477c7fb1b370bfcadc1a9373b416524a

Observation 2f950d82-b38b-482a-b3af-a31985f2dcf6 · inbound

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features cites this paper.

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:49:22.279848Z digest=sha256:ea31c188d2592b61a5adbb9a72977154975207f4c02ca75793c2ea90252e802f

Observation 5c7e599d-84db-411c-97b8-1fe0f5d5668f · inbound

Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP cites this paper.

Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:22:25.142222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T02:21:09.704462Z digest=sha256:e6c1ee04c71ced4c1ed82174cdae4b8f9f9b7a9ad5abcd0ea80ea6674931ba2b

Observation a5708f05-1ea0-4ae6-8afe-a53c1e48e17c · inbound

An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval cites this paper.

An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:02:13.620465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T22:59:36.542774Z digest=sha256:d15e7ef11a300d7754a54b731c79b95a8045ae7ca4e4e14841cff3715663a65e

Observation 0a0c0629-f13f-49fd-8b7a-bdcf99bef568 · inbound

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model cites this paper.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:6593d412dbe0baaeaff1f77552517a4f1aa60440766c9272ec1343507e3eb708

Observation b37da83f-52b4-4b70-b983-5f55f948a42b · inbound

Perception Encoder: The best visual embeddings are not at the output of the network cites this paper.

Perception Encoder: The best visual embeddings are not at the output of the network EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 129

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:21:15.801070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:21:15.681336Z digest=sha256:619f1d5828a33cec0b72ed1d9464f783a059ff5d497d75abda7e28e34fc13cf3

Observation 364b53bc-a8d8-4e6d-8713-91222701105e · inbound

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation cites this paper.

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:26:55.266321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:26:12.597756Z digest=sha256:b1f143365f24f8b2c89f1a643dc324ee0c9fb96974fded0a99bc95e6bb151d95

Observation 40473f12-a9fc-4851-a7b8-6e9de38e6e0f · inbound

Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning cites this paper.

Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:42:01.374084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T03:39:52.969100Z digest=sha256:7a8a743837849ccd7ba59803aaa942e1902d76c7eba1d45574a9901d67af9886

Observation affc2308-9aa7-48b9-a23b-d18a1657460c · inbound

SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs cites this paper.

SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T21:01:21.903186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:01:21.903186Z digest=sha256:0a13dbb90108eef0d8b12915aca0b8434b2e679e8c231ca23e2e2125c5bf6dc1

Observation e6a312fb-d024-40b9-b282-71b5beeada23 · inbound

Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models cites this paper.

Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T16:56:26.617118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:56:26.617118Z digest=sha256:39307ce8bf640fba055a77d7745b55b5e63c55d490bb4714eeb3efe23bcc4a5d

Observation 9c9abd3d-f647-47b3-830b-12f952e50b1c · inbound

MobileCLIP2: Improving Multi-Modal Reinforced Training cites this paper.

MobileCLIP2: Improving Multi-Modal Reinforced Training EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:59:19.599119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:59:19.599119Z digest=sha256:96f4d83d9381e2c54e828b8df0fa7459aa7bd4e0bc5b6cb94a76ed8c270d15b9

Observation 406f1b2e-2c25-44fd-b7c9-1c26f78867aa · inbound

Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding cites this paper.

Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 143

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T20:26:50.365543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T20:26:16.133860Z digest=sha256:903efb9e70054c8de75dcf9358fb1ee7c9cabfbb40648e0aa39ee5364b25eb36

Observation 6aceb296-ed2c-4b6d-90ac-699e9b6347f5 · inbound

Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering cites this paper.

Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:06:49.812306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T20:04:52.852253Z digest=sha256:07c61f1ebd10c38d8139c6c56cdb1bf7dd64d7e2b31387d7dd72d345d360b502

Observation 2d429898-6337-47ec-82bf-e4572a281f2c · inbound

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning cites this paper.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.731050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.731050Z digest=sha256:d9edfedb8ae4675c0c792a2664bde073485ffc5a48c9ad1021f629441cd61e87

Observation e627282e-42a3-4f9b-b6dc-ab948a7b2dd6 · inbound

EditIDv2: Editable ID Customization with Data-Lubricated ID Feature Integration for Text-to-Image Generation cites this paper.

EditIDv2: Editable ID Customization with Data-Lubricated ID Feature Integration for Text-to-Image Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T05:20:00.963413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:20:00.963413Z digest=sha256:172c3b561b6291d57f5c204366abccd9c4f69412b38a63ea84bcb8e8d72538bf

Observation 91419f0c-a081-47d8-a75b-c9473fdbcc9b · inbound

Recurrence Meets Transformers for Universal Multimodal Retrieval cites this paper.

Recurrence Meets Transformers for Universal Multimodal Retrieval EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.054038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.054038Z digest=sha256:8f0936e8632b65f2a1034172729a4fc97486c4e6c1f2b2c7dfe1aa95f9a46ac6

Observation 5f796c12-9c89-471f-8c58-1869dedd49f4 · inbound

FreeRet: MLLMs as Training-Free Retrievers cites this paper.

FreeRet: MLLMs as Training-Free Retrievers EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:01:23.521764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T13:00:31.952588Z digest=sha256:11f8440623506ad96d8c9f2885dc0d4382308581d47772883c54ed00db6e226e

Observation 917f7d0a-cc6d-4d83-a965-69ea11d1b269 · inbound

FreeRet: MLLMs as Training-Free Retrievers cites this paper.

FreeRet: MLLMs as Training-Free Retrievers EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:46.961197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:46.961197Z digest=sha256:8eadbabc5f67e6a6db5b96b7a0d8715aa95dad7f6b2b0bc1bde85642f3384edf

Observation 42f0c0bc-4e7b-4dbd-a477-6dc96c8ff0d6 · inbound

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model cites this paper.

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T10:18:43.621652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:18:43.621652Z digest=sha256:d941f839a9e1c6b58ed22c1776979f747083aaf2d48bd7c28f98fcc71afab588

Observation 309d3348-5a97-4c30-a655-16d62f2c23a9 · inbound

VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models cites this paper.

VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:22:24.001325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T05:22:05.125849Z digest=sha256:d22f3196c3e6f6704c644330fcd713ff362a436c612b69b9b61bf3b092ead328

Observation ae9baa82-be3a-43dd-8b0c-984158138234 · inbound

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds cites this paper.

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:08.392322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:08.392322Z digest=sha256:bd012f7e61990f63d5116ef5615534cec1ef25d734731f91ed480f582ea5a24a

Observation 454c9428-df48-4687-93bf-67bfdabbfa55 · inbound

Calibrated Multimodal Representation Learning with Missing Modalities cites this paper.

Calibrated Multimodal Representation Learning with Missing Modalities EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:10:22.645035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T22:08:07.217659Z digest=sha256:29cb7e34e7984268d9a027e69b6f8a15290c06e6c90cbabe36857755228ac7bb

Observation c4bd4a5a-8f24-4c82-b3c8-e7418fded368 · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:46.973924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:46.973924Z digest=sha256:a00c3c73344f15acef828ec58606a1bf62d7ce630a25755fbba1ad42bf47d450

Observation cdb73885-1c46-440f-ad65-7759c4bac2b1 · inbound

PowerCLIP: Powerset Alignment for Contrastive Pre-Training cites this paper.

PowerCLIP: Powerset Alignment for Contrastive Pre-Training EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:59:03.952297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:57:44.794342Z digest=sha256:aa1c5c96964af345e18a1014540925320246c598fcf67152eb34f9aaef65ff15

Observation ba39f140-8359-4390-8f64-4b034923e361 · inbound

CLIMP: Contrastive Language-Image Mamba Pretraining cites this paper.

CLIMP: Contrastive Language-Image Mamba Pretraining EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T11:22:14.718656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:22:14.718656Z digest=sha256:1deccd4d7c8cc8e201407c475089bc24fd30c8b1177f2b65816d8928eb5970d7

Observation 7b4f20f0-2d11-43ac-bc66-bc4774e5e052 · inbound

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation cites this paper.

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:47:45.501186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T10:46:17.411843Z digest=sha256:d9742a3323333143821b90cab3ccb60bb011c5f36b4aec7baaa5ef9e968e944f

Observation c5bcda11-180d-44ba-b6f2-16c95b643696 · inbound

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation cites this paper.

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T08:13:53.659640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:13:53.659640Z digest=sha256:c572fd955152122f6a29731c4e0308e8495e8e8a36a97bf1a20e58a8b5d6c4f1

Observation f17a7717-22af-45d5-9ee7-f604b22c5562 · inbound

CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining cites this paper.

CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:27:36.880704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T08:24:44.943709Z digest=sha256:5b10ae2d1f9bf13f69aefb5b0f72e943b69d8194cd06cc9b0b8c843fb346feaf

Observation 2ae307be-0700-4d1b-b25c-230300ffe8cb · inbound

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models cites this paper.

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:17:36.109866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T08:17:29.924860Z digest=sha256:23131aeac43f38daff18f130f9745650dbd063d2f19cc3663e40f5b5504d9702

Observation b76a954a-60de-4052-9867-d07859b42d8d · inbound

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models cites this paper.

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T05:34:29.156287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:34:29.156287Z digest=sha256:be44ef5c671b6c3aeba57a4a47d8887f91cb85b161bed0ce05f2cb12b8738546

Observation 18bf3ac8-75e0-4c52-9119-7bb7f850d2c3 · inbound

Mitigating Error Accumulation in Continuous Navigation via Memory-Augmented Kalman Filtering cites this paper.

Mitigating Error Accumulation in Continuous Navigation via Memory-Augmented Kalman Filtering EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:47:41.872584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T09:45:53.868748Z digest=sha256:e80d9fd6412a2107593b8e00ee04ecb6c7f9b91f013e75b51b115c4348f426a0

Observation 213a78c7-16c9-46c7-bafe-f8660ede7be5 · inbound

Vision Transformers Need More Than Registers cites this paper.

Vision Transformers Need More Than Registers EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:10:15.678117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:08:50.579190Z digest=sha256:b9afdef69c076cd849f86c9902b2b6061a5ebeac8df53ae1c6fad4ebb87674f3

Observation c703e467-541e-4475-8e23-6bbed5b26eeb · inbound

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum cites this paper.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:a90fcd7acdc393797cc16bee4638191a0dc34fd62ebb8099496ed8dd14eeab8f

Observation a1bafc80-6dde-46d0-a709-4f962c83870a · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-15T13:15:50.518716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:d4605f90b571db8c1abdfcb866af70fcfe43bb31c5c870c4a78a9f4c1b06b866

Observation a4dd4b93-5251-4c1a-839e-c3a7b5e5d2e1 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:253dd90a6c941a649b0ef40dd9e2aa9b107a0933a0cbe70db41c068eddbce351

Observation 4626e353-1038-4c0e-a310-470ee0b2c314 · inbound

Revisiting Model Stitching In the Foundation Model Era cites this paper.

Revisiting Model Stitching In the Foundation Model Era EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T22:19:38.973091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:19:38.973091Z digest=sha256:5b266ae0d478a59b68dbea835c06bbf3888404e48951f38ae6ae84114f797f57

Observation 25a2730e-4f28-4c42-ad70-a433ef7707ac · inbound

SteelDefectX: A Multi-Form Vision-Language Dataset and Benchmark for Steel Surface Defect Analysis cites this paper.

SteelDefectX: A Multi-Form Vision-Language Dataset and Benchmark for Steel Surface Defect Analysis EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:43:24.372856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:42:40.501053Z digest=sha256:73dfb1391b3e11816b7499fc1d8ecd9f15601d1747cb56d95d3a1cba0806ebef

Observation f6eb0dce-44dd-409f-ac3e-e5d9f962cbd4 · inbound

When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models cites this paper.

When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:19:28.032626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:19:04.468630Z digest=sha256:2badf1e4328a518d7cc8aa3d7ddcb4c12e1ee682ba1432688c6f9db2c910eb9a

Observation 99ad757c-a944-4b0c-8ff2-e602d0c9f040 · inbound

XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs cites this paper.

XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T05:42:17.220750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:42:17.220750Z digest=sha256:673ba2dcfa924eb642dab4a70227cbeb7cf33e35ce09c7ded0c485fa2984a5c9

Observation 5059d5a5-ed91-4b05-a0bb-bedd13ee31c4 · inbound

Video-Oasis: Rethinking Evaluation of Video Understanding cites this paper.

Video-Oasis: Rethinking Evaluation of Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T15:38:41.390945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:38:41.390945Z digest=sha256:975085c60688ca910971377711b201871af25c6956fdcde30fbac2cabdad2c60

Observation 51bcb73d-4e4a-49d5-aa19-4ffc23d77795 · inbound

RGB-Pointmap Pretraining for Unified 3D Scene Understanding cites this paper.

RGB-Pointmap Pretraining for Unified 3D Scene Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T21:23:17.889832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T21:19:49.421653Z digest=sha256:3c121ca4d03853de4f6a52651fe753f61ed4eeb7dbe093355d8ab7681183dfcc

Observation c74b8aad-87a8-4339-ac73-1a0c9e8d9472 · inbound

Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patch for Infrared Vision-Language Models cites this paper.

Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patch for Infrared Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:48:11.631135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T19:43:29.058335Z digest=sha256:289f44a99f45fddc0cb82cd34d92114278a71f2713fe8e1b958c4a89f14f8ce9

Observation 01c1ef13-9723-42db-817f-f4f29fca4bbd · inbound

Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patch for Infrared Vision-Language Models cites this paper.

Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patch for Infrared Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T16:52:02.612278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:52:02.612278Z digest=sha256:ef51379dcd2bc08cb4e3ada734e210e25883460ee2b7aa33dfa2c172b1ffb143

Observation 2596e2c1-cf21-440c-ba53-2ffff9b0dae0 · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 57

Resolution
malformed identifier
local_arxiv, observed 2026-05-13T20:33:17.151630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:766560c85ac19217dc996588346bba8712bd3d0a6a150063907b1bd985eab026

Observation c2268b87-e105-467f-8943-9fa41c8db29b · inbound

WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering cites this paper.

WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T19:59:10.657346Z digest=sha256:43699aaa51e48e27574102b79e1a0aa21460d2985649a237cd05055f53891a24

Observation 9a4b174d-15e8-4d43-9ff9-14b0352c9619 · inbound

OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance cites this paper.

OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:04:52.871199Z digest=sha256:efaf70486a8ca14944d5dfc762123a83af9b8efbecb34335beb46a3ccdd317a1

Observation 824bfa98-3039-43a3-bdf5-cdc6419f28ce · inbound

NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild cites this paper.

NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:42:52.084288Z digest=sha256:f2b210f8846b63590232432827477de75cdf7bb28bad5fcff158119bbc01611a

Observation f2ca36ba-bf71-4b66-99b2-967c260953f0 · inbound

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment cites this paper.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:5d0428a75bf6c5dcbd6c4e38a90a41e4c216c79b8e35ce3f45295f432a0b3364

Observation 6b520045-579d-42f9-a8ca-fe98c2fbca43 · inbound

Boosting Robust AIGI Detection with LoRA-based Pairwise Training cites this paper.

Boosting Robust AIGI Detection with LoRA-based Pairwise Training EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:56:11.099966Z digest=sha256:1a3bfccd25e2b793bcbed3cb5546faef71e766570e1c9eeaeb55d60861796c93

Observation e5545ba7-c333-444e-a0bb-55c4e3ec9a36 · inbound

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models cites this paper.

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:16:51.109113Z digest=sha256:4c46efaa3987703c74036dafb4e08f2a38c8554af49f9bac8763133861f775cb

Observation dfdbe15f-615a-485a-b799-a6975a1d15ff · inbound

Dual-Modality Anchor-Guided Filtering for Test-time Prompt Tuning cites this paper.

Dual-Modality Anchor-Guided Filtering for Test-time Prompt Tuning EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:11:42.935453Z digest=sha256:22fc4aed349b6910dd03f21200a169f74cf57b5422950436af9cc778ef6f00fc

Observation 9140e422-ebbd-4ccd-82a1-c3fed2d59e3c · inbound

Cross-Attentive Multiview Fusion of Vision-Language Embeddings cites this paper.

Cross-Attentive Multiview Fusion of Vision-Language Embeddings EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:56:08.616492Z digest=sha256:ddeec010113d4a3946b0e803da802d6957433cf86cd8b40d46b480f8e5e157f7

Observation 9726290c-7a77-4a8b-9ace-ed0bf7ed6695 · inbound

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks cites this paper.

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:58:28.202606Z digest=sha256:b2cf9ed54912edcf4aa61f287d653028ff4cb8d8940cdbc23ebeb36b88fb3368

Observation 5aaf68ef-9c46-437b-80c0-7e92ad3c1716 · inbound

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce cites this paper.

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T00:55:56.146885Z digest=sha256:6550b03e57a25a638bdcd0cb3340686642be6ac5640bd9b366931ca5d3f21c49

Observation cf904e8c-3cd1-404b-9a84-f73290ee5a9c · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:5bc6a87d8196fd110bb766efc2c655287c3ce29ba6b9189dac25cb7c817dab33

Observation b8bb0c24-34a2-4d37-8e17-fc7fd19f2d15 · inbound

MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment cites this paper.

MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T21:56:38.004300Z digest=sha256:91956e28f247e006027f1a34bd3581851b76965cd57269fd6e5c06226df96fc8

Observation 090bac04-f443-40d8-b78c-0b9efe46087e · inbound

Photonic Quantum Computing on Spin Memory Architecture with Tree-Encoded Fusion cites this paper.

Photonic Quantum Computing on Spin Memory Architecture with Tree-Encoded Fusion EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-05T01:20:37.312658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-05T01:15:04.990410Z digest=sha256:cf3512b2f4d7f832a57fdd09a56722e86ab4c9fa629cd7c4f5b891a021af2b54

Observation 21cd8b1a-969a-458e-b020-72f2f5d7dfe3 · inbound

Rethinking Cross-Domain Evaluation for Face Forgery Detection with Semantic Fine-grained Alignment and Mixture-of-Experts cites this paper.

Rethinking Cross-Domain Evaluation for Face Forgery Detection with Semantic Fine-grained Alignment and Mixture-of-Experts EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T21:46:18.489180Z digest=sha256:b9a174d9b938179c49dcaf8d7cba9ca0bd5750f53ca5d9ee21bf6e5434e1afc5

Observation e4614da5-a0ba-4aa0-8fe8-2fd7dd6a43f8 · inbound

BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering cites this paper.

BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T11:37:20.118862Z digest=sha256:05e5632721416023a2d645e257129ecd13d745fa99404fc0440873acbc342586

Observation 0af2a0bb-f1cc-4ab7-9b39-fce5a616ce30 · inbound

Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection cites this paper.

Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T08:41:12.351655Z digest=sha256:f7fef4365123d13af4dd532972abdc380ea363660559a9de06361129af409a57

Observation c9532299-39dc-434d-8e4f-e4eb38db7505 · inbound

Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics cites this paper.

Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T04:25:58.275705Z digest=sha256:0f036424866de4e004ffbb9a1ebb3e3a93c5e6b56495943eb2ec69df01ca348e

Observation 63c4e5bf-2640-455c-bb60-6d58b17f7ee6 · inbound

Bridging Perception and Action: A Lightweight Multimodal Meta-Planner Framework for Robust Earth Observation Agents cites this paper.

Bridging Perception and Action: A Lightweight Multimodal Meta-Planner Framework for Robust Earth Observation Agents EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T15:45:27.700503Z digest=sha256:8e03d84a513543c498d8e9ece36aca3a8d372763e0d3cb5ab2eae114cbf36766

Observation 97d78bb7-9df8-40fb-b21d-8f6ddf0cef5a · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:28:48.722301Z digest=sha256:387b29b9329d3dafacd525a21f3b17df15f01112d7b1c83134b6e298b77da96a

Observation c9b6b835-4073-45d7-b2b8-b274ab0e0209 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:02:27.788287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T06:57:25.015358Z digest=sha256:8333ac83a700064cb6a42345d76d07fe000cefacc3dc6428eb69186f8c5d6b47

Observation bce4d147-995c-4704-bbab-934d62077473 · inbound

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers cites this paper.

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:35:46.694857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T22:56:43.298141Z digest=sha256:010b33cccb727d8a44b2e8956476dfcf7ace8ef34ffff0596692654f16bfce2b

Observation 6f2a0980-4fe0-43ec-a40c-19e1f347e0b8 · inbound

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? cites this paper.

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:06:45.858231Z digest=sha256:08217119bba51ec9e54567959346cdda6b449c4835f9370395efc72d8d5c584f

Observation 45991b55-7280-4c1b-a614-91ef1c61d32a · inbound

MolSight: Molecular Property Prediction with Images cites this paper.

MolSight: Molecular Property Prediction with Images EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:54:22.133331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:05:21.133255Z digest=sha256:3d5fd25c042a3c323843def97c9506ca6cee35a3a5d4f58b277af445421470f4

Observation b2353877-f168-44ca-8daf-8701b3c72e0a · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 118

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.640200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:a3e4356a85e8589ec0cbdca36e19eb3239fe4c2a02c0213e75e0e1228e9e889c

Observation 6f653ca6-00a1-4094-95d5-0f989c5e488f · inbound

Same Image, Different Meanings: Toward Retrieval of Context-Dependent Meanings cites this paper.

Same Image, Different Meanings: Toward Retrieval of Context-Dependent Meanings EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:52:35.402683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:52:29.268016Z digest=sha256:33b4ec15d44a8feb811b6efe12b2d2890a17020c6cbbfbf063b27966edaebf53

Observation a44a9293-ed48-415e-ab13-e7ba2a6cfea1 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 106

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:22:56.360040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:3421576a29c890f33d5b5939bf8ee73bc0e9ea21ff585edfe50b55ca5db5e778

Observation 456a6c23-89ab-43f7-8f8b-d4f795c52cab · inbound

GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding cites this paper.

GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:47:58.705426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:44:28.495239Z digest=sha256:68e89aecefb7b3e8d66f4e09e2476d06ab54c81b89fcad9d9fcc6d09e067502c

Observation d2a2bc53-784f-4f62-afa5-b8704c1bbe24 · inbound

GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding cites this paper.

GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T05:15:50.090130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:15:50.090130Z digest=sha256:cd22ef6e2314c3d5202e48cb1872d3ffa983d0288782128e535f230505c069a7

Observation 4ff4f497-ab03-4cb1-b11d-9b5270e59472 · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:47:37.151940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:43:08.029165Z digest=sha256:7724fc96c488ecf647ae5019e97442642bce039bc3a57d3731ae13d91ccddb80

Observation 5c614f5c-af61-4ca5-a05f-db660ab649af · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:45:05.683078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:42:05.592911Z digest=sha256:457a5345aecf1548cab1fa0e66554b8eb1f2e30fe195c51952feb25e531905ec

Observation 0dd16ec6-311f-4086-a8c7-c294a7d3b426 · inbound

WOW-Seg: A Word-free Open World Segmentation Model cites this paper.

WOW-Seg: A Word-free Open World Segmentation Model EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:27:48.020448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T21:23:14.311122Z digest=sha256:1c6cea133ff74511bdaf28c2c7f15a583cebf1c8671df25204fcd8fafdf9eb10

Observation 48af3912-55bb-4a14-851d-6d30301f033b · inbound

What Matters for Grocery Product Retrieval with Open Source Vision Language Models cites this paper.

What Matters for Grocery Product Retrieval with Open Source Vision Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:03:15.406087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:59:53.366287Z digest=sha256:3e0f5b09a7358d7c7e14ace2a124dfff54cef20ffa70021f9a4fc92a6e228b0b

Observation cf63bafa-0192-4911-ae52-717ba1ed271d · inbound

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register cites this paper.

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:48:05.617493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T06:48:00.733211Z digest=sha256:dce1496547606a910a247a1cbd4afdbf1c8bde12850d60befd7d4fc8a75e1640

Observation c849cffa-b676-4395-ac69-365c97ec00e8 · inbound

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition cites this paper.

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:19:39.194693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T05:18:29.720630Z digest=sha256:20b1a9a37a8f47b04100d72f119e584a4429154357b0ff799630c5c7d4e9a7e5

Observation 276bc5df-82e1-4a2b-a522-693e0f47fd08 · inbound

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition cites this paper.

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:14:57.313973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:06:26.561698Z digest=sha256:47a676426505d96752d83643e54c87f04f05e3233d7cd65c69ea5db08b463d45

Observation b7e4db32-6239-4062-87ed-300494ba8dd5 · inbound

The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP cites this paper.

The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:03:51.394443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:56:02.724596Z digest=sha256:780fbaa1d6d13e04f17037ec798cbade12168fb5439073e9c65633e1a66dfc31