Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T01:54:21.943160Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 100 inbound Pith citation observations for arXiv:2303.15389.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T01:54:21.943160Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:01:21.903186Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
53 of 53 outbound references displayed
External citation measurements
80
pith, observed 2026-08-05T02:28:24.338817Z
Observation 5d4ac27b-e57c-4063-999b-dfd616550206 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale https://laion.ai/blog/giant-openclip/
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 905cb502-ae4f-4887-ac92-39bcffaf06ff · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale BEiT: BERT Pre-Training of Image Transformers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 47538e7b-3219-4811-a63f-f13c3d2dd47c · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Ob- jectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84c86ebd-289e-49b5-a5f1-2bc66792c690 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Birdsnap: Large- 5 config EV A-01-CLIP-g / EV A-02-CLIP-g+ image enc
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3be0339e-3f28-4a58-95f6-c2d4a21eb001 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Food- 101–mining discriminative components with random forests
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c713905-633c-42f9-9547-24cb2b680173 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Coyo-700m: Image- text pair dataset
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9227a8ef-d6a0-4d3c-9757-84e1b04262e7 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale A Short Note about Kinetics-600
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40df6e5f-c463-4583-b6c4-1466c3d0c92a · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale A Short Note on the Kinetics-700 Human Action Dataset
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3847e30a-d5d1-4bb2-803f-4f612dcd4e8a · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Quo vadis, action recogni- tion? a new model and the kinetics dataset
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c06272bf-e3a1-4305-bfd0-89ea29a742b7 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Train- ing deep nets with sublinear memory cost
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e0c82d9-1ca8-4a33-a341-5cf5fb20c7b0 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Remote sensing im- age scene classification: Benchmark and state of the art
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f2bd24a-7cf9-41b1-84e6-3b6d0e2f7775 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Cimpoi, S
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8929ef37-d1b6-4a04-8b9b-f0402469abd1 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 717a9c38-ba92-48d0-86f0-8dc0f1e20d74 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale An analysis of single- layer networks in unsupervised feature learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d8d4586e-eca9-4adb-857f-349b281b601d · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Fu, Stefano Ermon, Atri Rudra, and Christopher R´e
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6cfbdbf5-7aa9-4573-8694-628b8893427a · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Imagenet: A large-scale hierarchical image database
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72644616-38fc-45fd-bedc-7a895f2b1e88 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale An image is worth 16x16 words: Transformers for image recognition at scale
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c77d9479-4b93-49f0-8659-f4903cc3dc19 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale The pascal visual object classes challenge: A retrospective
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e6073abb-c3e6-41e7-9142-55638d434202 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale EVA-02: A Visual Representation for Neon Genesis
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4b240e77-6838-4723-a1a2-7770350f0f94 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e20181b7-f29b-458d-8159-f674dc8cb0a1 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Learning generative vi- sual models from few training examples: An incremental bayesian approach tested on 101 object categories
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ce706aa-5ae4-4af5-8c40-090708f9b9a4 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Challenges in repre- sentation learning: A report on three machine learning contests
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7058e54c-3bd3-4e16-95ea-db0cbdf8c405 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.IEEE J
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ed2b2a1-72cb-4008-bdaa-7c97f5bb18d2 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale The many faces of robustness: A critical analysis of out-of-distribution generalization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc7ce154-af0c-4833-b404-32e89e3ff53c · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Natural adversarial examples
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 68a7f5ce-8d31-450d-9300-2118792fe1cd · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Deep networks with stochastic depth
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc3cdfbf-a4a1-43fc-b352-bc6849e14b7c · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Openclip
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3fc2eed-9b09-4827-854a-056c7e6415a1 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Adam: A Method for Stochastic Optimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4429f131-efeb-4bf8-8c68-07cd4f61083c · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale 3d object representations for fine-grained categorization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a2d9850-3d62-46b4-8589-03158af307fd · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Learning multiple layers of features from tiny images
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 86370879-fa09-42fe-9550-642faf874be0 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Gradient-based learning applied to document recognition.Proceed- ings of the IEEE
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4157f1f2-ad2d-4c38-8bcc-eda18a0fd7fd · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7fa97d15-8688-401f-b5b5-20dcf49962f8 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Scaling language-image pre-training via masking
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc571e50-a4b3-46d6-8ac1-2cc18354ea40 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Mi- crosoft coco: Common objects in context
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4dfa8e20-811f-45cd-92d3-6e833563521d · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Decoupled weight decay regu- larization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 585dcffc-9241-46b2-b18f-85bb2e2024be · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Fine-Grained Visual Classification of Aircraft
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1717ca60-b062-4b7f-8ee2-58923a460f94 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Automated flower classification over a large number of classes
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4f03c5b8-54cc-4747-acdf-1c9d3e18005b · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Parkhi, Andrea Vedaldi, Andrew Zisserman, and C
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e8d43270-fa26-4155-9939-2e343b7b816b · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Learning transferable visual models from natural language supervision
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e2d96f6-9869-40b7-9a61-86d87ea4d916 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Zero: Memory optimizations toward training trillion parameter models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 38a04513-429b-44bd-8d97-2e2d7bda9b58 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8197d098-652e-44a6-945c-8404abe851bb · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Zero-Shot Text-to-Image Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e833415e-e83a-40c0-9403-c017d844f0b1 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Deepspeed: System optimizations enable training deep learn- ing models with over 100 billion parameters
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 24cbce26-2a3c-4a63-8a39-b3c00860d10f · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Do imagenet classifiers generalize to imagenet?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5ecfb2e-8233-4f49-be7d-837cd80cfe14 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation beddd684-84b6-4f4c-a0b7-0a12b308952f · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84b45e81-4f67-413c-9e81-f7a43c26b1b4 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0aac5f7d-690a-4287-87f5-ec929279514f · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d9c10ab-a031-40cc-8064-8973b00ecbac · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Rotation equivariant cnns for digital pathology
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93db47bd-4d8a-4d1c-a88a-11671c53c8c5 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Learning robust global representations by penalizing local predic- tive power
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e5eb02c-88a3-4689-bc21-19da9493582b · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Sun database: Large-scale scene recognition from abbey to zoo
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fbc6978b-0bef-4c9e-8c7e-b508f0ff55e1 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale Large batch optimization for deep learning: Training bert in 76 minutes
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 094735e7-f008-43ea-90a5-b0ef0af9bcb5 · outbound
EVA-CLIP: Improved Training Techniques for CLIP at Scale From image descriptions to visual denotations: New similarity met- rics for semantic inference over event descriptions
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae2f0075-525f-4e22-9d3a-765eb44e0ea0 · inbound
Sigmoid Loss for Language Image Pre-Training EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e12117b1-6c89-4e1c-a853-cc27b8436508 · inbound
VideoChat: Chat-Centric Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b02c8501-d249-4816-abf6-0acedb722186 · inbound
A Survey on Multimodal Large Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d75a59a5-7492-4341-8443-e0ed414b8f02 · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c9fbff2-1c1d-46ed-8d07-0a09f3a29daf · inbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f169a1d-6986-4a9a-849f-96dbd15aee9b · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 131
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0dafeb92-717b-4721-97ee-d5280200245b · inbound
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40e29025-82a3-4831-9d38-486880913f7b · inbound
Agent AI: Surveying the Horizons of Multimodal Interaction EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8e50dc18-bbd7-4efc-bb83-1dbe5e3c12a5 · inbound
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 221204e8-98ad-4dc9-9854-48309924cf39 · inbound
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b8d17ed1-6576-4d50-ad15-e05a10d25bd1 · inbound
BLINK: Multimodal Large Language Models Can See but Not Perceive EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c174f23-7ae4-45ae-9805-5c7e612debb2 · inbound
LVBench: An Extreme Long Video Understanding Benchmark EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 49d8516d-9262-4c1c-a24f-9aa1bdd97fa8 · inbound
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc55ca00-0a0b-4316-826b-277fe5ea1a14 · inbound
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 01cd2b6d-5285-4c09-9eb0-122aa8847b25 · inbound
E5-V: Universal Embeddings with Multimodal Large Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a75f5f2b-8ca2-4d28-8161-68348ad4b46a · inbound
CogVLM2: Visual Language Models for Image and Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9d723cf-4dce-4df9-ae41-a22f7e88f4af · inbound
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3c40a80-aaae-4644-a7e6-9cbeaecad50a · inbound
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c158e121-1f3b-4453-ba3a-598fe83fb317 · inbound
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c211992d-0b7a-4c40-83ca-94bf2190d47a · inbound
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4b1ba38-dc02-4bf6-9bf2-3ea499ab52d6 · inbound
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 187
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f950d82-b38b-482a-b3af-a31985f2dcf6 · inbound
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c7e599d-84db-411c-97b8-1fe0f5d5668f · inbound
Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5708f05-1ea0-4ae6-8afe-a53c1e48e17c · inbound
An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a0c0629-f13f-49fd-8b7a-bdcf99bef568 · inbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b37da83f-52b4-4b70-b983-5f55f948a42b · inbound
Perception Encoder: The best visual embeddings are not at the output of the network EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 364b53bc-a8d8-4e6d-8713-91222701105e · inbound
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40473f12-a9fc-4851-a7b8-6e9de38e6e0f · inbound
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation affc2308-9aa7-48b9-a23b-d18a1657460c · inbound
SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a312fb-d024-40b9-b282-71b5beeada23 · inbound
Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c9abd3d-f647-47b3-830b-12f952e50b1c · inbound
MobileCLIP2: Improving Multi-Modal Reinforced Training EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406f1b2e-2c25-44fd-b7c9-1c26f78867aa · inbound
Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6aceb296-ed2c-4b6d-90ac-699e9b6347f5 · inbound
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2d429898-6337-47ec-82bf-e4572a281f2c · inbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e627282e-42a3-4f9b-b6dc-ab948a7b2dd6 · inbound
EditIDv2: Editable ID Customization with Data-Lubricated ID Feature Integration for Text-to-Image Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91419f0c-a081-47d8-a75b-c9473fdbcc9b · inbound
Recurrence Meets Transformers for Universal Multimodal Retrieval EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f796c12-9c89-471f-8c58-1869dedd49f4 · inbound
FreeRet: MLLMs as Training-Free Retrievers EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 917f7d0a-cc6d-4d83-a965-69ea11d1b269 · inbound
FreeRet: MLLMs as Training-Free Retrievers EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f0c0bc-4e7b-4dbd-a477-6dc96c8ff0d6 · inbound
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 309d3348-5a97-4c30-a655-16d62f2c23a9 · inbound
VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae9baa82-be3a-43dd-8b0c-984158138234 · inbound
Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454c9428-df48-4687-93bf-67bfdabbfa55 · inbound
Calibrated Multimodal Representation Learning with Missing Modalities EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c4bd4a5a-8f24-4c82-b3c8-e7418fded368 · inbound
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb73885-1c46-440f-ad65-7759c4bac2b1 · inbound
PowerCLIP: Powerset Alignment for Contrastive Pre-Training EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ba39f140-8359-4390-8f64-4b034923e361 · inbound
CLIMP: Contrastive Language-Image Mamba Pretraining EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b4f20f0-2d11-43ac-bc66-bc4774e5e052 · inbound
R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5bcda11-180d-44ba-b6f2-16c95b643696 · inbound
R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f17a7717-22af-45d5-9ee7-f604b22c5562 · inbound
CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ae307be-0700-4d1b-b25c-230300ffe8cb · inbound
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b76a954a-60de-4052-9867-d07859b42d8d · inbound
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18bf3ac8-75e0-4c52-9119-7bb7f850d2c3 · inbound
Mitigating Error Accumulation in Continuous Navigation via Memory-Augmented Kalman Filtering EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 213a78c7-16c9-46c7-bafe-f8660ede7be5 · inbound
Vision Transformers Need More Than Registers EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c703e467-541e-4475-8e23-6bbed5b26eeb · inbound
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1bafc80-6dde-46d0-a709-4f962c83870a · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a4dd4b93-5251-4c1a-839e-c3a7b5e5d2e1 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4626e353-1038-4c0e-a310-470ee0b2c314 · inbound
Revisiting Model Stitching In the Foundation Model Era EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a2730e-4f28-4c42-ad70-a433ef7707ac · inbound
SteelDefectX: A Multi-Form Vision-Language Dataset and Benchmark for Steel Surface Defect Analysis EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6eb0dce-44dd-409f-ac3e-e5d9f962cbd4 · inbound
When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 99ad757c-a944-4b0c-8ff2-e602d0c9f040 · inbound
XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5059d5a5-ed91-4b05-a0bb-bedd13ee31c4 · inbound
Video-Oasis: Rethinking Evaluation of Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51bcb73d-4e4a-49d5-aa19-4ffc23d77795 · inbound
RGB-Pointmap Pretraining for Unified 3D Scene Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c74b8aad-87a8-4339-ac73-1a0c9e8d9472 · inbound
Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patch for Infrared Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 01c1ef13-9723-42db-817f-f4f29fca4bbd · inbound
Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patch for Infrared Vision-Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2596e2c1-cf21-440c-ba53-2ffff9b0dae0 · inbound
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c2268b87-e105-467f-8943-9fa41c8db29b · inbound
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a4b174d-15e8-4d43-9ff9-14b0352c9619 · inbound
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 824bfa98-3039-43a3-bdf5-cdc6419f28ce · inbound
NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f2ca36ba-bf71-4b66-99b2-967c260953f0 · inbound
TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6b520045-579d-42f9-a8ca-fe98c2fbca43 · inbound
Boosting Robust AIGI Detection with LoRA-based Pairwise Training EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e5545ba7-c333-444e-a0bb-55c4e3ec9a36 · inbound
Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dfdbe15f-615a-485a-b799-a6975a1d15ff · inbound
Dual-Modality Anchor-Guided Filtering for Test-time Prompt Tuning EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9140e422-ebbd-4ccd-82a1-c3fed2d59e3c · inbound
Cross-Attentive Multiview Fusion of Vision-Language Embeddings EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9726290c-7a77-4a8b-9ace-ed0bf7ed6695 · inbound
Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5aaf68ef-9c46-437b-80c0-7e92ad3c1716 · inbound
AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf904e8c-3cd1-404b-9a84-f73290ee5a9c · inbound
Exploring High-Order Self-Similarity for Video Understanding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b8bb0c24-34a2-4d37-8e17-fc7fd19f2d15 · inbound
MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 090bac04-f443-40d8-b78c-0b9efe46087e · inbound
Photonic Quantum Computing on Spin Memory Architecture with Tree-Encoded Fusion EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 21cd8b1a-969a-458e-b020-72f2f5d7dfe3 · inbound
Rethinking Cross-Domain Evaluation for Face Forgery Detection with Semantic Fine-grained Alignment and Mixture-of-Experts EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e4614da5-a0ba-4aa0-8fe8-2fd7dd6a43f8 · inbound
BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0af2a0bb-f1cc-4ab7-9b39-fce5a616ce30 · inbound
Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9532299-39dc-434d-8e4f-e4eb38db7505 · inbound
Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 63c4e5bf-2640-455c-bb60-6d58b17f7ee6 · inbound
Bridging Perception and Action: A Lightweight Multimodal Meta-Planner Framework for Robust Earth Observation Agents EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97d78bb7-9df8-40fb-b21d-8f6ddf0cef5a · inbound
jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c9b6b835-4073-45d7-b2b8-b274ab0e0209 · inbound
jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bce4d147-995c-4704-bbab-934d62077473 · inbound
jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f2a0980-4fe0-43ec-a40c-19e1f347e0b8 · inbound
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs? EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45991b55-7280-4c1b-a614-91ef1c61d32a · inbound
MolSight: Molecular Property Prediction with Images EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2353877-f168-44ca-8daf-8701b3c72e0a · inbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f653ca6-00a1-4094-95d5-0f989c5e488f · inbound
Same Image, Different Meanings: Toward Retrieval of Context-Dependent Meanings EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a44a9293-ed48-415e-ab13-e7ba2a6cfea1 · inbound
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 456a6c23-89ab-43f7-8f8b-d4f795c52cab · inbound
GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d2a2bc53-784f-4f62-afa5-b8704c1bbe24 · inbound
GeoFlowVLM: Geometry-Aware Joint Uncertainty for Frozen Vision-Language Embedding EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ff4f497-ab03-4cb1-b11d-9b5270e59472 · inbound
AttenA+: Rectifying Action Inequality in Robotic Foundation Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c614f5c-af61-4ca5-a05f-db660ab649af · inbound
AttenA+: Rectifying Action Inequality in Robotic Foundation Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0dd16ec6-311f-4086-a8c7-c294a7d3b426 · inbound
WOW-Seg: A Word-free Open World Segmentation Model EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48af3912-55bb-4a14-851d-6d30301f033b · inbound
What Matters for Grocery Product Retrieval with Open Source Vision Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf63bafa-0192-4911-ae52-717ba1ed271d · inbound
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c849cffa-b676-4395-ac69-365c97ec00e8 · inbound
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 276bc5df-82e1-4a2b-a522-693e0f47fd08 · inbound
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7e4db32-6239-4062-87ed-300494ba8dd5 · inbound
The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.