Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:14.795852Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 4 inbound Pith citation observations for arXiv:2506.02557.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:14.795852Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T22:07:35.021986Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T22:09:07.132811Z
98 of 98 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4d16e6ee-d894-4bda-8dcc-7acfeb66ae2e · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Tallyqa: Answering complex counting questions
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac6405e-c393-4fe7-be5e-44e9302a72b6 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8e08a7-39c5-472a-bf0b-f4305a43ef72 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Multi-label cluster discrimination for visual representation learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd2aea2c-c9b9-4bd1-894b-5ac1d6191e03 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae69c505-1c92-4049-8542-abc0a16de6f6 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93b3c931-1656-477b-91ee-e9678554b1ea · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models BE it: BERT pre-training of image transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15631b84-da0c-4158-890f-8af4e5d2f0cf · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Demystifying MMD GANs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 687fedb5-d726-43a1-a2cd-2cb12ad73329 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Domain prompt learning with quaternion networks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18af799d-c07f-4a26-8766-b41a6ea96b33 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Emerging properties in self-supervised vision transformers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae0c78d-9be3-4ee3-84a6-54fc288609ed · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bc761d1-b06c-4caa-a25b-ce9284b1b416 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aa90695-3be6-400e-943d-09083641276e · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Remote sensing image scene classification: Benchmark and state of the art
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bc4a2da-efac-488c-a0ec-b5892fa02b4a · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Describing textures in the wild
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f684b7-2458-4e0c-844a-5c245287180d · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Locality alignment improves vision-language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea5f798-9654-40d3-b615-2e300fe98c8e · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Vision transformers need registers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76f1de1f-3955-4271-956b-0e71f1f8478a · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29001f20-7d10-4a44-aa2e-3a212bc2928b · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Imagenet: A large-scale hierarchical image database
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f19e8af-fdf4-4222-b513-e6808e566af5 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Data Filtering Networks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1768fd8-cc1c-4a4b-8fb5-cac422a01f1c · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f27885c4-874f-4f66-80cd-ca21c6c67bfc · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models The Vendi Score: A Diversity Evaluation Metric for Machine Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe8321d5-a764-48b8-9d49-79270960917b · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e9ed5a2-454f-4ca2-afc6-0660046e0e4a · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clip-adapter: Better vision-language models with feature adapters
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01bb1b83-88d7-447e-a6f6-e6d5055328d6 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models 3dsam-adapter: Holistic adaptation of sam from 2d to 3d for promptable tumor segmentation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b00e8df-c653-435c-8b8a-207190ba5e94 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Boosting the visual interpretability of clip via adversarial fine-tuning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c80418-3b83-47c1-9d71-9fe959274c87 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models J., Erhan, D., Carrier, P
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6930017c-288a-4408-b09a-498b9b8db2e4 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ecbac2-a02a-4822-91ae-6e21c4010444 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Recovering low-rank matrices from few coefficients in any basis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 366ffc2e-2196-4e70-8d5f-1cc98c69c510 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a7fc68b-0e04-4eca-a9b0-df2e2203c2b7 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69abfb6-bd6c-44df-9b2b-e0d23c980b20 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models and Ozay, M
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 121c676e-fa5f-48ad-9c50-690fa20805bb · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Masked autoencoders are scalable vision learners
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a254a0e-9dcf-4e16-949a-9b1c1abc6c1c · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Introducing eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d7b9426-0970-4aae-b9d7-3d9b4b0597f1 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Natural adversarial examples
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf580be8-ad12-443a-9c32-81712a0531de · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Probability inequalities for sums of bounded random variables
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4416e393-5b99-45c6-9ea9-dba6335057b2 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc363e2b-3006-4734-a536-73a19962b57a · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec4452a5-ec39-44de-b667-91adde54424a · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models T., and Farnia, F
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eae863a3-ed91-4854-9112-b5ceaa723332 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models T., and Farnia, F
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbba5313-2aad-4768-8d8a-e44aa9ffa657 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206e2dba-54a7-45c6-8e8b-d3e48972387e · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad4df0c3-f888-4b8c-81b7-7cf62eae08cc · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models What's ''up'' with vision-language models? investigating their struggle with spatial reasoning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc348745-03d8-4a07-95c6-203ae6f6619c · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Studiogan: A taxonomy and benchmark of gans for image synthesis
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a1764d-f340-4029-ab33-673f15dd52c5 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Referitgame: Referring to objects in photographs of natural scenes
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd87e030-16f8-4a32-8da4-d7c713a2ebcb · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models A diagram is worth a dozen images
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d4571e4-a51e-4c97-a34b-51d71ddfc5ed · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models The hateful memes challenge: Detecting hate speech in multimodal memes
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f904f8d-1662-4281-94b0-c76d3ac765c4 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da99cb35-99f2-40e9-b0d1-58ece4309242 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Learning multiple layers of features from tiny images
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdc3bc36-8274-4109-a257-e0b499b0d560 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clip benchmark: Clip-like model evaluation, 2022
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 986ea017-a1a9-423e-a162-80aa191b2015 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab00bb5c-ea22-4fe4-b469-420a53fd3470 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Transfer learning in computer vision tasks: Remember where you come from
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e1a228a-284d-4834-acf0-c75ca3829e9a · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Evaluating object hallucination in large vision-language models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b55b8aa9-9a3c-492d-85ad-173108fd5013 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 941bbed5-eb97-4269-8ad7-47c1f58e61e1 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Visual spatial reasoning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95581358-a2dc-4a8f-8e01-5f39a1289963 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad658fa4-a1d9-4e11-88d2-2c6fa4d826b1 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Decoupled Weight Decay Regularization
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b9b5496-ad13-4cc1-b52b-2c82d1c81616 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Understanding Zero-Shot Adversarial Robustness for Large-Scale Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97d40f26-43b0-4ba1-b6c8-bbf546f4a1d4 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53d26859-b591-4ff0-8722-4807ae2a4770 · outbound
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76527488-a35d-443a-aab9-e2a81759f282 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Dinov2: Learning robust visual features without supervision
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dcf1e64b-a01b-4377-bd5c-3489b89a64c3 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5fae69-ee38-41bc-b5dc-14c10c800486 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74d19f35-21aa-471d-910a-247e7ee840c8 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards a scalable reference-free evaluation of generative models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43bac356-ca4d-42a1-81aa-2d9b908b5fc7 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models M., Vedaldi, A., Zisserman, A., and Jawahar, C
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2618be62-08dc-4987-b436-f6c8b8b33a5d · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b7853d-b122-4b04-98f4-4a8a043db632 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Am-radio: Agglomerative vision foundation model reduce all domains into one
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 918a3279-df71-4ef6-937d-790a1d9df14b · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Be More Diverse than the Most Diverse: Optimal Mixtures of Generative Models via Mixture-UCB Bandit Algorithms
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ac9159c-f5d3-405d-81b3-fcf795583ee8 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Improved zero-shot classification by adapting vlms with text descriptions
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e5c131a-6d20-4eff-b873-ca932678061f · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models CLIP meets Model Zoo Experts: Pseudo-Supervision for Visual Enhancement
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebf584df-d99f-4edd-a28a-dc9ee716d36c · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b6a7b7-2294-4b70-ba5a-04f1a26dbd1c · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f02e0a-268b-47b9-8f9d-0ec95bc9ce56 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Finetuning Text-to-Image Diffusion Models for Fairness
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5efff50b-5682-40c5-b5d9-0ed382beb549 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c03fab-1dc6-414c-9dd9-3cf325655264 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards vqa models that can read
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5eeb4e4-5a06-4026-845f-44b8f6fe2d19 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a43fc90d-4703-49a7-8e49-220310d8585c · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models L., Taylor, E., and Loaiza-Ganem, G
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92360aea-3301-4b3a-8ae3-14ba88afdfd5 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75fce3a4-6058-4509-a68a-2ce192512d2d · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Winoground: Probing vision and language models for visio-linguistic compositionality
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a11379d-27de-4437-81c4-d11d76d335b6 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8958d4da-88b8-4d83-a640-75d27891eb51 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models S., Linmans, J., Winkens, J., Cohen, T., and Welling, M
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd473008-39b0-4582-86b0-9cf57d3e5756 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clip the gap: A single domain generalization approach for object detection
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5740819-6d52-4e95-8a60-38da9a6d39f4 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models S., Steiner, A
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7a586b6-5dc2-43c8-bc52-d27d87f85af7 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65186142-1aa3-4919-9f62-ca235b46f4c6 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Diffusion feedback helps clip see better
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62d63c46-d754-4f26-828f-ce1d23533941 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a43c192a-2ac7-447c-a69d-9cc493a6cfcf · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Demystifying CLIP data
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c367537a-4415-49b2-9169-6420db8b0cea · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Explicit inductive bias for transfer learning with convolutional networks
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 984b5b48-16a6-4047-ac45-5dbea0e582a2 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6135983-a5c7-4bad-a1fd-8200c9309782 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models C., and Berg, T
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74155806-e21b-4baf-881f-117b867164ee · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 215a7a43-1c7f-4c53-bc74-17924b037985 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations, 2023
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2cb6e895-9f87-4c8d-8c4d-69b22d6ba8ec · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Sigmoid loss for language image pre-training
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bfa7c29-6693-46f8-b2ca-bd214c111d9c · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unveiling Differences in Generative Models: A Scalable Differential Clustering Approach
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f147c74c-a66a-443c-a7e1-333c75204a53 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models An Interpretable Evaluation of Entropy-based Novelty of Generative Models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc3159f4-8721-495b-b9dd-4cb65b7c68ef · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Tip-adapter: Training-free adaption of clip for few-shot classification
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd3c76e1-ac43-4256-a6c3-cb6b3b7ec446 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models H., Zhou, L., Dai, X., Yuan, L., Li, Y., et al
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aabcead0-936a-4ad9-9929-e5d6cdbdf76a · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models C., and Liu, Z
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef46bbe8-0ec8-4a86-b913-da9bd885d844 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Rethinking Centered Kernel Alignment in Knowledge Distillation
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32880931-6e19-4a41-8cc9-ebf2088f0d25 · outbound
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models write newline
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d62736ac-a819-41bb-a7c7-eb0abf546778 · inbound
UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3fff1ec-2d20-4ddc-ad54-cc96cf462af6 · inbound
Latent Denoising Improves Visual Alignment in Large Multimodal Models Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b84b43d-5e64-4629-a595-4b924cbc06b4 · inbound
A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 885178fa-8e7d-4394-a495-a60a3a927002 · inbound
A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.