Pith. sign in

Paper Citation Record · LEDGER

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

As of 8 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 4 inbound Pith citation observations for arXiv:2506.02557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02557 v1

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:14.795852Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T22:07:35.021986Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T22:09:07.132811Z

Reference resolution

98 of 98 outbound references displayed

  • verified exact4
  • verified fuzzy35
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d16e6ee-d894-4bda-8dcc-7acfeb66ae2e · outbound

This paper cites Tallyqa: Answering complex counting questions.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Tallyqa: Answering complex counting questions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.076810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.076810Z digest=sha256:6a3525846b8aadec7b4c87b14dcdda6248bacc14e4075f1a7414d140a019da14

Observation 9ac6405e-c393-4fe7-be5e-44e9302a72b6 · outbound

This paper cites Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.168645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.168645Z digest=sha256:67e14721872294cfb417af7444fb498507523086f8a8fd44a709006c091e14f4

Observation 5b8e08a7-39c5-472a-bf0b-f4305a43ef72 · outbound

This paper cites Multi-label cluster discrimination for visual representation learning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Multi-label cluster discrimination for visual representation learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.268183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.268183Z digest=sha256:a450d4fa31c3a1e982bf8a2e90087c3ff9e3d3b599ff078ab4f04779181de131

Observation dd2aea2c-c9b9-4bd1-894b-5ac1d6191e03 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.323907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.323907Z digest=sha256:0827ac27492ffbc1b55c2d268f2485a90f0387a435a52e7329b93a69e658e1d8

Observation ae69c505-1c92-4049-8542-abc0a16de6f6 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.361069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.361069Z digest=sha256:3ef19fc82f85f077aa05c4ad66503896983e4f63bb8cb233148f57d67acb5bab

Observation 93b3c931-1656-477b-91ee-e9678554b1ea · outbound

This paper cites BE it: BERT pre-training of image transformers.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models BE it: BERT pre-training of image transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.366454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.366454Z digest=sha256:385349000aa3c46184f28c99a4cdfcca38c4e2a4746a7d5961b759848f65502f

Observation 15631b84-da0c-4158-890f-8af4e5d2f0cf · outbound

This paper cites Demystifying MMD GANs.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Demystifying MMD GANs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.372034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.372034Z digest=sha256:b72487b168c0f9a68b096db85a6a461760084f78857f1d62f099a9d87744a237

Observation 687fedb5-d726-43a1-a2cd-2cb12ad73329 · outbound

This paper cites Domain prompt learning with quaternion networks.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Domain prompt learning with quaternion networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.379895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.379895Z digest=sha256:6878fd58bff01a7d61158004babeb21c818d06f1595b278366b0ce486e839c3d

Observation 18af799d-c07f-4a26-8766-b41a6ea96b33 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Emerging properties in self-supervised vision transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.388470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.388470Z digest=sha256:c47d46ac86c49922f4ba561a918edf7e1d13617b77f0d8a6359b91b53ad13689

Observation 4ae0c78d-9be3-4ee3-84a6-54fc288609ed · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.395742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.395742Z digest=sha256:7f77e3bfd2454683e38aae5185229ab22e6e2baa9c6c03cc7a8247ba12b75ad7

Observation 5bc761d1-b06c-4caa-a25b-ce9284b1b416 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.404320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.404320Z digest=sha256:b3c7e6dec0126b18bbb120b7235119450965a1b3f2ca4d7fa92197edd708c3bf

Observation 9aa90695-3be6-400e-943d-09083641276e · outbound

This paper cites Remote sensing image scene classification: Benchmark and state of the art.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Remote sensing image scene classification: Benchmark and state of the art

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.413877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.413877Z digest=sha256:404fe358a4edc5f57b986220a07bfa066f7ac2727b4851f2f34cfe6f2f51ecfb

Observation 2bc4a2da-efac-488c-a0ec-b5892fa02b4a · outbound

This paper cites Describing textures in the wild.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Describing textures in the wild

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.424318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.424318Z digest=sha256:d7115d2e9fded2d2763cc88c038b20b1e21bd9a53e5faae327ede5873d03f676

Observation 01f684b7-2458-4e0c-844a-5c245287180d · outbound

This paper cites Locality alignment improves vision-language models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Locality alignment improves vision-language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.433254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.433254Z digest=sha256:ebe7dca639d5ced2a8ae6d540798ebf658a1749608a366bc4045c1e13d9f993d

Observation 9ea5f798-9654-40d3-b615-2e300fe98c8e · outbound

This paper cites Vision transformers need registers.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Vision transformers need registers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.442273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.442273Z digest=sha256:f399783c1cd164cb39da3b535e7e09879a98dbeb3527a9598afdd584eaa46e77

Observation 76f1de1f-3955-4271-956b-0e71f1f8478a · outbound

This paper cites FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.449072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.449072Z digest=sha256:4800c6c1a31ec960e77f888de7869fa40102c96f4e869e3810c3bf0232a5c785

Observation 29001f20-7d10-4a44-aa2e-3a212bc2928b · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Imagenet: A large-scale hierarchical image database

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.458149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.458149Z digest=sha256:a766508314283c22b617051b43aef77a01784409886d8d9019cecb2e59309507

Observation 0f19e8af-fdf4-4222-b513-e6808e566af5 · outbound

This paper cites Data Filtering Networks.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Data Filtering Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.466593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.466593Z digest=sha256:4b2b747dc397b8014ee3cb97867eb04d821e2984ecb26f1314d80648ca44c1b6

Observation e1768fd8-cc1c-4a4b-8fb5-cac422a01f1c · outbound

This paper cites Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.474213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.474213Z digest=sha256:cda36fa35ae05563f4ec8267cdd0ea4bf0bb5e8d90ddcaff3e2da38ea4b06bca

Observation f27885c4-874f-4f66-80cd-ca21c6c67bfc · outbound

This paper cites The Vendi Score: A Diversity Evaluation Metric for Machine Learning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models The Vendi Score: A Diversity Evaluation Metric for Machine Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.482039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.482039Z digest=sha256:03cae7645056426255dc725584002a1fe8083ed975f404f701d2740886b7bc9d

Observation fe8321d5-a764-48b8-9d49-79270960917b · outbound

This paper cites Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.489672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.489672Z digest=sha256:275a15fdd26a00b8d0ec5f8f08384a16a0b39474e1acb562cd035f7fba84bc6f

Observation 5e9ed5a2-454f-4ca2-afc6-0660046e0e4a · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clip-adapter: Better vision-language models with feature adapters

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.498808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.498808Z digest=sha256:9b5552e4eaaef2b0bb99fa713243785675e5c6d4c3991a87a77c5aec15524d25

Observation 01bb1b83-88d7-447e-a6f6-e6d5055328d6 · outbound

This paper cites 3dsam-adapter: Holistic adaptation of sam from 2d to 3d for promptable tumor segmentation.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models 3dsam-adapter: Holistic adaptation of sam from 2d to 3d for promptable tumor segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.507737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.507737Z digest=sha256:67d71f6abe9018df0decf70f12ec7120d84c76ae5fcacec833d40f833d70d0e2

Observation 1b00e8df-c653-435c-8b8a-207190ba5e94 · outbound

This paper cites Boosting the visual interpretability of clip via adversarial fine-tuning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Boosting the visual interpretability of clip via adversarial fine-tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.514659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.514659Z digest=sha256:26747956b801dfafc1744d0d56544648ea5b6f7cbb1dc4f473e5f36b82aaa2ba

Observation d9c80418-3b83-47c1-9d71-9fe959274c87 · outbound

This paper cites J., Erhan, D., Carrier, P.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models J., Erhan, D., Carrier, P

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.523302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.523302Z digest=sha256:2ea451f817a16754139655a61c0af61bf1f19909b1e7692ed8dd3af99aa982aa

Observation 6930017c-288a-4408-b09a-498b9b8db2e4 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.531936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.531936Z digest=sha256:c7783ff2d36c66a22870c456238eed407ee243822cff4be97eaac28164eb3bc0

Observation f2ecbac2-a02a-4822-91ae-6e21c4010444 · outbound

This paper cites Recovering low-rank matrices from few coefficients in any basis.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Recovering low-rank matrices from few coefficients in any basis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.543285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.543285Z digest=sha256:886d0499e67b6f6ecee7e24d117bb9795768a569bf5bf663932041ff254dd6a5

Observation 366ffc2e-2196-4e70-8d5f-1cc98c69c510 · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.322547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.551090Z digest=sha256:865fc10080734833e56a91d55ed8f30a3d8667459ff5e278e58265bb4608971e

Observation 3a7fc68b-0e04-4eca-a9b0-df2e2203c2b7 · outbound

This paper cites J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.559218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.559218Z digest=sha256:3fdde33ee42aa1edec884cf0815cd0b874563285af07400388b84d188b517cf7

Observation b69abfb6-bd6c-44df-9b2b-e0d23c980b20 · outbound

This paper cites and Ozay, M.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models and Ozay, M

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.293617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.569083Z digest=sha256:a96227b643eb2bd12a31d95966162846279fa971464b885a168fd0d3270ca5a8

Observation 121c676e-fa5f-48ad-9c50-690fa20805bb · outbound

This paper cites Masked autoencoders are scalable vision learners.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Masked autoencoders are scalable vision learners

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.577818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.577818Z digest=sha256:39aabfdb171804778c2b8ccfaebb802fd5a80b84e0d2e3bea2277038746ac730

Observation 2a254a0e-9dcf-4e16-949a-9b1c1abc6c1c · outbound

This paper cites Introducing eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Introducing eurosat: A novel dataset and deep learning benchmark for land use and land cover classification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.263624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.585481Z digest=sha256:21ed3273cd3a005f1e57db8fe21078462f6f7dee21ccdfd5bfd44c59b16d0a2f

Observation 4d7b9426-0970-4aae-b9d7-3d9b4b0597f1 · outbound

This paper cites Natural adversarial examples.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Natural adversarial examples

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.245014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.594682Z digest=sha256:5240ae8795c7ecb9a1df418a303b715c21eecbbd58ed9d477bed4b09ba534e02

Observation cf580be8-ad12-443a-9c32-81712a0531de · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Probability inequalities for sums of bounded random variables

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.227681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.603893Z digest=sha256:7fd7306f0396e948efba9a9af970a77f615a84d169fd3fcdf845be4539ff3641

Observation 4416e393-5b99-45c6-9ea9-dba6335057b2 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:16.203710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.611022Z digest=sha256:ff247fcdd07ea93b44941fa8eaec190251ea2b0f8bf70872e527f1bb32e45f18

Observation dc363e2b-3006-4734-a536-73a19962b57a · outbound

This paper cites J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.622913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.622913Z digest=sha256:8eafbdda516783d92be68475380df4a1953548555d1625800ec539919dbed5bf

Observation ec4452a5-ec39-44de-b667-91adde54424a · outbound

This paper cites T., and Farnia, F.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models T., and Farnia, F

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.171988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.629823Z digest=sha256:915ff3282136abb143d8041b0569e8a91eec1006e9739c7e0fad2df1a12b6201

Observation eae863a3-ed91-4854-9112-b5ceaa723332 · outbound

This paper cites T., and Farnia, F.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models T., and Farnia, F

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.156234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.639656Z digest=sha256:789355790e59a19ccd5d79b6ec9a1601fb15431a020d27266dfb24b9f07f50e4

Observation dbba5313-2aad-4768-8d8a-e44aa9ffa657 · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.676257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.676257Z digest=sha256:fa5d81d80322f8fc65155242b75ff8881bfe74156b448b1cec0b8ff6a4a84990

Observation 206e2dba-54a7-45c6-8e8b-d3e48972387e · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.137088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.710170Z digest=sha256:1f0d847b3dc22c71d30a3aa59613ecd862b9f727ca1a0669ab16c63722b6a9e4

Observation ad4df0c3-f888-4b8c-81b7-7cf62eae08cc · outbound

This paper cites What's ''up'' with vision-language models? investigating their struggle with spatial reasoning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models What's ''up'' with vision-language models? investigating their struggle with spatial reasoning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.118412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.745597Z digest=sha256:02f6ec0cdf8ebf07103d14b3abe170a15e8d72d82a97e1b3c7ee6851165690b3

Observation bc348745-03d8-4a07-95c6-203ae6f6619c · outbound

This paper cites Studiogan: A taxonomy and benchmark of gans for image synthesis.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Studiogan: A taxonomy and benchmark of gans for image synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.769979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.769979Z digest=sha256:0ebe756dbebbd4106f7ded67bfdcb51bc3358fc47a44a58b6a5b86c8c55d2e3c

Observation e2a1764d-f340-4029-ab33-673f15dd52c5 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Referitgame: Referring to objects in photographs of natural scenes

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.098644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.788031Z digest=sha256:355102b3473bdfdee9abbb3e0ad2f86cbb85c5381a1910b501e9272942f46212

Observation fd87e030-16f8-4a32-8da4-d7c713a2ebcb · outbound

This paper cites A diagram is worth a dozen images.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models A diagram is worth a dozen images

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.076987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.814240Z digest=sha256:d91b5fca7fc0ecfeaf6d986c07a5961a298550da7ad5adae3b220a45217534e5

Observation 1d4571e4-a51e-4c97-a34b-51d71ddfc5ed · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.058144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.852337Z digest=sha256:6ce650ac7fda8d460eee01c3b5b78eb357148e934da05a815827326414baf24c

Observation 4f904f8d-1662-4281-94b0-c76d3ac765c4 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.878916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.878916Z digest=sha256:cd4b5870a5271aa8ae33d8d0283476bf7070a672f72fa90a482a55f1c74c2247

Observation da99cb35-99f2-40e9-b0d1-58ece4309242 · outbound

This paper cites Learning multiple layers of features from tiny images.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Learning multiple layers of features from tiny images

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.897490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.897490Z digest=sha256:1ec1c590593e986591d5801d4974ed76a0c81770b0a9c9f7608958144cba642f

Observation fdc3bc36-8274-4109-a257-e0b499b0d560 · outbound

This paper cites Clip benchmark: Clip-like model evaluation, 2022.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clip benchmark: Clip-like model evaluation, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.019546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.931847Z digest=sha256:d6a719b0e13d42f976c2917197f767b8b9aa0cb8d05b9a3b91b6ee2ea117e36d

Observation 986ea017-a1a9-423e-a162-80aa191b2015 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.958430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.958430Z digest=sha256:b403ea830245d544ea90786b95eddcf633c27260bb829ad087d45449a186221c

Observation ab00bb5c-ea22-4fe4-b469-420a53fd3470 · outbound

This paper cites Transfer learning in computer vision tasks: Remember where you come from.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Transfer learning in computer vision tasks: Remember where you come from

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.987195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.985748Z digest=sha256:f3755ac80d053a03a1fca35eb45b7d27cf24457fecd42469ec456d94912f54b4

Observation 0e1a228a-284d-4834-acf0-c75ca3829e9a · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Evaluating object hallucination in large vision-language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.971091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.022976Z digest=sha256:d955ed10a3c4d25995a1eb78e602e1ad8b1389036532f16fc4604fed2ce781de

Observation b55b8aa9-9a3c-492d-85ad-173108fd5013 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.058425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.058425Z digest=sha256:2f71c234af463f9009834c32105b1b2d6ae0ef5e95b5b555a4a5d9563c6d64c9

Observation 941bbed5-eb97-4269-8ad7-47c1f58e61e1 · outbound

This paper cites Visual spatial reasoning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Visual spatial reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.943536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.093939Z digest=sha256:fe07c7791f581a22bdfcd05730ef5d217ba430d77939ca2bd521b04f3d67451b

Observation 95581358-a2dc-4a8f-8e01-5f39a1289963 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.130078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.130078Z digest=sha256:6b23913fa12823e8d56443027ec0c11d8b5115e4d217ae46c949c26be9e429de

Observation ad658fa4-a1d9-4e11-88d2-2c6fa4d826b1 · outbound

This paper cites Decoupled Weight Decay Regularization.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Decoupled Weight Decay Regularization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.157884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.157884Z digest=sha256:7a25989d5ec94c7f585c5cca8397f32347e3333507c3ad4a74311874ea1fba01

Observation 2b9b5496-ad13-4cc1-b52b-2c82d1c81616 · outbound

This paper cites Understanding Zero-Shot Adversarial Robustness for Large-Scale Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Understanding Zero-Shot Adversarial Robustness for Large-Scale Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.191544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.191544Z digest=sha256:be68f14302fcb5e723702afe7b325064039fe0cd5dbdfe90d604572f93ec41d1

Observation 97d40f26-43b0-4ba1-b6c8-bbf546f4a1d4 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.214004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.214004Z digest=sha256:72a7a94f0238ddd854717ec98feb5261209e8bb4d5103fe5672828b77a2b81b0

Observation 53d26859-b591-4ff0-8722-4807ae2a4770 · outbound

This paper cites Y., et al.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Y., et al

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.234636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.234636Z digest=sha256:357f85d85465e1a9595fd8aa67c29f6f7b9838c5e40ec02826182ce635478ea3

Observation 76527488-a35d-443a-aab9-e2a81759f282 · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Dinov2: Learning robust visual features without supervision

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.894122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.250170Z digest=sha256:b1d194046263a73c6a2b4abbecfd80bcb881f89c73fc29d692dd8de737ed4ce9

Observation dcf1e64b-a01b-4377-bd5c-3489b89a64c3 · outbound

This paper cites Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.272780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.272780Z digest=sha256:e7071fbb52caebe454470141688b975b1b03d677c366c3d05e2b323266a95f9d

Observation 1a5fae69-ee38-41bc-b5dc-14c10c800486 · outbound

This paper cites Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:15.081496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.289306Z digest=sha256:ae2972240d55b70bcb32422f99fcf2d924ebf28b105bd3ccf7c8809081a2f224

Observation 74d19f35-21aa-471d-910a-247e7ee840c8 · outbound

This paper cites Towards a scalable reference-free evaluation of generative models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards a scalable reference-free evaluation of generative models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.875218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.310303Z digest=sha256:d99a5727f6b6195a6bf623ecbf30a90424dd4003b5be708a1c0c67501eb0bac0

Observation 43bac356-ca4d-42a1-81aa-2d9b908b5fc7 · outbound

This paper cites M., Vedaldi, A., Zisserman, A., and Jawahar, C.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models M., Vedaldi, A., Zisserman, A., and Jawahar, C

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.857821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.315818Z digest=sha256:2fa6f0c46f78e4af74539518acca2d86af08bedc20808a49dc2809cf4fcb9952

Observation 2618be62-08dc-4987-b436-f6c8b8b33a5d · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.327491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.327491Z digest=sha256:a722b697b67a1b096c48a22db640606a2da1d7bec23c821356bb71629b05b8db

Observation 36b7853d-b122-4b04-98f4-4a8a043db632 · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.828165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.337840Z digest=sha256:8f4b81b9261de98d58c53274ea108bb24574327c4a7d465737ff0e6f66805c7c

Observation 918a3279-df71-4ef6-937d-790a1d9df14b · outbound

This paper cites Be More Diverse than the Most Diverse: Optimal Mixtures of Generative Models via Mixture-UCB Bandit Algorithms.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Be More Diverse than the Most Diverse: Optimal Mixtures of Generative Models via Mixture-UCB Bandit Algorithms

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:15.058806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.349429Z digest=sha256:7fcbfcc790a04c908e378b0e43c7be4d19b4502d43d97faf4f57f618405f9f18

Observation 6ac9159c-f5d3-405d-81b3-fcf795583ee8 · outbound

This paper cites Improved zero-shot classification by adapting vlms with text descriptions.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Improved zero-shot classification by adapting vlms with text descriptions

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.809295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.362400Z digest=sha256:4c335452cd9efeb6f8354d7c0e56f303dee518a65111525ee2b0366923dfdbdc

Observation 4e5c131a-6d20-4eff-b873-ca932678061f · outbound

This paper cites CLIP meets Model Zoo Experts: Pseudo-Supervision for Visual Enhancement.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models CLIP meets Model Zoo Experts: Pseudo-Supervision for Visual Enhancement

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:15.035097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.375483Z digest=sha256:c2b9d4d08edbe738d056cb0a4dce874b60eabde52cf7c5f04e59a457d589a390

Observation ebf584df-d99f-4edd-a28a-dc9ee716d36c · outbound

This paper cites Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.409174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.409174Z digest=sha256:d62eb63c208fa7f62fe192ad8a83a9214c93c17303b42d18fd3ef75ae1d5fba6

Observation b6b6a7b7-2294-4b70-ba5a-04f1a26dbd1c · outbound

This paper cites MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.425638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.425638Z digest=sha256:497d18847ff4432452d57761a2e38a1d206618ca50b2537bc3f033a8e87a365f

Observation 73f02e0a-268b-47b9-8f9d-0ec95bc9ce56 · outbound

This paper cites Finetuning Text-to-Image Diffusion Models for Fairness.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Finetuning Text-to-Image Diffusion Models for Fairness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.450325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.450325Z digest=sha256:e6da6f7ff94445040e0e7378a02c9ead5995ea346ee42822fee7802f3e6c6e8b

Observation 5efff50b-5682-40c5-b5d9-0ed382beb549 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.486555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.486555Z digest=sha256:094e3c7553a1b93cfd29b85fe67c7826bc3af1811b41fbfc9fb5b5aa351ad4cf

Observation 86c03fab-1dc6-414c-9dd9-3cf325655264 · outbound

This paper cites Towards vqa models that can read.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards vqa models that can read

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.514075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.514075Z digest=sha256:d7856bd1f268c9da6295b9d0578a6e12b7a99784d21874b75a5c26e4afec2ae3

Observation d5eeb4e4-5a06-4026-845f-44b8f6fe2d19 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:15.779942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.532330Z digest=sha256:c7f09345556e256124a589996bf0c2b893269989c910f09bae02cbdd145f49a6

Observation a43fc90d-4703-49a7-8e49-220310d8585c · outbound

This paper cites L., Taylor, E., and Loaiza-Ganem, G.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models L., Taylor, E., and Loaiza-Ganem, G

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.758280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.567592Z digest=sha256:7b0d45fe6bc4ab8926e14323472026c5c221a3bd944cc7430aded1124aaaccdd

Observation 92360aea-3301-4b3a-8ae3-14ba88afdfd5 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.595926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.595926Z digest=sha256:7a77f06a19bc944901a3d627d1b67b154ca42441df21c9042754385ae933bddf

Observation 75fce3a4-6058-4509-a68a-2ce192512d2d · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.740074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.633074Z digest=sha256:4162e49ca81a5b06985a27ff10b58b242d11deec0aaad78d48306d46571655ba

Observation 8a11379d-27de-4437-81c4-d11d76d335b6 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.668972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.668972Z digest=sha256:8243dc3d36f97d414c23642c1b0889280fb14f2d34fa67d9fd57526f25b34c90

Observation 8958d4da-88b8-4d83-a640-75d27891eb51 · outbound

This paper cites S., Linmans, J., Winkens, J., Cohen, T., and Welling, M.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models S., Linmans, J., Winkens, J., Cohen, T., and Welling, M

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.709139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.688699Z digest=sha256:9158563521236799dd5271fc7d507e095dac9000b3d3207d85e0d84491d1a539

Observation bd473008-39b0-4582-86b0-9cf57d3e5756 · outbound

This paper cites Clip the gap: A single domain generalization approach for object detection.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clip the gap: A single domain generalization approach for object detection

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.691181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.693990Z digest=sha256:b31b20f1a7c70dca355460b14c8f5fa757bba23348cc99f8e66ec6f20f70a834

Observation e5740819-6d52-4e95-8a60-38da9a6d39f4 · outbound

This paper cites S., Steiner, A.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models S., Steiner, A

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.673241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.699343Z digest=sha256:b572867226bc3f2711e91049794264b21a0ba93d65eec05a1aa045694485818d

Observation b7a586b6-5dc2-43c8-bc52-d27d87f85af7 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:15.655991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.708069Z digest=sha256:c24a9bfc95ae52fa6ea380ca3dda8044e715d74b4ebcf73d2eee46d4d9e251b5

Observation 65186142-1aa3-4919-9f62-ca235b46f4c6 · outbound

This paper cites Diffusion feedback helps clip see better.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Diffusion feedback helps clip see better

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.640372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.714678Z digest=sha256:cf76817c822badc8c0e9cf73047b8bc73585807ba4f777f4b6f3a11195fc7d7f

Observation 62d63c46-d754-4f26-828f-ce1d23533941 · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.720443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.720443Z digest=sha256:4da77f1ee979b995bfa4c317cfbb4a186452371835d75e51673efb03089780e0

Observation a43c192a-2ac7-447c-a69d-9cc493a6cfcf · outbound

This paper cites Demystifying CLIP data.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Demystifying CLIP data

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.623105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.725923Z digest=sha256:1bfcff8ebf4b14b91d7c85d6861f460e4ef508b3ba6bda8c55454d5ca844954e

Observation c367537a-4415-49b2-9169-6420db8b0cea · outbound

This paper cites Explicit inductive bias for transfer learning with convolutional networks.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Explicit inductive bias for transfer learning with convolutional networks

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.606495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.731229Z digest=sha256:4bfd92650edc75c35ecd962b3ecae38cfc003b9d8c9883512071901879a68026

Observation 984b5b48-16a6-4047-ac45-5dbea0e582a2 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.587517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.736867Z digest=sha256:cdc2c8a92aafde1fcf0a449c3d826eb7f5dc19e00c85a58d1018267a88256d9e

Observation a6135983-a5c7-4bad-a1fd-8200c9309782 · outbound

This paper cites C., and Berg, T.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models C., and Berg, T

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.570210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.741696Z digest=sha256:f152d11cb71ac7af84bb0e24997ac611a5e9370abf64e8dc3ab11a7188cfa4c5

Observation 74155806-e21b-4baf-881f-117b867164ee · outbound

This paper cites Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.554372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.746368Z digest=sha256:a864b0bd465245ced2cab797987c6e248efde69fd9c29f6047160bc6d41c1029

Observation 215a7a43-1c7f-4c53-bc74-17924b037985 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations, 2023.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations, 2023

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.538506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.752572Z digest=sha256:af6940f57151b001c370c5ea3ca9598845064932a566a0f4e9839e6fb306a08a

Observation 2cb6e895-9f87-4c8d-8c4d-69b22d6ba8ec · outbound

This paper cites Sigmoid loss for language image pre-training.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Sigmoid loss for language image pre-training

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.757778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.757778Z digest=sha256:a42d17c6de46e6ce85107809305645392b16cc7751584799db738cb6f34dee98

Observation 4bfa7c29-6693-46f8-b2ca-bd214c111d9c · outbound

This paper cites Unveiling Differences in Generative Models: A Scalable Differential Clustering Approach.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unveiling Differences in Generative Models: A Scalable Differential Clustering Approach

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.763166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.763166Z digest=sha256:46550c4aad57dcee423af83366da9baa1da7e912a9ee794b1dca19be69dae35f

Observation f147c74c-a66a-443c-a7e1-333c75204a53 · outbound

This paper cites An Interpretable Evaluation of Entropy-based Novelty of Generative Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models An Interpretable Evaluation of Entropy-based Novelty of Generative Models

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:14.871675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.769062Z digest=sha256:b09fcfa55eff6886c072f6e1397a4ba68dccccf840f5291b660152b43285ec78

Observation dc3159f4-8721-495b-b9dd-4cb65b7c68ef · outbound

This paper cites Tip-adapter: Training-free adaption of clip for few-shot classification.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Tip-adapter: Training-free adaption of clip for few-shot classification

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.510047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.774232Z digest=sha256:023e3a8703e5da906a434d8f9109c97d3565a4ef449dc41a3811a6c41f9fbca8

Observation bd3c76e1-ac43-4256-a6c3-cb6b3b7ec446 · outbound

This paper cites H., Zhou, L., Dai, X., Yuan, L., Li, Y., et al.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models H., Zhou, L., Dai, X., Yuan, L., Li, Y., et al

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.489436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.780451Z digest=sha256:c7c7c5a9680ae294ea5862b0c673934577cb1539182056519d4bc92a3c7b215d

Observation aabcead0-936a-4ad9-9929-e5d6cdbdf76a · outbound

This paper cites C., and Liu, Z.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models C., and Liu, Z

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.786001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.786001Z digest=sha256:129dd592a5405542f349dba363c0b78d9038e97ffe264a05ddc32d73320c217e

Observation ef46bbe8-0ec8-4a86-b913-da9bd885d844 · outbound

This paper cites Rethinking Centered Kernel Alignment in Knowledge Distillation.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Rethinking Centered Kernel Alignment in Knowledge Distillation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.790735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.790735Z digest=sha256:0e9cbb0a13abf997b8f237c9cbf5f7dbb28bef4b150c693fd6bbc022c97472c5

Observation 32880931-6e19-4a41-8cc9-ebf2088f0d25 · outbound

This paper cites write newline.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models write newline

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.795852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.795852Z digest=sha256:fe08283c25c1a5d7f76cf71c07c5bee6da0a16ba9881c3c29d60f662458c8d39

Pith citing papers

Observation d62736ac-a819-41bb-a7c7-eb0abf546778 · inbound

UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels cites this paper.

UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:32:52.003486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:32:36.729122Z digest=sha256:a776791377e7b847076808dc8da77d5f17c4a72084d8eb2e1809b3097c20efb2

Observation a3fff1ec-2d20-4ddc-ad54-cc96cf462af6 · inbound

Latent Denoising Improves Visual Alignment in Large Multimodal Models cites this paper.

Latent Denoising Improves Visual Alignment in Large Multimodal Models Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:09:26.700458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T23:07:54.806529Z digest=sha256:8d0d7b0938cec37b346bc64e0b6d29982bc006bc685f92ae618351195997d0a3

Observation 5b84b43d-5e64-4629-a595-4b924cbc06b4 · inbound

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning cites this paper.

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:51.075371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:14:03.809826Z digest=sha256:933f76879d04cfd85962407f1513b56a2e380219725b0b20ca6f89bbf1c554f5

Observation 885178fa-8e7d-4394-a495-a60a3a927002 · inbound

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning cites this paper.

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.135657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:07:35.021986Z digest=sha256:909a580759a2943ea80a26cca5ed0b973dc8778a18d3ab836f95fc4751b832e9