Pith. sign in

Paper Citation Record · LEDGER

From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2310.08825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.08825 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:10:13.086299Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:58.025162Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c5cea414-c347-43ed-8d89-640ad01738a2 · inbound

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models cites this paper.

Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T20:27:19.970450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:27:19.970450Z digest=sha256:20e3625328e8c4c9e0f23fe5af9ae79ddb38b51ebc7b77f64a8c13cef934241b

Observation 87f820bf-13af-4c63-9f8a-c55ea640ef73 · inbound

Libra: Leveraging Temporal Images for Biomedical Radiology Analysis cites this paper.

Libra: Leveraging Temporal Images for Biomedical Radiology Analysis From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T10:20:44.821297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:20:44.821297Z digest=sha256:77a86b557b4a9882ba6b8bf976d61b67004731f4f619c0ab5d14302faf1ccce4

Observation d7a15a1a-c8b9-4686-adfe-973c94009a05 · inbound

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding cites this paper.

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:14:09.681095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:14:09.681095Z digest=sha256:b0352f583499c4ef4bb6d4e94a4c2ba5409278c33f959fd0e0837b62fa4a1ca1

Observation b45d62d5-a49d-40b9-925b-a6e2f82afa10 · inbound

CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology cites this paper.

CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:21:46.527734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:21:46.527734Z digest=sha256:fd474610296f2dd29875aaadb477a2580d4610b6e7721671a1f49427cfc82f8c

Observation f191937e-84f9-4da9-bf32-adca4d8e3361 · inbound

ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing cites this paper.

ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:47:10.243345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:47:10.243345Z digest=sha256:96379c2227cd3aa0944e1aa885a4f7caf24ca42bc4af3c4c3d9fcf4ccdf52842

Observation 4bd58b9c-39b1-48e2-ad57-c9350f2b7e0d · inbound

Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models cites this paper.

Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:14.324383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:14.324383Z digest=sha256:e1c555c2ec0bc281b7433a38303630f49e35746c0f72c983a5092a3dd61b9e8c

Observation c7b5b60e-23e2-47cf-953a-30ef807b3ed0 · inbound

Diffusion Instruction Tuning cites this paper.

Diffusion Instruction Tuning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.875384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.875384Z digest=sha256:1762fe1cc953d3d334db5850ba2fa511577d41df181d5b976594be747cfce19f

Observation 1b2c9241-8d5b-4c41-983b-dcd4728b5c64 · inbound

Toward Generalizable Forgery Detection and Reasoning cites this paper.

Toward Generalizable Forgery Detection and Reasoning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:27:12.473517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T22:26:49.062978Z digest=sha256:14527f72c42ec6149d9ba0bb29d7d7fc1f80582e7f7a5effcd355c5b0aba9414

Observation 4372cebe-1cb5-48a4-8acd-b28cb473a856 · inbound

RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction cites this paper.

RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:12:08.638336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T21:09:07.088042Z digest=sha256:89c27b3fedd364ae901ec83b2714adce2c2351762fce647bbf7bf7f57f04d38a

Observation dbba5313-2aad-4768-8d8a-e44aa9ffa657 · inbound

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models cites this paper.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.676257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.676257Z digest=sha256:6ada38938ae28a381b6e1f9cbb0cd1be81380fce597bc2324c233c0d6f4e7f65

Observation ea4cd2ef-2034-4a1b-8b7a-f9b908f0b8ea · inbound

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor cites this paper.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:31.897432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:31.897432Z digest=sha256:aee05d9b20f3ce0b99904f2c5046abf72dd9a5ee350ee91a8f14f67e0bd6973e

Observation 31d9f30d-3254-4f1c-8878-c08df639758a · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.907280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.907280Z digest=sha256:48fe643b4180e1ddfc9985ac8b7b98513b4952cc203be6bc838e40a0a3277859

Observation 6d5705cf-9e95-4c13-baa3-b7bdb6ecc9bd · inbound

CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation cites this paper.

CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:00:04.849995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:00:04.849995Z digest=sha256:24c26c55e99e3751bdba44339323b0fcaead7c6cbefe5d2d5378ea5852beb76e

Observation 7de859b7-0815-4f8a-b1e1-53695a0ea7af · inbound

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework cites this paper.

HEAL: A Hypothesis-Based Preference-Aware Analysis Framework From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:35.340576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:35.340576Z digest=sha256:fc89ef6e56c902dd219f305d5f8869f17f5b74cc705357335d80235444e695f6

Observation 6c61d66c-4586-4d15-846f-be83e0498925 · inbound

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection cites this paper.

Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:33:18.602691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:33:18.602691Z digest=sha256:50cda79fefe99016d3cc07c928fdeb9adc5ed0482513035f50b660ce41bca508

Observation a92a6dea-f979-4319-b35a-d63074e2910a · inbound

TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank cites this paper.

TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:40.129707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:40.129707Z digest=sha256:da2a234b58a6f068d8aebda51c09b87890e16d9cd608124e7405457fd9f52fc4

Observation a54931fd-ea48-473d-a5de-b2c9e84bb2d0 · inbound

Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations cites this paper.

Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T17:26:38.328095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:26:38.328095Z digest=sha256:579416778215c350a92466e6e1709d2cb7c7d7f30fa5b4a1975dcdfab0a32aaf

Observation 41b63a2f-a473-4cd1-a457-f6f7c807a93b · inbound

NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion cites this paper.

NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:00:29.744034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T07:58:25.211030Z digest=sha256:2278ea240fff4c584778e8fccceef3b30d37fc195e078c23e83c4fdfef25bb39

Observation da0fb4f2-cd52-41d2-a5f9-f584920f3438 · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:03.179935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:03.179935Z digest=sha256:37cb85009a865815ed5b9b15badea8a26b28577e5439114f2e033e73abf91d6c

Observation c2906415-2ba2-4c97-b381-bcdb3c75e281 · inbound

Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks cites this paper.

Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:21:17.260679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T21:19:37.970460Z digest=sha256:e43e933f5b9f4bdd61b389ca8f0801dbcea8f0b29778c1f6a657354222840752

Observation 008c6eab-522e-4181-8713-264b2bb71c45 · inbound

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models cites this paper.

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T09:37:07.017396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:37:07.017396Z digest=sha256:0432bdea8a63f9528b7b76d75c1c11582d1b2ab0225f7dfb978cedd2e8792113

Observation f10a2e3c-0f71-4694-8865-08eabaf7a67d · inbound

Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference cites this paper.

Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T04:37:27.791489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:37:27.791489Z digest=sha256:c1177f66022a7b974bf26dfa5fa5045b10a837c08c33a329df4db48c1a73f521

Observation 7b8810df-b7bb-4a6a-85ec-1bd91ca6ba43 · inbound

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders cites this paper.

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:11.852907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:11.852907Z digest=sha256:6a188e0162e47bee1e1a3e02f581dd9b043af0e5835f566546a86e684a261fb9

Observation bba09ca1-9f4f-42cc-8fde-16e58a55610e · inbound

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation cites this paper.

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:10:15.922063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T19:07:46.733450Z digest=sha256:06235aef0a844fe088aad6225098b255ff6056dcdba3fb45f4d58480363c823c

Observation 14df1184-58e6-4ca9-a7e3-8590e06f4864 · inbound

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation cites this paper.

InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T20:35:05.020062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:35:05.020062Z digest=sha256:ac5537d198fc93c1e6cc586631d2ed755034dedce05bef0bbf31922fe0c8b353

Observation cfd77212-2fac-4242-bb44-7c4465f2cec7 · inbound

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning cites this paper.

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.602770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T22:04:18.591594Z digest=sha256:52eedcaa598b77311f12b42ad837a5534cde1580a7999b8e5c582e7776d9bfef

Observation 6ab8a71c-97e8-40cf-87fd-de9bf6a49861 · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.149321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:fdb587dbf13e1b9011c1cc93d892f080263754105be7dc162c7d113748e92b7d

Observation 639975b0-9fe4-424d-9db0-ddbfa385ee24 · inbound

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models cites this paper.

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:01:05.567231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:20:19.653806Z digest=sha256:d167fafe7441f5307d58e871ab3eae9d7a0d53ab26d36b6ddfc3e9edf7bfb2bc

Observation a3803d08-b0de-4b1f-a851-f4046f935b27 · inbound

G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval cites this paper.

G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:50:19.964062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T11:50:00.850857Z digest=sha256:25f506efe711d5e0ff4c2193c740621ae8dfbe56e10a80cacb44a6d954d30d47

Observation c38471b3-d21d-4cb4-abf8-5df62635316d · inbound

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding cites this paper.

Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:01:49.307858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T06:58:19.094492Z digest=sha256:b85cc1813161fbc8dfe7d25d4c72a8d960ac5099462cda9bb29255bf8f44ac44

Observation 3680a95e-b128-46bf-88c7-600e96d7739d · inbound

Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction? cites this paper.

Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction? From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:11.344431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T09:08:28.326598Z digest=sha256:2d3600c46c333f3fc146a171124e094dd10edcba3cff21ab0da4447279d7fb01

Observation 58db8de2-4657-4f0c-b486-b091a2229339 · inbound

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety cites this paper.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:00.953629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T16:00:32.413225Z digest=sha256:dfed8fd5d8054e9ab0ff89592d0c4c0183db609c753886ec1c99657fb86f90e0

Observation bcc9241a-18f3-4ff0-848f-ec5f2df21fe7 · inbound

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks cites this paper.

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:17:02.333089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T01:13:55.067785Z digest=sha256:66674e9d046c6a494721eb67ac32070a87ce55ec9db91854c676c9abb7e96e4c

Observation 02cfb42a-25d1-4b9b-bb4f-5c3f5c1a48a5 · inbound

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks cites this paper.

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.405020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-30T23:37:21.982973Z digest=sha256:74628d6fa11645a29aa5ff28fb8153621c0b8d89afbca8a41b053a88657d9f57

Observation 507e19a5-b352-4bdb-a469-82e92d246550 · inbound

New Wide-Net-Casting Jailbreak Attacks Risk Large Models cites this paper.

New Wide-Net-Casting Jailbreak Attacks Risk Large Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:48:23.297324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T14:46:52.070071Z digest=sha256:13c9054e8c6dc1df094e38fa95de7d574ef75bdb491e31f2a9fd51b28f6a9888

Observation 3c2bf51d-b6ca-4524-8d43-f8951ff71cf6 · inbound

Mechanisms of Object Localization in Vision-Language Models cites this paper.

Mechanisms of Object Localization in Vision-Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:48:05.803191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T06:46:20.141227Z digest=sha256:742019d7c96b6dc35c45b60c8278fc3c655474bf074c52e9afbca1c3ebf8e810

Observation fdd383ec-bbde-4b1a-965e-ad04a6bb396d · inbound

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens cites this paper.

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.510067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T05:07:12.965100Z digest=sha256:9ca6cddf8aa26f9335e9bbdcbacb9592fb86b808f622422394d8dc1f8c333377

Observation 817649ba-9fa7-452b-af17-eb8b87251fd9 · inbound

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation cites this paper.

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.271586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T13:33:12.116833Z digest=sha256:e6e7299f0824981959bb1bdb2b5cf2dac39f610925c93443a93eb0ec953c1662

Observation 55e98e71-8362-4592-b03c-286675d156b2 · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:59.346516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:55446d280cc19f5df021b3ff0310dca60419424475d43f7607e6bb90a9b09812

Observation 2687d8ab-1295-47af-b1ca-0fb12f8f773d · inbound

Investigating The Security of Modern AI and Cloud Infrastructure cites this paper.

Investigating The Security of Modern AI and Cloud Infrastructure From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:42.327590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T11:31:39.910784Z digest=sha256:a3b5f9298cbac1a8ffa1c2a74844193682466993e62f92cfbb05c51c2920601e

Observation 4ad0ef00-57c0-48b7-bbd6-c519d2da434e · inbound

EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL cites this paper.

EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:05:36.899599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-01T08:57:37.907354Z digest=sha256:80c29b757323d78586292d6806adb974b38a203bb8f7233825790b3670ae2af2

Observation a31b896d-ed4e-4428-8bb3-3714b9e5800b · inbound

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference cites this paper.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.026479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:05015354086e3c10535935cfd945737ea67108cb442a9cff3927d74abedf4c08

Observation 309ad413-30c8-43a7-b124-7c7ef35c92d5 · inbound

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning cites this paper.

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:10:13.086299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:10:13.086299Z digest=sha256:ad38d014f481a3a841bfc01dbdbf2654262ef2105c34fe96f3aafce12c4d6da4