Pith. sign in

Paper Citation Record · LEDGER

Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2402.12336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.12336 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:38:17.480616Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:59:07.063832Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d2f5efb4-3dbf-442c-b739-0e378715e711 · inbound

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs cites this paper.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.480616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.480616Z digest=sha256:0a0f931bc0b032d714d29d85c5fb44e4bba560a4cc6e0feb1e23a00b0b79963e

Observation 0b0d825d-4486-4c91-b147-d572b13c2511 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 148

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:19.969645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:19.969645Z digest=sha256:46b9ae1cd11af3cd9532a28bdbb19b7f61a3a0292a52efbd4c5be77e323add70

Observation 60a5a2dd-ed63-4bae-94e8-2a741ef83965 · inbound

Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment cites this paper.

Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.555198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:31.555198Z digest=sha256:436c95943c118179156bc71ba38288756f4a1f6192174a1d6fa675c63e091c25

Observation ebf584df-d99f-4edd-a28a-dc9ee716d36c · inbound

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models cites this paper.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.409174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.409174Z digest=sha256:bd1f0ef0b1cde360634f9474a218b349ce35ab1d76f640274d2fcb81a9fd960d

Observation 7bf6b250-1b9c-40ed-92e4-0b3daabdcbca · inbound

Diffusion-based Cumulative Adversarial Purification for Vision Language Models cites this paper.

Diffusion-based Cumulative Adversarial Purification for Vision Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:59:19.962834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:59:19.962834Z digest=sha256:42ad2d567a56dd9621d746348bbae91a60af6baad19a905de5dee63004836f3e

Observation 0f6951c7-d207-492e-a4ee-d99974500a89 · inbound

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents cites this paper.

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:45.196277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:45.196277Z digest=sha256:b89e72b4206e113f40081fd8d392eb4a5481d1bab32bcf2f1a479c2e19731461

Observation 40c52470-0a62-4685-8bf2-293a3fe76fa5 · inbound

Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models cites this paper.

Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:59.332470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:59.332470Z digest=sha256:3f20704ac6a77f9edab61f13330d38184e78b6717050dbf6fcbd49ef8119b323

Observation 44628e35-95c1-46cb-a172-88e77b633620 · inbound

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding cites this paper.

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T11:54:43.232843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:54:43.232843Z digest=sha256:69ddeb5db967b29a0d8d6bbcd60a47081e2c62a66f133130676900a040fefc58

Observation 42f5acdd-0d3a-4206-9f1f-8a3e676722b5 · inbound

Beyond the Textual: Generating Coherent Visual Options for MCQs cites this paper.

Beyond the Textual: Generating Coherent Visual Options for MCQs Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:37.853389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:17:37.853389Z digest=sha256:0707db80582e2db27ee91c54e9b3baeea6ea50bd7794d0658dacd3255fa4ccc9

Observation 9d741b0b-7ea8-47c8-b0a1-15ab48bef964 · inbound

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models cites this paper.

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:26:33.866893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:23:13.318310Z digest=sha256:6f2a3f124be62e7cebc3d72f14c7444437d58a66fb2971286c2857229ca08ef1

Observation 1f6c570e-8d54-4f01-bab1-e703eb2c4c22 · inbound

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models cites this paper.

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:30:39.404154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T21:28:59.898029Z digest=sha256:bccb5f6c8828a865eeff4ac183a1cfb9498537f3be558521f8868af5cbcb12a4

Observation 79c270ba-1df5-424b-9319-3061f87632e6 · inbound

Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting cites this paper.

Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:50.528290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:50.528290Z digest=sha256:7e5c15052360b81c60ebbc4ec9933d559a9d51670e4ae0a4417c549c91cca751

Observation 07b7e299-dd12-43fa-9994-0c7e0a87f8b5 · inbound

Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings cites this paper.

Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:19:00.649555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T04:15:38.632039Z digest=sha256:338bb2774eceac7a7682e8f7ac141f8016a8de4ae255dd34c257dc65f2d7c301

Observation 47352615-a4e4-4a93-941d-2173c1eb5178 · inbound

Pay Less Attention to Function Words for Free Robustness of Vision-Language Models cites this paper.

Pay Less Attention to Function Words for Free Robustness of Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:33:44.684493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T00:33:13.977360Z digest=sha256:fc704e50d6b25c978c0222d3387b672cca1d67b4d8175302f406ba6941274f33

Observation 18c41af6-dc16-4dfd-9daf-0e3ec428f094 · inbound

Hierarchically Robust Zero-shot Vision-language Models cites this paper.

Hierarchically Robust Zero-shot Vision-language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:31.436342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:25:24.008261Z digest=sha256:db0514c90ecc6f016bd6ddd233a581acdb5ef28aea897e2a43e0eb0213f56047

Observation 2de6a4d2-d14f-4c8e-a78b-2f7dec54061f · inbound

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models cites this paper.

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:08.096327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T14:24:48.999632Z digest=sha256:3937e1e676363a40ce706108e10942bdbac18a2e92e669922d5a8c75c1129abe

Observation bf82e486-a7ba-410e-829d-9b0784ee6381 · inbound

TARO: Temporal Adversarial Rectification Optimization Using Diffusion Models as Purifiers cites this paper.

TARO: Temporal Adversarial Rectification Optimization Using Diffusion Models as Purifiers Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:26:30.625340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:52:25.317140Z digest=sha256:0caf49720238fa227bde35dcec382fe725ad08553037b3d27a5727e9a70e4991

Observation 8fad4a87-bd7c-45d5-90b0-8083bf4bd21b · inbound

AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models cites this paper.

AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:38:56.014305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T19:37:09.983024Z digest=sha256:9ba9c5ba992c02885631f0a38b11c33926c38dba9a4a4ebf7d63effc18ac522a

Observation e2c317cd-cc9f-4d88-a85a-0070af298475 · inbound

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models cites this paper.

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:44:01.424531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:40:26.803098Z digest=sha256:a007cf85a45f9bb47464a183f037d450793ff68ff60583e27e535dfcebb04807

Observation 88d00781-61dc-4169-be5e-9483b0047586 · inbound

Investigating Adversarial Robustness of Multi-modal Large Language Models cites this paper.

Investigating Adversarial Robustness of Multi-modal Large Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:27.643485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:11:34.152223Z digest=sha256:39213f5718b1ac99356d91a7ff115294863ba014127db5d311f028e8c194d1d7

Observation 8c1a1912-7c5a-4c20-9366-be2a7de70aca · inbound

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models cites this paper.

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.844167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:04:30.654255Z digest=sha256:947952209f318af7888251ba825b763750cf1893840671c022c30df05e1e5fed

Observation 15bcbbc8-c1d1-4e41-b3bb-1cf236b83bef · inbound

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models cites this paper.

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:16:34.811575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T10:10:38.261635Z digest=sha256:737d092f7f3ef9e1426f5fc5ecae7b85fa63c77912269c1f936dd5d5e7b8a4b4

Observation 89952cc8-2263-4c1c-b85d-c46a9375ba8c · inbound

Semantic Robustness Certification for Vision-Language Models cites this paper.

Semantic Robustness Certification for Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:59:07.065828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:32:54.522220Z digest=sha256:6b53a6b6b6dc67220d6b24c1b5fe4462f70c759d85293ddb97f738b4bdcd0cc4

Observation 538d2546-01bd-4239-a8ac-34713d4da38b · inbound

Rethinking Brain Decoding with CLIP: The Role of Adversarial Robustness cites this paper.

Rethinking Brain Decoding with CLIP: The Role of Adversarial Robustness Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T04:24:39.660564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:24:39.660564Z digest=sha256:714a10c7463b79a3d9464071174c5743f6806d8264f9a257891ac19c5e94a965

Observation 4ac2d30b-6440-46da-a1cd-4dd36c0c6ef1 · inbound

A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning cites this paper.

A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T01:16:40.492300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:16:40.492300Z digest=sha256:9f29b41d153a0f6dd82be456bc2bbb1bf70f64a065ebfb7bce6a65907096e06c

Observation 45b3fb0f-1089-4e95-9caf-ba4113ac01b8 · inbound

Unifying Adversarially Robust Model Experts in Vision-Language Models cites this paper.

Unifying Adversarially Robust Model Experts in Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T23:31:48.797271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:31:48.797271Z digest=sha256:246a681c51596b771cecaff8e6735dd54ea1db0025424f50b89a4f1c8e839d34

Observation 8da3932e-745e-4220-96c4-9f02a0a37b5f · inbound

Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models cites this paper.

Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:18.042445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:51:18.042445Z digest=sha256:984ba72fcb0a53151b536be9b66e6ec626f412cc04d9f37d5150ee40435d21fd