Pith. sign in

Paper Citation Record · LEDGER

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2505.22943.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22943 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:03:27.907091Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd77411f-eeaf-40ac-9e8f-71093929505e · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:32.135976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:24.192467Z digest=sha256:3b01cafad4beccc3bed89dfe90bf6a88d2330bb1aa07fedfa63411ad3b6ada66

Observation 5db80748-a755-4cb2-9b17-b6edff055445 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:31.611148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:24.587198Z digest=sha256:aeca966b3c3d95dcfa41b9c05fbe4412fd25e13770bf0e83abedce275d512eef

Observation 56651823-46b8-4779-a46d-1b7018d143d8 · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:31.880859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:24.293166Z digest=sha256:16a70cf8c478ebc40f0da1eac8f7134c20ccb762a09c0752c1d0b581c641807d

Observation 58d51f95-3331-4538-815e-5b0e1a3ada5c · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:31.396163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:24.670254Z digest=sha256:ba4450f3e190b1a2c126dc1515f927a3b55e74854614e3a23ae3f6665dbd1998

Observation 5d65069e-17ec-4bb4-b5ca-45d5188d3b11 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:31.128502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:25.028185Z digest=sha256:b07ffa78be0252505306ec3f766129d0e1690d0d3752df82b33e9b020d429743

Observation 5c30c388-c76f-45d1-acf5-f15ada91d3e5 · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:30.900311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:25.138186Z digest=sha256:85717ee1d33c8121007f9863ff0824cba5243e171a972b7515940f5aee08492c

Observation 1f698652-5f2b-4cd2-acf3-b4e4b0700469 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:30.665279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:25.477645Z digest=sha256:617cf3d57eb9b0f68d5f4d522a24ad1dc09382d1792394dc27cd0c630d47b2ee

Observation eaf17d0d-44ba-4a1f-8ce8-a6e3eb1e24fd · outbound

This paper cites two" to.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates two" to

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:30.401493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:25.587960Z digest=sha256:00723999f5c40c00046857bf3d640e012b8cd1912ebf4da6de7f9dbc19039c73

Observation 5f03481d-a958-4467-9716-b4e9626181f5 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:30.177243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:25.933385Z digest=sha256:570f96f857d0f31231e19cab130c9bc6ef41f193eadf10ecdbd9533c26c19bb2

Observation f56e1532-cafb-48d5-bcfe-8733701b67fb · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:30.026424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:26.063298Z digest=sha256:26c5735f48677916fda61aa0712ae546c2cb1d6a8ea5abb47b287066b6d324e7

Observation 3e6f94a8-8ec3-4f1f-a9e1-0b143c4c1122 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.890834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:26.435869Z digest=sha256:73c988541c664030394b8965e947166024537a80375497e0d2d339ae302b736d

Observation 360cc445-622b-4b9b-ae47-086087b6188e · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:29.759656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:26.589257Z digest=sha256:832a45b22c616087c2590a651d52ea095df240fe8008bd0ab4edb61ab8f62baa

Observation 90763670-9554-4175-8df5-b4e9b07818e3 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.636168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:26.871455Z digest=sha256:80fab4c06a13cccc013a53dd0d4e3241e44bd201204ae2bd776d21a1070cfa1b

Observation 0648eaf3-d0a0-4439-b498-8f596280b9a2 · outbound

This paper cites woman looking at elephant.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates woman looking at elephant

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.505309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:26.942960Z digest=sha256:c47cfdbcf2a68c7529736eca653751002b2e11b38dcca0fd4b4c8699c3d15c52

Observation 112439a1-5232-4c2d-8a64-da0c18786d75 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.341367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.146824Z digest=sha256:0b5aa471a3d3dde0a53ebad7383d242cec777712ab8669e820fc015a90d74fec

Observation 6cac1fc0-a57c-493e-b09f-064839ae7bb8 · outbound

This paper cites a red apple and a purple grape.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates a red apple and a purple grape

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.213278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.200308Z digest=sha256:11de8549c294db0f3c75a629f36c6a7cf62e4cc5f0f37728645006d63556e8bb

Observation 3b3bd504-1f00-48e4-84cc-87588b9b5f4e · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:32.609322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.270150Z digest=sha256:874fd372c6f155aa07b0ee0ab36123e592c2bf190820acd39c1491583c6bb85d

Observation 478ee17e-cd4a-4e98-9482-dc8752cbd321 · outbound

This paper cites no", "not.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates no", "not

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:32.379731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.339596Z digest=sha256:2f8a4c6178ee95152cd28e4702e50c1941e14827d6eb6d071d60fd5f5c0e7254

Observation 1c154188-28ca-42c0-86c3-a77aa5f3b54f · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.090619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.402973Z digest=sha256:ba6a81ddfbac35ae77747c75ead6baf01da88a46bbfc9062d0c9520c7bbc68b9

Observation b1b27cf8-4516-4adf-a3b5-f9fc878b3fbd · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:28.793263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.535721Z digest=sha256:c99ec26c7f1476e59e9d3b5902f50de8a337e5d55c4b74780189fe719a60dd2e

Observation 50407dc4-ed1f-4d02-984f-e5ef0c0ea697 · outbound

This paper cites Are two sentences contradictory based on the video?.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Are two sentences contradictory based on the video?

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:28.640131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.595774Z digest=sha256:cfcdcf02e4e70f7844f6b8864ffca9d96b1e0218762a4ab8073fa604b54efa4f

Observation 9a438ecd-66f0-4a14-a012-8fac3d5871c0 · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:28.947788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.686828Z digest=sha256:debe831af402e408591e2b7f2c096ac7e48c3cf6f8f14af2fde7127a60f427c3

Observation 4c33dc6f-5c22-4462-a0b3-b717711d64c3 · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:28.490465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.765510Z digest=sha256:e1101a38ef631a7ef3574971bc98a03d73165e14fd8ed81120ce239bdd5ed1fb

Observation 454eeccb-7337-46aa-8a25-72ba1e2e3407 · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 44

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:03:28.340863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.832893Z digest=sha256:469311a428426eec444499a471c43b0648877c2ce14938bca106d770636095e2

Observation 4af1bbe1-ff21-40f9-b9f3-43595f78a485 · outbound

This paper cites Does this image I match the following captionT? Answer Yes or No directly.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Does this image I match the following captionT? Answer Yes or No directly

Reference 2017

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:03:28.184069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:27.907091Z digest=sha256:0d119bec7b86506039565f7b1664dabc20a18d8f7ea61e296be949f880e130b4

Observation 73a6ba4c-b22a-42ca-8bb8-42c084618fbc · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:23.784740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:03:23.784740Z digest=sha256:4863c12f969fcf57b8d5309dd4842c0ef3dfb0563b6b3ce7b9cf7f015e371a7d

Observation 9c05eac7-4813-495a-8713-457e65196fcd · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:32.800619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:03:23.861538Z digest=sha256:c4d11b3ec4cd564b9fa0a7a25a1c51a00dd766db18dac92cc8448301bc981a80

Observation 1a9df769-951b-4c9b-a0bf-5577854c22e7 · outbound

This paper cites The Llama 3 Herd of Models.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:23.704320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:03:23.704320Z digest=sha256:1751be79a622ff39a3b79789eb1bfc4254d4b61e262a093353fb28cf28ea7bf1

Pith citing papers

No inbound Pith citation observations are available.