Pith. sign in

Paper Citation Record · LEDGER

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2505.22943.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22943 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:03:27.907091Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd77411f-eeaf-40ac-9e8f-71093929505e · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:32.135976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:24.192467Z digest=sha256:ec9da6cc384fa3248cbdc7185d27b400a2df82a1f3929e7dd2cb954d13a8f930

Observation 5db80748-a755-4cb2-9b17-b6edff055445 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:31.611148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:24.587198Z digest=sha256:816e2206a114ae66687441b5a477d3ea950b189943e8e27ea7d304ceef280fea

Observation 56651823-46b8-4779-a46d-1b7018d143d8 · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:31.880859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:24.293166Z digest=sha256:89ae59d579bf204798e129abe077183cba83a16485b817e2e267783cd4c82dcd

Observation 58d51f95-3331-4538-815e-5b0e1a3ada5c · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:31.396163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:24.670254Z digest=sha256:38b9110302ba94d708a3878d11900c0681823041abae142231096e777febe7cd

Observation 5d65069e-17ec-4bb4-b5ca-45d5188d3b11 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:31.128502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:25.028185Z digest=sha256:29d3b6d70b52f593e659cddf2c1cf6f9c6a9d820da8043b9d859f2142823c749

Observation 5c30c388-c76f-45d1-acf5-f15ada91d3e5 · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:30.900311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:25.138186Z digest=sha256:3a378f2c2b4d320247fa0b0157d629a29f87bd24e949aa4fe8e05ac498825f1f

Observation 1f698652-5f2b-4cd2-acf3-b4e4b0700469 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:30.665279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:25.477645Z digest=sha256:428ec3d24c9b604118faa07ad72b5c32c97de7ef01be6625768594ccc5ed219e

Observation eaf17d0d-44ba-4a1f-8ce8-a6e3eb1e24fd · outbound

This paper cites two" to.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates two" to

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:30.401493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:25.587960Z digest=sha256:9bd20857ebe3e495fc50da92f84b87207c402884593fc7ea70db4492f702bf24

Observation 5f03481d-a958-4467-9716-b4e9626181f5 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:30.177243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:25.933385Z digest=sha256:b53b03a51bdfeb23d094e0ab3768e586daa9406cca4e5b545010441dd0e29747

Observation f56e1532-cafb-48d5-bcfe-8733701b67fb · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:30.026424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:26.063298Z digest=sha256:f3d547792871b7b3a417deebb87a9904e909b3fd368b4db7623f1697521495c4

Observation 3e6f94a8-8ec3-4f1f-a9e1-0b143c4c1122 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.890834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:26.435869Z digest=sha256:206b828ec8594af51d63310d702f92b3616b17d1d841b9595c2627a41b0eafc4

Observation 360cc445-622b-4b9b-ae47-086087b6188e · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:29.759656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:26.589257Z digest=sha256:a456ce542597931faae3e9a51f7c2364b5205cac99fb7964229adc5526002679

Observation 90763670-9554-4175-8df5-b4e9b07818e3 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.636168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:26.871455Z digest=sha256:c7a651080b5ebc42d856ddb0dee7dbfe0ae573380e80f16e823f18774d2d8820

Observation 0648eaf3-d0a0-4439-b498-8f596280b9a2 · outbound

This paper cites woman looking at elephant.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates woman looking at elephant

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.505309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:26.942960Z digest=sha256:9d369cb311abaa915e81c67bd7b90e7571ad639465d503198bd202fe18b9478d

Observation 112439a1-5232-4c2d-8a64-da0c18786d75 · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.341367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.146824Z digest=sha256:55cc9e187f487b6a2e084cea95705291155c7ac0da8e3d44d52d24701a00347e

Observation 6cac1fc0-a57c-493e-b09f-064839ae7bb8 · outbound

This paper cites a red apple and a purple grape.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates a red apple and a purple grape

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.213278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.200308Z digest=sha256:c8adfb50e19c6f7d656b54377ad262a1e451a3a3dde26d2092ecf991b522ec9e

Observation 3b3bd504-1f00-48e4-84cc-87588b9b5f4e · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:32.609322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.270150Z digest=sha256:7509623d0c972cf240f57c52c55c16c15646f298d2d109245365ad334f1e9ba5

Observation 478ee17e-cd4a-4e98-9482-dc8752cbd321 · outbound

This paper cites no", "not.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates no", "not

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:32.379731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.339596Z digest=sha256:57b42dd83c1b0a399cd5a0190c502ff22c5a607b023262bb5674b579ae6f55b3

Observation 1c154188-28ca-42c0-86c3-a77aa5f3b54f · outbound

This paper cites Generated Caption:.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Generated Caption:

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:29.090619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.402973Z digest=sha256:ca47dff934d24726cac689d3a43e59d9256a2971f9c2b3bd3d76313c0aec5f5f

Observation b1b27cf8-4516-4adf-a3b5-f9fc878b3fbd · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:28.793263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.535721Z digest=sha256:0dcf13bd0eff9810bb49dc889b240f9c41ead62312712b25c87ba7e3d40e944c

Observation 50407dc4-ed1f-4d02-984f-e5ef0c0ea697 · outbound

This paper cites Are two sentences contradictory based on the video?.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Are two sentences contradictory based on the video?

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:03:28.640131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.595774Z digest=sha256:acc25d381b65550e5cfbf8de94a0fe13c0f71e1e2a242cf4fe6642ee98f79275

Observation 9a438ecd-66f0-4a14-a012-8fac3d5871c0 · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:28.947788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.686828Z digest=sha256:f04cc30ff1b6156db62188cefe2bf6151e98f0310602a2ef73ce8143bb9639ff

Observation 4c33dc6f-5c22-4462-a0b3-b717711d64c3 · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:28.490465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.765510Z digest=sha256:62dc36bc3a9a2694473ab054e4fc38830d13bf635bdfb95b69c14b703c0adba5

Observation 454eeccb-7337-46aa-8a25-72ba1e2e3407 · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 44

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:03:28.340863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.832893Z digest=sha256:7b3f30c5e0ac7c3303bac7a1ace339ff142a808cf19c6956e2eecc38195a6266

Observation 4af1bbe1-ff21-40f9-b9f3-43595f78a485 · outbound

This paper cites Does this image I match the following captionT? Answer Yes or No directly.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Does this image I match the following captionT? Answer Yes or No directly

Reference 2017

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:03:28.184069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:27.907091Z digest=sha256:2b535a86f7bf69ff7601c397229ce5c95d4052e31d13fecca3966c4acaf307f3

Observation 73a6ba4c-b22a-42ca-8bb8-42c084618fbc · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:23.784740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:03:23.784740Z digest=sha256:4863c12f969fcf57b8d5309dd4842c0ef3dfb0563b6b3ce7b9cf7f015e371a7d

Observation 9c05eac7-4813-495a-8713-457e65196fcd · outbound

This paper cites an unresolved cited work.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:03:32.800619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:03:23.861538Z digest=sha256:3f47f43bb51e5ba7700175d1930f92f3a3420028dd0d700b512f7953e416c0c8

Observation 1a9df769-951b-4c9b-a0bf-5577854c22e7 · outbound

This paper cites The Llama 3 Herd of Models.

Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:03:23.704320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:03:23.704320Z digest=sha256:1751be79a622ff39a3b79789eb1bfc4254d4b61e262a093353fb28cf28ea7bf1

Pith citing papers

No inbound Pith citation observations are available.