Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Models Do Not Understand Negation

As of 11 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 8 inbound Pith citation observations for arXiv:2501.09425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09425 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:08:46.326666Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:58.344234Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T00:18:29.517494Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved11
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 876019cf-4e11-484d-93b7-b82286f09e37 · outbound

This paper cites Aug- mented reality meets computer vision: Efficient data gen- eration for urban driving scenes.

Vision-Language Models Do Not Understand Negation Aug- mented reality meets computer vision: Efficient data gen- eration for urban driving scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.570497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.008226Z digest=sha256:3d758322e273871cf55d0134286bf6ae0f4290f27d170c8bad64d4c982e494f2

Observation 76626a12-b8c4-4540-861e-c24ddca9336b · outbound

This paper cites Effective conditioned and composed im- age retrieval combining clip-based features.

Vision-Language Models Do Not Understand Negation Effective conditioned and composed im- age retrieval combining clip-based features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.548229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.014749Z digest=sha256:e33c827e991ca7e960a6db4caa6648421f4cffeaa443370e00645af97f036cce

Observation 5d08414e-b6b6-492a-ac86-dbac81d7a1fe · outbound

This paper cites FitCLIP: Refining large- scale pretrained image-text models for zero-shot video un- derstanding tasks.

Vision-Language Models Do Not Understand Negation FitCLIP: Refining large- scale pretrained image-text models for zero-shot video un- derstanding tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.517244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.020123Z digest=sha256:860d8c826f56af1a319cbffca94d08ef15b5e57231f0b5d540a98bf9876d918a

Observation adc40a9b-4ad3-44b7-96a0-b98ec2b7ff5d · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Vision-Language Models Do Not Understand Negation Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.031905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.031905Z digest=sha256:b3235bf29b730bba120fbb817b3ac68e98081ab08dd1e970debad589902dde15

Observation 19db9dd0-2961-4ed7-ad64-5bae5646be84 · outbound

This paper cites Learning semantic segmentation from synthetic data: A geo- metrically guided input-output adaptation approach.

Vision-Language Models Do Not Understand Negation Learning semantic segmentation from synthetic data: A geo- metrically guided input-output adaptation approach

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.453953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.037424Z digest=sha256:8874ebbfb6faf548ab3b137250dfdb2f9f38d5a4fdd909b6c32fc55cc9454058

Observation 3e0cc5c6-38ab-4ba6-855b-8a80c4030f3d · outbound

This paper cites The Llama 3 Herd of Models.

Vision-Language Models Do Not Understand Negation The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.042836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.042836Z digest=sha256:15a0a0c85327708e90c672349dea7b71f841f6bcda8fdc8142f4175a518b488f

Observation 97d27b26-b7c2-45d5-ab5c-dbe351a3529a · outbound

This paper cites The pascal visual object classes (voc) challenge.

Vision-Language Models Do Not Understand Negation The pascal visual object classes (voc) challenge

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.429941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.048534Z digest=sha256:458df60a1a2b49f9f2e3b9e67df1705fe790f23ed3c216d926c5105048b90a22

Observation 11613278-2c26-4b96-9ae0-c64f4c515b2a · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

Vision-Language Models Do Not Understand Negation Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.054919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.054919Z digest=sha256:7c77bc4a60401558073fb55135eef20fe783dff6b6c7e5fc16486690af0a66b3

Observation 43de77e5-ed52-44ba-8ed4-68d1bedef7ab · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Vision-Language Models Do Not Understand Negation Datacomp: In search of the next generation of multimodal datasets

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.402544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.061090Z digest=sha256:f9e1e7e24ca11635ffe0cfe1d80ee2c0bb96e70cf96bc8c52f16a45a229980b3

Observation e1117c67-0df2-4e42-a899-fb9d4b5c7bcd · outbound

This paper cites This is not a dataset: A large negation benchmark to challenge large language mod- els.

Vision-Language Models Do Not Understand Negation This is not a dataset: A large negation benchmark to challenge large language mod- els

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.382504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.076557Z digest=sha256:a9bbd3075adf9fb7618c5ba604ffa58ef450d636a7b16dffe4eae5dd7546bb5c

Observation 504dc068-fbb3-4026-b1d4-87f3c4e9d131 · outbound

This paper cites Shortcut learning in deep neural networks.

Vision-Language Models Do Not Understand Negation Shortcut learning in deep neural networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.355708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.081422Z digest=sha256:18a22e16edbc5d89598f3331ed4c5c4b8ca70e2b322587d964a685986afb4876

Observation 6ad3e1d3-2a2b-4509-be2f-5b1e12a3db5e · outbound

This paper cites SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?.

Vision-Language Models Do Not Understand Negation SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.086165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.086165Z digest=sha256:f2fb4155f2e23131e2dfd83e8dff1206147c3415398924a869157fd21ffaa22b

Observation 863bde2d-1a1e-433d-bfcd-09a3e99ecd6e · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:08:47.335828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.091030Z digest=sha256:0f09c8ee6ef82a2c0e3187792079f2059b7909d4e1110506fad4acb5e18e2e7c

Observation fd484441-67ad-4796-b2bf-0cc61cb951c7 · outbound

This paper cites Quilt-1m: One million image-text pairs for histopathology.

Vision-Language Models Do Not Understand Negation Quilt-1m: One million image-text pairs for histopathology

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.314683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.096959Z digest=sha256:0824b60569d92285995dee52b6db84b35ec3a68a4d8a443fb91c79ad04561f2b

Observation 73531a14-883c-4bee-842c-7c0fc3a33525 · outbound

This paper cites Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison.

Vision-Language Models Do Not Understand Negation Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.296473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.103106Z digest=sha256:47f69d4344b62b4715e6e9e8e960327b06b42c8e0db3ff5a504a882a076273e3

Observation a0594ace-e525-4638-b328-70feab20beaa · outbound

This paper cites Generative models as a data source for multiview representa- tion learning.

Vision-Language Models Do Not Understand Negation Generative models as a data source for multiview representa- tion learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.277193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.108782Z digest=sha256:f8fcfa3eab9af04342ea70d5d0c8c507864f7d4c0721fa108f96dd92b46132f2

Observation 6825bbb5-161f-4b27-8c15-85c4acb02932 · outbound

This paper cites The power of negation in english: Text, context and relevance.

Vision-Language Models Do Not Understand Negation The power of negation in english: Text, context and relevance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.257383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.113421Z digest=sha256:ef1c6d7c2c61bde12118b100291ede4efa5aacb602c667c7585d0834e067051b

Observation b40567b7-2ebc-44d1-a989-c1c016b6515d · outbound

This paper cites Negation in syntax–on the na- ture of functional categories and projections.

Vision-Language Models Do Not Understand Negation Negation in syntax–on the na- ture of functional categories and projections

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.234939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.118134Z digest=sha256:58def4e1e794dd54b58565d0295d861e7913083cd2bc42729549f0330cf5b0ce

Observation e4b8857b-2a24-492a-8b01-221b1f34db22 · outbound

This paper cites Naturalbench: Evalu- ating vision-language models on natural adversarial samples.

Vision-Language Models Do Not Understand Negation Naturalbench: Evalu- ating vision-language models on natural adversarial samples

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.214576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.122815Z digest=sha256:ad1179d8d099c05da3789f4dfae5da60c1a9c0588068acc90f776beab649bd77

Observation 51eb013c-e984-4379-8ec6-8ce2bcc29324 · outbound

This paper cites Compre- hending and ordering semantics for image captioning.

Vision-Language Models Do Not Understand Negation Compre- hending and ordering semantics for image captioning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.194283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.128486Z digest=sha256:9112d320df4554c05e7f0e27a987a82f1b43cfdd2a344ccf8c9e5f7c8f5fb2ae

Observation e2d73bba-3d18-4099-bf07-0007c3bfce36 · outbound

This paper cites Cross-modal retrieval and semantic re- finement for remote sensing image captioning.

Vision-Language Models Do Not Understand Negation Cross-modal retrieval and semantic re- finement for remote sensing image captioning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.174919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.133410Z digest=sha256:fe3ba2d63363a2a91de39b6291d923e93aef98e022d62b84870e2463088cb4ca

Observation 9c371f82-ffe9-4c86-b92b-0595593f0e24 · outbound

This paper cites Microsoft coco: Common objects in context.

Vision-Language Models Do Not Understand Negation Microsoft coco: Common objects in context

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.156422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.138330Z digest=sha256:23f67b72af3960376bcb0678235fd0a0f8a1fcbddb00c9a56c69753dff09b3f9

Observation c376ff10-6159-490e-b8f8-24a60b6a017e · outbound

This paper cites A visual- language foundation model for computational pathology.

Vision-Language Models Do Not Understand Negation A visual- language foundation model for computational pathology

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.134814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.143444Z digest=sha256:88d0b678cc003d0159a54a4951194fa10cadc374a0a6c667fcad861d31fe2d8c

Observation 92c26dc6-4b83-4553-a6ca-f3662ab335d6 · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

Vision-Language Models Do Not Understand Negation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.149164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.149164Z digest=sha256:a829d67048e5533437c179be1e1239f0bb59db1f4fe4a67f6cc14ce35a0c0b56

Observation 46d0c712-e0cf-4ceb-9a8a-be0410d8928e · outbound

This paper cites Fine-tuning llama for multi-stage text retrieval.

Vision-Language Models Do Not Understand Negation Fine-tuning llama for multi-stage text retrieval

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.105551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.154599Z digest=sha256:a263196ecba75fc30600c7efa574f1efed8ab1faf0875d47adfd372374597dbc

Observation e298a188-b5e0-4d97-b981-26f8a39358a5 · outbound

This paper cites Crepe: Can vision-language foundation models reason compositionally? In CVPR, 2023.

Vision-Language Models Do Not Understand Negation Crepe: Can vision-language foundation models reason compositionally? In CVPR, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.077283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.159371Z digest=sha256:3100fd884d5c5d58a2c227766662808dca36860fba17805acc0bf6df4e81e431

Observation 367d25f1-9eed-41c0-bc92-6b3d86ddcd67 · outbound

This paper cites Simple open-vocabulary object detection with vi- sion transformers.

Vision-Language Models Do Not Understand Negation Simple open-vocabulary object detection with vi- sion transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.046531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.164498Z digest=sha256:92fb403c374788e9fda5ff8a35007d526d0aeec555a9f80f123f194b254a2c8b

Observation 7bb530a6-5b4e-447e-983a-dc3ff370dd56 · outbound

This paper cites Recent advances in processing negation.

Vision-Language Models Do Not Understand Negation Recent advances in processing negation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.021702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.169611Z digest=sha256:40b118587caffb1d82d63a8a3525e21171358144a8e6daf8b50f5e3a4567178f

Observation 98baac9d-107f-4d94-9a8e-fc3102221a5b · outbound

This paper cites Effect of negation in sentences on sentiment analy- sis and polarity detection.

Vision-Language Models Do Not Understand Negation Effect of negation in sentences on sentiment analy- sis and polarity detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.990672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.174089Z digest=sha256:341a9302f0698225e858b0afb093bc6f16b46f205aa6c89ce94d612533d51f7e

Observation 1e9c7f92-9608-4b9a-bbb0-8f78ca2618b7 · outbound

This paper cites Clip-it! language-guided video summarization.

Vision-Language Models Do Not Understand Negation Clip-it! language-guided video summarization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.967778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.178920Z digest=sha256:34b19b1eb26c9f5390e99a3eaa1ed300e21b8519c6aac8365e134e6369e56335

Observation 8835517b-641d-4731-9c1e-a04012018817 · outbound

This paper cites Multi-Stage Document Ranking with BERT.

Vision-Language Models Do Not Understand Negation Multi-Stage Document Ranking with BERT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.185183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.185183Z digest=sha256:62c736ec9c1fe83ad63ceb9d1b1fafa5a080fd6337cadabb78dd33fc93c2a4ce

Observation 5869f8b3-04a3-4e44-a339-5be9332767af · outbound

This paper cites Synthesize diagnose and optimize: Towards fine- grained vision-language understanding.

Vision-Language Models Do Not Understand Negation Synthesize diagnose and optimize: Towards fine- grained vision-language understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.947397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.190882Z digest=sha256:220527bff5be9e86dcadf52f32a13ca7c14e736853788b475fcb1d5bd8e49ac4

Observation 3e17222d-e85e-4157-8b69-7f9b531211b2 · outbound

This paper cites On guiding vi- sual attention with language specification.

Vision-Language Models Do Not Understand Negation On guiding vi- sual attention with language specification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.926944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.198307Z digest=sha256:c73b2f6215ba02a6167bc8e0d9b7b9210d739b5ea3bb2e8bafbb5a599f261bfb

Observation 877627b6-c6ca-4b13-89d2-df8a6aa91d57 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Vision-Language Models Do Not Understand Negation Learn- ing transferable visual models from natural language super- vision

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.909588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.203619Z digest=sha256:671f75db9e5ab0f21ef9a2aa2f0e19369354c177a1d05e8764f97c165e9f21b1

Observation 1a058905-4956-4332-b0bb-e67c043cadab · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

Vision-Language Models Do Not Understand Negation Denseclip: Language-guided dense prediction with context- aware prompting

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.891949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.208659Z digest=sha256:11f2960728ed823560efb1019ae0cb504d35696d1d96559ef24b7d2c5d3e43e0

Observation b2cf1896-10f9-4896-a3a8-551b340ce5b3 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

Vision-Language Models Do Not Understand Negation Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.875485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.214930Z digest=sha256:9d71b2847df14c98434f1305f1c6c13f1e836cd1f5b87b0ed156f454437dd32b

Observation 986b4a43-0592-49e9-b54d-5fdb299a5c7d · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Vision-Language Models Do Not Understand Negation High-resolution image syn- thesis with latent diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.855890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.221528Z digest=sha256:cd496f83837ded948242fa70f46bc4c6f6aca3b053e785554c91fa70105c04d0

Observation 850b4137-6f56-4bb0-8303-dfb1b0199aa7 · outbound

This paper cites Clip for all things zero-shot sketch-based image retrieval, fine- grained or not.

Vision-Language Models Do Not Understand Negation Clip for all things zero-shot sketch-based image retrieval, fine- grained or not

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.837669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.226461Z digest=sha256:359e08601e2320388fbd43bab8549338b49da8ee6edddc082f36298c78cad02e

Observation 1a841611-2507-45f8-b381-b5259404120b · outbound

This paper cites LAION-5b: An open large-scale dataset for train- ing next generation image-text models.

Vision-Language Models Do Not Understand Negation LAION-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.819673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.231797Z digest=sha256:f8f36b7f073e6bcaf757c7b0bedd0696708667d863e73e7e0bcfa427c1fdf1bb

Observation b1117bc4-936e-443f-b7f3-fcd2367020e7 · outbound

This paper cites How much can clip benefit vision-and-language tasks? In International Conference on Learning Representa- tions.

Vision-Language Models Do Not Understand Negation How much can clip benefit vision-and-language tasks? In International Conference on Learning Representa- tions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.779307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.238053Z digest=sha256:e5a2d0f80f433ce00517443ed452ed147db0541fa26e143027cc88e994bab772

Observation 89029596-1604-4d2e-8baa-02b6076886a1 · outbound

This paper cites Proposalclip: Unsupervised open-category object pro- posal generation via exploiting clip cues.

Vision-Language Models Do Not Understand Negation Proposalclip: Unsupervised open-category object pro- posal generation via exploiting clip cues

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.759276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.244630Z digest=sha256:4717c9f7f1d07016cfe30af82b00bd61b46d80d4e0f255f1a57215b2fbdb40d8

Observation 15f7bafa-b741-49ac-964c-95d392e8eba8 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Vision-Language Models Do Not Understand Negation Cliport: What and where pathways for robotic manipulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.739522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.250104Z digest=sha256:4e2673587baad7576bf4e41f2783f0b76163d988293c7c8fac218cfb42d317e1

Observation 6b9ddda6-4d36-47c2-8059-4939e8fb9485 · outbound

This paper cites Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations.

Vision-Language Models Do Not Understand Negation Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.254987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.254987Z digest=sha256:e83aeebfffd57ec2e9fe6e57e6032d8f155d3243ca97df05d52b664ca60d558a

Observation e6d9dcd2-4121-4365-b480-4979e54b9d75 · outbound

This paper cites Stablerep: Synthetic images from text-to- image models make strong visual representation learners.

Vision-Language Models Do Not Understand Negation Stablerep: Synthetic images from text-to- image models make strong visual representation learners

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.707901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.260007Z digest=sha256:595cacdbafcec29f72c9ef05763061fa49b6812dffe4d63b374a7b9497f759c9

Observation 3cd798c9-eaeb-4969-820e-39ad73b97f21 · outbound

This paper cites Learning vision from mod- els rivals learning vision from data.

Vision-Language Models Do Not Understand Negation Learning vision from mod- els rivals learning vision from data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.670835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.265791Z digest=sha256:9627d8615c64a6880c276f2262be1979546ede17e608233761a4598d145a552d

Observation bacd2688-c153-4551-9cd1-9d3ea0705ac0 · outbound

This paper cites Expert-level detection of pathologies from unannotated chest x-ray images via self- supervised learning.

Vision-Language Models Do Not Understand Negation Expert-level detection of pathologies from unannotated chest x-ray images via self- supervised learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.653013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.271490Z digest=sha256:95f088a51788c4cc5a488040b908ba2dc4ab5803dee6fb80f6f7d406510cbf41

Observation 505de08f-ebd5-4b72-bb76-e06e2100f58f · outbound

This paper cites Language models are not naysayers: an anal- ysis of language models on negation benchmarks.

Vision-Language Models Do Not Understand Negation Language models are not naysayers: an anal- ysis of language models on negation benchmarks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.635529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.277242Z digest=sha256:b69200c0519c77b151a2a215151fea56fda60383451e720fd2a558b39bfaf557

Observation 57d0e0a7-a39f-419a-8995-ab824dc0e9ca · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Vision-Language Models Do Not Understand Negation Msr-vtt: A large video description dataset for bridging video and language

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.283190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.283190Z digest=sha256:72db96c41222efa43278d7a903fcf268e0b483565b64ad018a8e5f6c450bf29b

Observation 134bce19-b0fc-4a7e-8e3a-d53033da6fed · outbound

This paper cites Real-fake: Effective training data synthesis through distribution matching.

Vision-Language Models Do Not Understand Negation Real-fake: Effective training data synthesis through distribution matching

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.607105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.289143Z digest=sha256:82e9fdb5bc692f28fd946511a05c92b8f09e1c64d28a0fffeaa815bb224d0a79

Observation 28592488-7af3-4134-a2ea-ce92e7c9575c · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In ICLR, 2023.

Vision-Language Models Do Not Understand Negation When and why vision- language models behave like bags-of-words, and what to do about it? In ICLR, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.588431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.294239Z digest=sha256:5bbb9ec76d0f5d75801ab69142cca5be1a313144152d67e0b4f6df94c0ca7580

Observation 56c7ca48-e364-4fab-afa8-8f56e3500ae9 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

Vision-Language Models Do Not Understand Negation Lit: Zero-shot transfer with locked-image text tuning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.569731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.298798Z digest=sha256:f6f45808e829ff326507828d914a594e54148c78a2b11e9cab4df94546eef1b3

Observation 3b85b920-c49b-45cb-b1c9-3d6d2ee36ec1 · outbound

This paper cites Sigmoid loss for language image pre-training.

Vision-Language Models Do Not Understand Negation Sigmoid loss for language image pre-training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.552358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.304481Z digest=sha256:eddeec9f9c2d0260277fad1a399ee1585a3b5e7496e36a76b1ce8961308ca7eb

Observation 5bdf1839-48ec-41c2-8e4d-f6c92c722db5 · outbound

This paper cites BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.

Vision-Language Models Do Not Understand Negation BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.310828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.310828Z digest=sha256:f37631ca1a198e0bf097e50f375d6bc0d76fb759f16fb5511176a60cd2e58444

Observation 31792e35-f68e-4868-919a-52efb1ade9b3 · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:08:46.517398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.321808Z digest=sha256:8fc04a329b4ee18e15441655c727cb2c00047e49c8e10611ee9ffbec92de60d9

Observation eebb94a9-a692-433c-88eb-a37f3f06fbda · outbound

This paper cites Yes.” over “No.

Vision-Language Models Do Not Understand Negation Yes.” over “No

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.500166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.326666Z digest=sha256:9328caef2263cc593ea6c290bb3df9fd563a2c1c67556f45d525fe15f2f14308

Observation 6ff600fa-2015-4ed7-996a-b3c7c3a277ae · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-10T20:08:47.498368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.025541Z digest=sha256:b4545b17aca34ae596377818c59b30a1069f92b54c4a5b672042a382d9c40d3f

Observation 1367e721-a6e2-480a-80d9-672a22dd9629 · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 2023

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T20:08:46.534065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.316277Z digest=sha256:509c52452030eac67dce5d2e932946b0976915c59d424750e5538f04d51cb5b5

Pith citing papers

Observation 6a51e034-03ea-4726-a7a1-89714e6234a0 · inbound

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP cites this paper.

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP Vision-Language Models Do Not Understand Negation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:58.344234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:58.344234Z digest=sha256:8149bc0a878aa9b7354cc887c18abf6a7a7771849aa91622aecc09ef66e1447f

Observation 93ca8fa6-82d7-43c1-853c-34dbd8976c31 · inbound

NegVQA: Can Vision Language Models Understand Negation? cites this paper.

NegVQA: Can Vision Language Models Understand Negation? Vision-Language Models Do Not Understand Negation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:09.271049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:01:09.271049Z digest=sha256:729db4c0d5ef6ab8c0e9ec1a62e9002bac7177f22844a80ec0f444991915fded

Observation af5a841d-b63e-4378-9185-5f06f40dddbf · inbound

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs cites this paper.

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:28:21.003431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:28:21.003431Z digest=sha256:6836d12e21159deb067558925dfdc52918f8a14c4bc7efdb83638cd937ac134b

Observation 65823167-c3b5-467e-9557-9740e1845312 · inbound

Disparities In Negation Understanding Across Languages In Vision-Language Models cites this paper.

Disparities In Negation Understanding Across Languages In Vision-Language Models Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:03.990536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T03:31:50.234201Z digest=sha256:74a2c44eba0c605b0bb16c3c28a4b80b3ce9d1e143f6d011cf05dca383f94092

Observation 8f28e7c4-ea6d-489f-935b-6c6cc2a6296d · inbound

Trace Mutation in Human-LLM Dialogue: The Transcript as Forensic and Mitigation Surface cites this paper.

Trace Mutation in Human-LLM Dialogue: The Transcript as Forensic and Mitigation Surface Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:18:29.524884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:18:14.010081Z digest=sha256:01efad3d499c32f557bcc9d450c827c9bc2f595aa980cefd4ec5a84ec5611387

Observation 2471455d-3f34-45b8-8239-b49ecff50ba2 · inbound

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs cites this paper.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:09.727332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:2574fa50e407e3a35c9b1ef472ef8b49cd6bfe879f0a3b81b129e61d5d49c290

Observation 10e8c895-0807-4b6d-a917-130295408816 · inbound

Uneven Evolution of Cognition Across Generations of Generative AI Models cites this paper.

Uneven Evolution of Cognition Across Generations of Generative AI Models Vision-Language Models Do Not Understand Negation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:58.108557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T01:00:59.086900Z digest=sha256:ad655b38fac3b1113075d6330ac22628500d72e0abb8e93f7064e6f494e5070b

Observation 64c36674-3a95-4912-9f7c-3014d8fc0d3a · inbound

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution cites this paper.

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution Vision-Language Models Do Not Understand Negation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:16:49.314099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:16:49.314099Z digest=sha256:1ca284e0f3619f504366af6b25588a293fb6628d0d23d1e7dccb3f45f2e5b77f