Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Models Do Not Understand Negation

As of 15 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 10 inbound Pith citation observations for arXiv:2501.09425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09425 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:08:46.326666Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:08:58.575726Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T00:18:29.517494Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved11
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 876019cf-4e11-484d-93b7-b82286f09e37 · outbound

This paper cites Aug- mented reality meets computer vision: Efficient data gen- eration for urban driving scenes.

Vision-Language Models Do Not Understand Negation Aug- mented reality meets computer vision: Efficient data gen- eration for urban driving scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.570497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.008226Z digest=sha256:c18c8f8bc95df37b696ff2c2938d8b49fd67422d3e86e2d2e3c0346edaddb6f0

Observation 76626a12-b8c4-4540-861e-c24ddca9336b · outbound

This paper cites Effective conditioned and composed im- age retrieval combining clip-based features.

Vision-Language Models Do Not Understand Negation Effective conditioned and composed im- age retrieval combining clip-based features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.548229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.014749Z digest=sha256:600e7adc3bfe459f975e91a06aead2a746bc4cb4db3619adc286b28a2f277ba4

Observation 5d08414e-b6b6-492a-ac86-dbac81d7a1fe · outbound

This paper cites FitCLIP: Refining large- scale pretrained image-text models for zero-shot video un- derstanding tasks.

Vision-Language Models Do Not Understand Negation FitCLIP: Refining large- scale pretrained image-text models for zero-shot video un- derstanding tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.517244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.020123Z digest=sha256:2f12f00f629b09ffd8b48693e2a7d44633d5c259e731e79f068d505e2acb2af8

Observation adc40a9b-4ad3-44b7-96a0-b98ec2b7ff5d · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Vision-Language Models Do Not Understand Negation Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.031905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.031905Z digest=sha256:7f54999cdc33a4caadd7b5db60442a560ce2f875e3762d839b094d5546e6f32d

Observation 19db9dd0-2961-4ed7-ad64-5bae5646be84 · outbound

This paper cites Learning semantic segmentation from synthetic data: A geo- metrically guided input-output adaptation approach.

Vision-Language Models Do Not Understand Negation Learning semantic segmentation from synthetic data: A geo- metrically guided input-output adaptation approach

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.453953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.037424Z digest=sha256:91ae44413ce5fc6ee7f20b36e5bd81b75eba683f24c7f60e33885e2406ed62bb

Observation 3e0cc5c6-38ab-4ba6-855b-8a80c4030f3d · outbound

This paper cites The Llama 3 Herd of Models.

Vision-Language Models Do Not Understand Negation The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.042836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.042836Z digest=sha256:18e948b9b2c1fba26e7f2137eb953dae674ad405c13b38a30c0857884db3d4a4

Observation 97d27b26-b7c2-45d5-ab5c-dbe351a3529a · outbound

This paper cites The pascal visual object classes (voc) challenge.

Vision-Language Models Do Not Understand Negation The pascal visual object classes (voc) challenge

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.429941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.048534Z digest=sha256:7c86f32f43d0668e6298058eccef3be4808e85a70bcff2eb0600e12c39a38e00

Observation 11613278-2c26-4b96-9ae0-c64f4c515b2a · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

Vision-Language Models Do Not Understand Negation Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.054919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.054919Z digest=sha256:3d412650a5b3601f7f7432e8d3694decd30314ae7df96a665568107285234d7f

Observation 43de77e5-ed52-44ba-8ed4-68d1bedef7ab · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Vision-Language Models Do Not Understand Negation Datacomp: In search of the next generation of multimodal datasets

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.402544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.061090Z digest=sha256:7ecddab482e9666fd0b0916af36336b55e78d8d90d392726c03940a56616d332

Observation e1117c67-0df2-4e42-a899-fb9d4b5c7bcd · outbound

This paper cites This is not a dataset: A large negation benchmark to challenge large language mod- els.

Vision-Language Models Do Not Understand Negation This is not a dataset: A large negation benchmark to challenge large language mod- els

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.382504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.076557Z digest=sha256:0a08ec139e8caa5147be91b858689733b899be0d0be61c98d8f2270397b45dac

Observation 504dc068-fbb3-4026-b1d4-87f3c4e9d131 · outbound

This paper cites Shortcut learning in deep neural networks.

Vision-Language Models Do Not Understand Negation Shortcut learning in deep neural networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.355708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.081422Z digest=sha256:bea280e571154cb67c356fc0a0ea7da3cbc4d2f451cbb58fc7cae3589a654d17

Observation 6ad3e1d3-2a2b-4509-be2f-5b1e12a3db5e · outbound

This paper cites SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?.

Vision-Language Models Do Not Understand Negation SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.086165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.086165Z digest=sha256:e417ccd124711206272e1253546cd9e5b99f3fadb3610badfa4c0f5ead4d603f

Observation 863bde2d-1a1e-433d-bfcd-09a3e99ecd6e · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:08:47.335828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.091030Z digest=sha256:5b950eb69f820da83d54c9d03fc0fe1a0afce44b2f7faba7927e620c186a4f99

Observation fd484441-67ad-4796-b2bf-0cc61cb951c7 · outbound

This paper cites Quilt-1m: One million image-text pairs for histopathology.

Vision-Language Models Do Not Understand Negation Quilt-1m: One million image-text pairs for histopathology

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.314683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.096959Z digest=sha256:02bb278be8ab617dfdf83f6abf448c12fc2c9e01d70e07f593bd79fdd8b5d140

Observation 73531a14-883c-4bee-842c-7c0fc3a33525 · outbound

This paper cites Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison.

Vision-Language Models Do Not Understand Negation Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.296473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.103106Z digest=sha256:8331664b0ca81fd2842b1eb07cc8940cb85bef31c3d6b29d237196685823e0b6

Observation a0594ace-e525-4638-b328-70feab20beaa · outbound

This paper cites Generative models as a data source for multiview representa- tion learning.

Vision-Language Models Do Not Understand Negation Generative models as a data source for multiview representa- tion learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.277193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.108782Z digest=sha256:9a45f4b4b777315036006d384e4a1cf885a39d61649573aca11696a639621621

Observation 6825bbb5-161f-4b27-8c15-85c4acb02932 · outbound

This paper cites The power of negation in english: Text, context and relevance.

Vision-Language Models Do Not Understand Negation The power of negation in english: Text, context and relevance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.257383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.113421Z digest=sha256:57c42996be14f118a2a62373a5ff2b64b112d6442f14bd1a2242e453c0468196

Observation b40567b7-2ebc-44d1-a989-c1c016b6515d · outbound

This paper cites Negation in syntax–on the na- ture of functional categories and projections.

Vision-Language Models Do Not Understand Negation Negation in syntax–on the na- ture of functional categories and projections

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.234939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.118134Z digest=sha256:60b088517688c3e1b65e9984deda0ade7f49ff681b37a7e3d2bc311a77e24664

Observation e4b8857b-2a24-492a-8b01-221b1f34db22 · outbound

This paper cites Naturalbench: Evalu- ating vision-language models on natural adversarial samples.

Vision-Language Models Do Not Understand Negation Naturalbench: Evalu- ating vision-language models on natural adversarial samples

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.214576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.122815Z digest=sha256:0c0362174ccf856dafc0cda1b3165f19e8853ceb99a326ce0b379a935a098272

Observation 51eb013c-e984-4379-8ec6-8ce2bcc29324 · outbound

This paper cites Compre- hending and ordering semantics for image captioning.

Vision-Language Models Do Not Understand Negation Compre- hending and ordering semantics for image captioning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.194283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.128486Z digest=sha256:475c117cf3905c26906475ad86d8e81468059d7c5d007bddc97cc0d25e67d9ac

Observation e2d73bba-3d18-4099-bf07-0007c3bfce36 · outbound

This paper cites Cross-modal retrieval and semantic re- finement for remote sensing image captioning.

Vision-Language Models Do Not Understand Negation Cross-modal retrieval and semantic re- finement for remote sensing image captioning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.174919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.133410Z digest=sha256:eb7b9f6c449e533708a82bb44db6638998acec5741360f230adadc4ab279b34d

Observation 9c371f82-ffe9-4c86-b92b-0595593f0e24 · outbound

This paper cites Microsoft coco: Common objects in context.

Vision-Language Models Do Not Understand Negation Microsoft coco: Common objects in context

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.156422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.138330Z digest=sha256:57c6a68005410faff32a54b1d00f7a10b7f13fc8afbda64c9f518fce927911e9

Observation c376ff10-6159-490e-b8f8-24a60b6a017e · outbound

This paper cites A visual- language foundation model for computational pathology.

Vision-Language Models Do Not Understand Negation A visual- language foundation model for computational pathology

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.134814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.143444Z digest=sha256:018201370a785dadd73e27e65ecdb3c1131ec272c31b8b63fe536915f8c6d4e6

Observation 92c26dc6-4b83-4553-a6ca-f3662ab335d6 · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

Vision-Language Models Do Not Understand Negation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.149164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.149164Z digest=sha256:efb23dbc0bc224b72c92399bde3d26b5bf9be0b0bb951ddb23402bfbc5e53a9c

Observation 46d0c712-e0cf-4ceb-9a8a-be0410d8928e · outbound

This paper cites Fine-tuning llama for multi-stage text retrieval.

Vision-Language Models Do Not Understand Negation Fine-tuning llama for multi-stage text retrieval

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.105551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.154599Z digest=sha256:8ed7692aeffe705301a5dea8b84a88612040a58addc5944e02947ced0c8951e6

Observation e298a188-b5e0-4d97-b981-26f8a39358a5 · outbound

This paper cites Crepe: Can vision-language foundation models reason compositionally? In CVPR, 2023.

Vision-Language Models Do Not Understand Negation Crepe: Can vision-language foundation models reason compositionally? In CVPR, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.077283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.159371Z digest=sha256:ae958e8003c752fca534f8515ccc971295cfe3b6c66f915588557a75155887a3

Observation 367d25f1-9eed-41c0-bc92-6b3d86ddcd67 · outbound

This paper cites Simple open-vocabulary object detection with vi- sion transformers.

Vision-Language Models Do Not Understand Negation Simple open-vocabulary object detection with vi- sion transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.046531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.164498Z digest=sha256:2c16a5dd4cf67beed77afa767dcd204d7d8dd98a8e384f08c10dc16886fa0940

Observation 7bb530a6-5b4e-447e-983a-dc3ff370dd56 · outbound

This paper cites Recent advances in processing negation.

Vision-Language Models Do Not Understand Negation Recent advances in processing negation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.021702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.169611Z digest=sha256:f4e0eecef4552a2167ad32a89b2e3deac17f22f3e89a083b7f1a53d18dc9512e

Observation 98baac9d-107f-4d94-9a8e-fc3102221a5b · outbound

This paper cites Effect of negation in sentences on sentiment analy- sis and polarity detection.

Vision-Language Models Do Not Understand Negation Effect of negation in sentences on sentiment analy- sis and polarity detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.990672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.174089Z digest=sha256:fb438fab51fbc6efc11440451eed714b35205c556463f2bb997886e22be20d8e

Observation 1e9c7f92-9608-4b9a-bbb0-8f78ca2618b7 · outbound

This paper cites Clip-it! language-guided video summarization.

Vision-Language Models Do Not Understand Negation Clip-it! language-guided video summarization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.967778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.178920Z digest=sha256:19a8338218f9a8513a8823c450302d073b4b4753cf296be0b7cbe2d30aa80262

Observation 8835517b-641d-4731-9c1e-a04012018817 · outbound

This paper cites Multi-Stage Document Ranking with BERT.

Vision-Language Models Do Not Understand Negation Multi-Stage Document Ranking with BERT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.185183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.185183Z digest=sha256:ffd6ea411efab08782680220f894da9e62a721a302abe86331448d28ae5dd5a8

Observation 5869f8b3-04a3-4e44-a339-5be9332767af · outbound

This paper cites Synthesize diagnose and optimize: Towards fine- grained vision-language understanding.

Vision-Language Models Do Not Understand Negation Synthesize diagnose and optimize: Towards fine- grained vision-language understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.947397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.190882Z digest=sha256:38c32a6407e2f6466cc18692486805fec810c878d412ab72a47131b99fca7ae5

Observation 3e17222d-e85e-4157-8b69-7f9b531211b2 · outbound

This paper cites On guiding vi- sual attention with language specification.

Vision-Language Models Do Not Understand Negation On guiding vi- sual attention with language specification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.926944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.198307Z digest=sha256:c7e1b06786b81c9aea3e283c8c8db52eaf139d65fb531686eb347a82e81dbaa9

Observation 877627b6-c6ca-4b13-89d2-df8a6aa91d57 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Vision-Language Models Do Not Understand Negation Learn- ing transferable visual models from natural language super- vision

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.909588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.203619Z digest=sha256:b0a55f489166945dffb23e182e19b3bb3fffe426ff9d7e1d794a0c4f9e839725

Observation 1a058905-4956-4332-b0bb-e67c043cadab · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

Vision-Language Models Do Not Understand Negation Denseclip: Language-guided dense prediction with context- aware prompting

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.891949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.208659Z digest=sha256:774d779b36a00e28b7c5552e1b7c8da0e731ee2c0d63ad3dc2065b0eea988a9e

Observation b2cf1896-10f9-4896-a3a8-551b340ce5b3 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

Vision-Language Models Do Not Understand Negation Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.875485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.214930Z digest=sha256:f8ff852bc620a6a217bfcb061f52d99972832450bc05d4b347ccb4abdd4137f3

Observation 986b4a43-0592-49e9-b54d-5fdb299a5c7d · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Vision-Language Models Do Not Understand Negation High-resolution image syn- thesis with latent diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.855890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.221528Z digest=sha256:cd16e80f2a8cb9698243421fb98de03fec286df06ea07f5405e7bc6db2e9a0cd

Observation 850b4137-6f56-4bb0-8303-dfb1b0199aa7 · outbound

This paper cites Clip for all things zero-shot sketch-based image retrieval, fine- grained or not.

Vision-Language Models Do Not Understand Negation Clip for all things zero-shot sketch-based image retrieval, fine- grained or not

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.837669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.226461Z digest=sha256:5580d557c7da250da683e3658dd255fb79e1181b5ae59dc25a456515a000672f

Observation 1a841611-2507-45f8-b381-b5259404120b · outbound

This paper cites LAION-5b: An open large-scale dataset for train- ing next generation image-text models.

Vision-Language Models Do Not Understand Negation LAION-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.819673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.231797Z digest=sha256:23d8f0e850d09ceb73bbd60e8a8e712036b800509905df3fcdf1eaca551e5366

Observation b1117bc4-936e-443f-b7f3-fcd2367020e7 · outbound

This paper cites How much can clip benefit vision-and-language tasks? In International Conference on Learning Representa- tions.

Vision-Language Models Do Not Understand Negation How much can clip benefit vision-and-language tasks? In International Conference on Learning Representa- tions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.779307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.238053Z digest=sha256:adb5dfd3c81e8b7c158f5867cf00aa7efd0f9dae15283abbe3407d762b510db2

Observation 89029596-1604-4d2e-8baa-02b6076886a1 · outbound

This paper cites Proposalclip: Unsupervised open-category object pro- posal generation via exploiting clip cues.

Vision-Language Models Do Not Understand Negation Proposalclip: Unsupervised open-category object pro- posal generation via exploiting clip cues

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.759276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.244630Z digest=sha256:0bee1762719d9c79a0f9c32e97541735488dc8b3f2e4f837a322210de099780c

Observation 15f7bafa-b741-49ac-964c-95d392e8eba8 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Vision-Language Models Do Not Understand Negation Cliport: What and where pathways for robotic manipulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.739522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.250104Z digest=sha256:6bc372b4f05807af42a6b592602781b0c67a0130044fa28c904ca0d5cf813060

Observation 6b9ddda6-4d36-47c2-8059-4939e8fb9485 · outbound

This paper cites Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations.

Vision-Language Models Do Not Understand Negation Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.254987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.254987Z digest=sha256:fd4f9e675665c114748d4220197ac3a3ced1bceda50de877cff265dc3f38535e

Observation e6d9dcd2-4121-4365-b480-4979e54b9d75 · outbound

This paper cites Stablerep: Synthetic images from text-to- image models make strong visual representation learners.

Vision-Language Models Do Not Understand Negation Stablerep: Synthetic images from text-to- image models make strong visual representation learners

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.707901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.260007Z digest=sha256:b9fe25861ec7b4b90646d0eccfb2cc2b0869e418da6791e99cbf893c57c81764

Observation 3cd798c9-eaeb-4969-820e-39ad73b97f21 · outbound

This paper cites Learning vision from mod- els rivals learning vision from data.

Vision-Language Models Do Not Understand Negation Learning vision from mod- els rivals learning vision from data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.670835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.265791Z digest=sha256:9f9bc1f9655ad2a2d20ef91409c09e5054f45601e52f4dfe715d0db2a00f3e88

Observation bacd2688-c153-4551-9cd1-9d3ea0705ac0 · outbound

This paper cites Expert-level detection of pathologies from unannotated chest x-ray images via self- supervised learning.

Vision-Language Models Do Not Understand Negation Expert-level detection of pathologies from unannotated chest x-ray images via self- supervised learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.653013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.271490Z digest=sha256:31608bddc16e25a768a1aa699437cbc9d8d82850c682bc8b720e7278d1cc67a7

Observation 505de08f-ebd5-4b72-bb76-e06e2100f58f · outbound

This paper cites Language models are not naysayers: an anal- ysis of language models on negation benchmarks.

Vision-Language Models Do Not Understand Negation Language models are not naysayers: an anal- ysis of language models on negation benchmarks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.635529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.277242Z digest=sha256:a4628e07aa54e7637a436d663e52e302e96984405ecaa28f33a3aea3671addbb

Observation 57d0e0a7-a39f-419a-8995-ab824dc0e9ca · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Vision-Language Models Do Not Understand Negation Msr-vtt: A large video description dataset for bridging video and language

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.283190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.283190Z digest=sha256:571107f747d075b6d0c4f2500da3227fe04dbefdda6d92bfd37f10b668b9108b

Observation 134bce19-b0fc-4a7e-8e3a-d53033da6fed · outbound

This paper cites Real-fake: Effective training data synthesis through distribution matching.

Vision-Language Models Do Not Understand Negation Real-fake: Effective training data synthesis through distribution matching

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.607105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.289143Z digest=sha256:93aa0a85771de1eeab52ecd9063df8934ab13aaf1ce9a4cd523e87021caa6913

Observation 28592488-7af3-4134-a2ea-ce92e7c9575c · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In ICLR, 2023.

Vision-Language Models Do Not Understand Negation When and why vision- language models behave like bags-of-words, and what to do about it? In ICLR, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.588431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.294239Z digest=sha256:f39e8c8b8f047a3ada39b8d46acbc86a38903bd463f9a742d76c7762ecf1023c

Observation 56c7ca48-e364-4fab-afa8-8f56e3500ae9 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

Vision-Language Models Do Not Understand Negation Lit: Zero-shot transfer with locked-image text tuning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.569731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.298798Z digest=sha256:06a22903706ef0338326a3eab40cfe53c2dadc040380831cce349151ebe677d6

Observation 3b85b920-c49b-45cb-b1c9-3d6d2ee36ec1 · outbound

This paper cites Sigmoid loss for language image pre-training.

Vision-Language Models Do Not Understand Negation Sigmoid loss for language image pre-training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.552358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.304481Z digest=sha256:e391fdaaf605c815ff34236ea22118a390dbd3d9f6184c22ac478827af83ed7f

Observation 5bdf1839-48ec-41c2-8e4d-f6c92c722db5 · outbound

This paper cites BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.

Vision-Language Models Do Not Understand Negation BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.310828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.310828Z digest=sha256:3a97c6dd71f4a69969eb518e9d35fe2db0006b50f6b9f483b240173db4dbf15c

Observation 31792e35-f68e-4868-919a-52efb1ade9b3 · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:08:46.517398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.321808Z digest=sha256:465c0b1cd4fb33a0735b6abcddeb4cfab14bd7206530e5bbecf98b37905bf08d

Observation eebb94a9-a692-433c-88eb-a37f3f06fbda · outbound

This paper cites Yes.” over “No.

Vision-Language Models Do Not Understand Negation Yes.” over “No

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.500166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.326666Z digest=sha256:c57a48398163b877e0fea0a5dc3acce51ea4e55f682710d16646737ac153caed

Observation 6ff600fa-2015-4ed7-996a-b3c7c3a277ae · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-10T20:08:47.498368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.025541Z digest=sha256:2343eaaf74be1b1d4dc63c3cc01861832985d818501cb237bfa20d62d8a5968e

Observation 1367e721-a6e2-480a-80d9-672a22dd9629 · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 2023

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T20:08:46.534065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T20:08:46.316277Z digest=sha256:d1c1175a92a1e6c284996c3651a6c53de4e3be7d0684f9af52dd2228640761ed

Pith citing papers

Observation 4139ed13-aaf4-457f-826e-c0a487ba724c · inbound

On the Compatibility of Generative AI and Generative Linguistics cites this paper.

On the Compatibility of Generative AI and Generative Linguistics Vision-Language Models Do Not Understand Negation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:40:04.817443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:40:04.817443Z digest=sha256:be1f45e764572300d8c8c3a8d30298c06519b9186a5cd2dcb131aa698cfc81e8

Observation 6a51e034-03ea-4726-a7a1-89714e6234a0 · inbound

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP cites this paper.

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP Vision-Language Models Do Not Understand Negation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:58.344234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:58.344234Z digest=sha256:3c4a024e0ad9659c77a5802cbe01c4f00988244a50b47d3496436d7474947040

Observation 93ca8fa6-82d7-43c1-853c-34dbd8976c31 · inbound

NegVQA: Can Vision Language Models Understand Negation? cites this paper.

NegVQA: Can Vision Language Models Understand Negation? Vision-Language Models Do Not Understand Negation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:09.271049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:01:09.271049Z digest=sha256:b4b0cde317f4efd65b52f9c91306a12b509e4f0f7f427ea2c4a3587248c7135c

Observation 08c993ae-bb63-403f-bdf3-428abeaada82 · inbound

Negation-Aware Test-Time Adaptation for Vision-Language Models cites this paper.

Negation-Aware Test-Time Adaptation for Vision-Language Models Vision-Language Models Do Not Understand Negation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:08:58.575726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:08:58.575726Z digest=sha256:ef0048cecbb340e8cff69a7bf42240baa882600849690704ddf7bb62bc20f4a1

Observation af5a841d-b63e-4378-9185-5f06f40dddbf · inbound

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs cites this paper.

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:28:21.003431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:28:21.003431Z digest=sha256:1ac30f0bde2d7a3f1361b7a8e1157ca395e69ae7e6b3efe5c08dca6ede2a47de

Observation 65823167-c3b5-467e-9557-9740e1845312 · inbound

Disparities In Negation Understanding Across Languages In Vision-Language Models cites this paper.

Disparities In Negation Understanding Across Languages In Vision-Language Models Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:03.990536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T03:31:50.234201Z digest=sha256:2b3bc22cacdbd90ba8c76ad77acaf982521549fdb685bd44850b1d48dc53c984

Observation 8f28e7c4-ea6d-489f-935b-6c6cc2a6296d · inbound

Trace Mutation in Human-LLM Dialogue: The Transcript as Forensic and Mitigation Surface cites this paper.

Trace Mutation in Human-LLM Dialogue: The Transcript as Forensic and Mitigation Surface Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:18:29.524884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T00:18:14.010081Z digest=sha256:4350b394c306bf232d52abed26265deb4b6705e2b0aec5924676d0849c4fb23c

Observation 2471455d-3f34-45b8-8239-b49ecff50ba2 · inbound

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs cites this paper.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:09.727332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:e098ee7a77df93c583fec05a1a44b3c0938d5c2db4ca3ed349627bd32c70c0f4

Observation 10e8c895-0807-4b6d-a917-130295408816 · inbound

Uneven Evolution of Cognition Across Generations of Generative AI Models cites this paper.

Uneven Evolution of Cognition Across Generations of Generative AI Models Vision-Language Models Do Not Understand Negation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:58.108557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T01:00:59.086900Z digest=sha256:ad164288aefec6828cae56b4c2cd8183c27e1bf1560ff78fd5ea0de339312a3d

Observation 64c36674-3a95-4912-9f7c-3014d8fc0d3a · inbound

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution cites this paper.

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution Vision-Language Models Do Not Understand Negation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:16:49.314099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:16:49.314099Z digest=sha256:1898c03efee6186a8908bad629d6b9c9f0c5e32fd7ad2c098c5f42673a15cf29