Pith. sign in

Paper Citation Record · LEDGER

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.15576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15576 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:17:42.908883Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec091c7e-84e8-4160-8c8d-34a155b859f5 · outbound

This paper cites Distill- ing knowledge from text-to-image generative models im- proves visio-linguistic reasoning in clip.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Distill- ing knowledge from text-to-image generative models im- proves visio-linguistic reasoning in clip

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:47.435271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.136324Z digest=sha256:4528e34cbbef11bbfef355e6e6648234a48cc1e45301bf0510569519cd6bc920

Observation 92b560c6-34bd-476c-b3a9-d2147cb97ed8 · outbound

This paper cites Cross-modal common representation learning by hybrid transfer network.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Cross-modal common representation learning by hybrid transfer network

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:46.255343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.338903Z digest=sha256:825dedbd389150ac69c561180ea3f9884dc9bfe5cb341d53778cfa6f6422817a

Observation 02561694-c93f-40cf-8a5b-ae8bedb68762 · outbound

This paper cites Structure-clip: Towards scene graph knowledge to en- hance multi-modal structured representations.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Structure-clip: Towards scene graph knowledge to en- hance multi-modal structured representations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:46.073325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.386167Z digest=sha256:94d32b3c0afdaebc0f5d4f234c273c0c5e41693116fcf4530f8c207349100d47

Observation e666f391-3fb3-411a-9ee2-73ff1fef7a39 · outbound

This paper cites Adaptive prompt-based semantic embedding with inspire potential of implicit knowledge for cross-modal retrieval.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Adaptive prompt-based semantic embedding with inspire potential of implicit knowledge for cross-modal retrieval

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:45.795270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.435335Z digest=sha256:e7300ad662982d9e735df72248dfc73fc43c98fff27d10f5961f319a3f1b31e3

Observation 85db73d7-d937-4086-91b1-bef36901bcea · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:45.322222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.522432Z digest=sha256:63a7d056e47534f02a66705d3cc4172e2c42d3c17b10f911d1b972344b74f091

Observation b3abd114-552d-46df-bc4c-e6de579628bf · outbound

This paper cites Microsoft coco: Com- mon objects in context.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Microsoft coco: Com- mon objects in context

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:41.562217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:41.562217Z digest=sha256:9eb53eae8fc83eb50aa63be1fb8623b77a85ad61df1fa19d2ac10024b0993083

Observation 02da6426-1133-42ef-abc0-1f88d3ffe75b · outbound

This paper cites Image segmentation using text and image prompts.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Image segmentation using text and image prompts

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:44.823487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.689003Z digest=sha256:533c9bd3068e3435976a22cdfb6d33a6be566e135a11f37ab09ab529028059c2

Observation 72b8af50-9fce-4487-bdc8-3f60a2142d9e · outbound

This paper cites AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:17:43.137060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.738117Z digest=sha256:38a5aa4ab011857d45f6e0a0b1066f80be3c5a1dc74a737a3b57a972744fe9bd

Observation 6b11b91a-b0be-477e-896a-6acd0a4a5ea5 · outbound

This paper cites Textattack: A frame- work for adversarial attacks, data augmentation, and ad- versarial training in nlp.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Textattack: A frame- work for adversarial attacks, data augmentation, and ad- versarial training in nlp

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:44.653688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.771940Z digest=sha256:5d102da3eda5b39699640ae285ac0c234099306627103bc2a9668e50f01cf5d4

Observation 5d4cd703-11a3-4ee7-9db8-1a5e6f0ca69d · outbound

This paper cites Valse: A task-independent benchmark for vision and language models centered on linguistic phe- nomena.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Valse: A task-independent benchmark for vision and language models centered on linguistic phe- nomena

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:44.362580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.867429Z digest=sha256:0d1cbdad36c64fa251ff0c33c088f36a9402521b06189e8e9c7405e9ba412bb8

Observation 0cbf1f8a-a4bb-4e26-8d84-6aa6fd31065c · outbound

This paper cites Learning transferable visual models from nat- ural language supervision.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Learning transferable visual models from nat- ural language supervision

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:44.179538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.987278Z digest=sha256:b4b8295ba5baadeaddb0ec56cf8c856e7965cbf64555ab62170c23f5ec114951

Observation cb4a62ac-5a42-473d-9441-1154ee5e00a6 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:43.907616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:42.271735Z digest=sha256:329d2070f830e6f7025bd2a215840aceee5494a04a15345dbe07fdfcfbba34a9

Observation 84b12e22-3366-4d4b-8043-9a845ee50fa2 · outbound

This paper cites Image as a foreign language: Beit pretrain- ing for vision and vision-language tasks.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Image as a foreign language: Beit pretrain- ing for vision and vision-language tasks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:43.788994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:42.417377Z digest=sha256:e4a851f83ae2903b32408e680ee9914cfc0481c57eb8e0f43b06f2127fb9b942

Observation a201851e-5b78-4fa4-8326-dfef90c5a8e3 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Groupvit: Semantic segmentation emerges from text supervision

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:43.593934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:42.546581Z digest=sha256:043bed74dea63bcb71bee9cebf89861385dc3a21cb45db1097fc080485012177

Observation f15431b8-7660-43e9-abf1-13e05e3da31f · outbound

This paper cites CREPE: open-domain ques- tion answering with false presuppositions.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models CREPE: open-domain ques- tion answering with false presuppositions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:43.376201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:42.679353Z digest=sha256:0421eab062bf3793c6939b674b68ce10df4b8b7a9b344f8dcaa2ea2f380f8a36

Observation dff164eb-6c6a-4acf-8cb4-bee651993ea6 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:42.791864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:42.791864Z digest=sha256:758f4d34c34f36a6b6a2608659e5a08a329ac7e1eda497c4e1696c63a0495925

Observation e203c0d3-7549-434a-9fba-2910c9f836f2 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:42.908883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:42.908883Z digest=sha256:06bfd4bfcda76ace2ffe97f047a2999fa5b6bb2e3d9545a6214aec237843fce2

Observation 330b64f6-7ef8-41e1-ab72-c3da8578474a · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:41.609162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:41.609162Z digest=sha256:74919b7343325a25d057a7bdf8fcaef5fc6a5e2488244c3f62c1fdebd83a2718

Observation e2954376-bc86-4cf7-a5db-f3cedd1d4238 · outbound

This paper cites Sugar- crepe: fixing hackable benchmarks for vision-language compositionality.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Sugar- crepe: fixing hackable benchmarks for vision-language compositionality

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:46.523237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.309512Z digest=sha256:fe3d3b75ff257450b4745a1b1b0b79437002b7738a80eba6e5a26057b847f2a0

Observation 8a557498-957f-4da6-bc18-93cd963a7d59 · outbound

This paper cites Smith, Yejin Choi, and Hannaneh Ha- jishirzi.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Smith, Yejin Choi, and Hannaneh Ha- jishirzi

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:45.063534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.654189Z digest=sha256:d6d3bea13a57389ab4377f6af5efe22076da6015dc7b829d9519415cd9b9f692

Observation 7ccaf003-a3a7-45dc-8058-ea0b90f595cf · outbound

This paper cites Chils: Zero- shot image classification with hierarchical label sets.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Chils: Zero- shot image classification with hierarchical label sets

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:44.506009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.832915Z digest=sha256:b917a91f0d438b2e944fcda059e7821c3f021516c54c552d5f8e014b786c1248

Observation a189848c-4dbe-48c3-b195-b6e52c4a40f6 · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:44.043181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:42.143732Z digest=sha256:0ccc04de4cf45ae6d7d26553f0355bf227558c32712679278a5770cc320f2f7f

Observation c418a5b3-7238-4f24-87f9-f94242e88f99 · outbound

This paper cites spacy 2: Natural lan- guage understanding with bloom embeddings, convolu- tional neural networks and incremental parsing.To appear,.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models spacy 2: Natural lan- guage understanding with bloom embeddings, convolu- tional neural networks and incremental parsing.To appear,

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:46.855712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.274371Z digest=sha256:9fb9dde79e48ede30495877bc86535f6825fba6e7e57873b350b4a8898ecf4b0

Observation 6ef37cb2-d20d-4616-82b2-d37e257a9bfd · outbound

This paper cites CyCLIP: Cyclic contrastive language-image pretraining.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models CyCLIP: Cyclic contrastive language-image pretraining

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:47.015340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.221150Z digest=sha256:9d77ac0668fdc235ad92e8299192596bf96a7bb4f73f367ba024707b90104f57

Observation 8e556a4b-9df3-43d6-a5e9-a246ffaee2c1 · outbound

This paper cites Going beyond nouns with vision & language models using synthetic data.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Going beyond nouns with vision & language models using synthetic data

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:47.220575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.177955Z digest=sha256:5ceb5b19aec21d245916b2c6ea6b988a9ef39e35c586c4b3b9dd70443b357ee8

Observation 00bc3172-cf4b-41a1-a036-4a5a0eb713bf · outbound

This paper cites Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation.

Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:17:45.559757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:41.478450Z digest=sha256:eaec771f3f2d9626c7aa268dadf6ba6677745600c848620f0ddb7327899f772a

Pith citing papers

No inbound Pith citation observations are available.