Pith. sign in

Paper Citation Record · LEDGER

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2507.04699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04699 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:45:53.015838Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e541e223-0112-4f2b-bbc9-a9d01ba6a773 · outbound

This paper cites GPT-4 Technical Report.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:51.395705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:51.395705Z digest=sha256:c20d6bd1256d381a1e6e4723045a415c0333b7adf041d6bdd58b1d0dd1976ff3

Observation 227d747f-b29e-408f-a3b9-4f62ce9cb1cb · outbound

This paper cites Vismin: Visual minimal-change understanding.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Vismin: Visual minimal-change understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.566352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:51.925349Z digest=sha256:86364ee2d671531bcd79562f59aebab80c0f5422ed67712c10ce41cf61aa6d41

Observation 3880f207-f9cf-4cb5-b486-27be3270d16f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.329853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.329853Z digest=sha256:aa49d6a9959caecf8f92128b9607ae7dc01966bdc28aa87c1f4f2eba03265562

Observation 1c97579e-5679-49a4-bffb-f94465bededd · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.383239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.383239Z digest=sha256:0d2d9ed678f81355d3a0cfd8ab7a02b08d579aaeed349932cde6e4faa031692a

Observation 9a65d044-fa8b-44fe-9df2-d56f09da7c57 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Clip2scene: Towards label-efficient 3d scene understanding by clip

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.555147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.413537Z digest=sha256:dbf78837215a6127aaffceefe58be5739ec0d396b653144297cb250f34c6a026

Observation 92458099-43b3-46b0-9043-596746f5c15e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets A simple framework for contrastive learning of visual representations

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.543996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.487593Z digest=sha256:8a2e65557a4cc9aa64d3196eab91c933482a7db4de52e6978871be2db918f3e7

Observation 3f519064-4f4e-4d59-9604-802108f68bf0 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.563298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.563298Z digest=sha256:576973f66644681ec85a3add76ce7c04c8c5e083df5ee444dbc5997c1025c453

Observation d1ed6828-d923-4682-b1cd-7e32a88ecf1c · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.624039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.624039Z digest=sha256:58cee8ffb90df7e4a3d7ef2612bb1f4a12afb47e9172090465d8c3a16d653ac8

Observation 255f7f21-4cc0-4c52-b8bb-4f8f972f554a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.695270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.695270Z digest=sha256:e22d0eb32411b135c4f65fad7a88ed2130687dbc6da0aae828a19d958c8cdc39

Observation 907d0124-c006-4743-8192-02908b83132e · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.531705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.724114Z digest=sha256:20aca25e840100e5ee91f1d8fc9e3556d2515832d8d5555456089c4837bb2b21

Observation 8fe1138b-f0b8-424d-8daf-e48f4b130349 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.761831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.761831Z digest=sha256:c0997767548f18ae7eb88bc4542ebb5730b78a32968a932f08d53f131fa0bc72

Observation 2870a00f-fa24-45d3-8a1f-15b453fb94fe · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36:76137–76150, 2023.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36:76137–76150, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.520358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.779677Z digest=sha256:8e5b835347e1451df2f53937e0387bc1bb68db96ddebb676ac0c0b29be472fcd

Observation 53581e1d-5568-463a-a2ed-2a178875b130 · outbound

This paper cites Teaching structured vision & language concepts to vision & language models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Teaching structured vision & language concepts to vision & language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.509035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.873754Z digest=sha256:47db816b46966178689406c679b347e654e72372c679ff8c4bb6ac266af2a858

Observation b3cbef5f-286e-4ca8-abc8-4dc10f2b40d7 · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.497514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.904104Z digest=sha256:f2d11925039d8e684f39c913d6b4e326698a15516916efca331ab3c6760d4e8a

Observation ff1e9ff5-461c-410f-99de-bdf4a9bf2b57 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.484864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.907558Z digest=sha256:7c386443cf806aea03800974d03934de691156d74b132a2f4184a93793365cb1

Observation bc526a0e-7a79-435c-b2c1-9d0b1222b699 · outbound

This paper cites Advances in deep concealed scene understanding.Visual Intelligence, 1(1):16,.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Advances in deep concealed scene understanding.Visual Intelligence, 1(1):16,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.472951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.910679Z digest=sha256:431a7333e69f9a89c030faf1de7f058fd41c8f2fdfd734823ea566526f5b1e4f

Observation 7883666e-4ec6-4b91-a616-a090dfa88957 · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.913989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.913989Z digest=sha256:869977f56479fe606b04be2ab433bee1080da6f021fe022e2abb95ea1ca629f6

Observation ff800481-a222-45c8-80db-7d4383db1b2e · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Visual program- ming: Compositional visual reasoning without training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.917752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.917752Z digest=sha256:19d05ea16c6a6233a48ae01896e3f3dcc32e35e9cf20e1508dd2139844d79f97

Observation 034cbaee-71a7-48be-bae9-c85b71a30a3a · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.453433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.921818Z digest=sha256:2a651f5efc7e16f2117e496204d24e57ee2fc007f43ec961808cb7efb5a4505f

Observation 37d2f436-fce2-4aec-9234-3c5a9aa599ab · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets LoRA: Low-Rank Adaptation of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.925907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.925907Z digest=sha256:e068b42cf6737752ed9261c516600ef74099c78332637442e94e358a3edb70d9

Observation 784193e0-7fa7-4c41-970e-390885a53cf6 · outbound

This paper cites Semantic to Structure: Learning Structural Representations for Infringement Detection.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Semantic to Structure: Learning Structural Representations for Infringement Detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.929510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.929510Z digest=sha256:df59a78a658901a13e940d02c8537f98058ca0ddc961952106fe85ce4f31899a

Observation 56f713ab-86a7-4915-b5d1-57a18790bbda · outbound

This paper cites Secret lies in color: Enhancing ai-generated images detection with color distribution anal- ysis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Secret lies in color: Enhancing ai-generated images detection with color distribution anal- ysis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.441738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.933305Z digest=sha256:9d60478fd4f948360bc4400c0db9da11f0a1d8df19c58479c2cec34ee9dddc21

Observation d19b3da0-5082-4a78-aa13-25f335ec430d · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.Advances in Neural Information Processing Systems, 35:9287–9301,.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Elevater: A benchmark and toolkit for evaluating language-augmented visual models.Advances in Neural Information Processing Systems, 35:9287–9301,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.936493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.936493Z digest=sha256:7f3d3295ffdeb4e4ea6932433805e636f545fec3f04f45b6ab9d6b788816cc01

Observation 3398b2d4-3be4-426a-a529-90823f509958 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.939887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.939887Z digest=sha256:56b8bc8798bafef4fb6701a63f39f3d7e99eab28b604784b6848a2bd53755d0e

Observation d0c50298-4162-40d0-a531-ad5605347184 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.415920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.943579Z digest=sha256:a3223673925816c8a35b2e452bd1bb58a66de102918fd56c0ee92428d14396be

Observation 8a3fd584-54a8-4495-8687-75df5a833d64 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Vila: On pre-training for vi- sual language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.403872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.947716Z digest=sha256:97aaff58715076a685548a7d9afc4b82234649d994eb04923ea12d6876be989a

Observation e46b6ba9-fbad-431f-a29a-0b7d99b39aa6 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.392740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.951149Z digest=sha256:1c83a053a0cd51d145842bb3c0393171cf36bf065bb4a7e0c71d679545dfbefc

Observation 2280f4f3-b2b0-4be1-838b-7c667339728c · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.956170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.956170Z digest=sha256:cd80a2414b75ea946c56bb25c8f2312825cdacb3cd4f73c938b6a3b00617cf03

Observation 0c1b60e1-3386-462d-a2d7-9aa857a4ddf5 · outbound

This paper cites Synthesize diagnose and optimize: Towards fine- grained vision-language understanding.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Synthesize diagnose and optimize: Towards fine- grained vision-language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.381845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.960323Z digest=sha256:6e0dfc2589eb85ad3d75d498763bf414a755d5c8067a37635d8c0e4e18606122

Observation d93327d3-3d0c-4963-85bc-6477f84f3fa1 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.963677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.963677Z digest=sha256:21a19db85fe7bc233b805147aa9ef6d116bc60e54279ba1e58b2c62be2c72c32

Observation 66df1fac-28a3-490d-ab5e-9321e882a381 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.370750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.967774Z digest=sha256:08cd81f9bcf01f8a6d342be84a9f049d5fb68c3af99787b6ed853d2bfa34cf8a

Observation 5f425da8-9735-4cc6-9291-eed2103a7e35 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.971449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.971449Z digest=sha256:ea11f2ecf91e68272fe915d975cb70a695e37dee28a91c330c8a9b51719f624b

Observation e4bcd0cc-39f7-4c3a-9b0d-7c18cb48a168 · outbound

This paper cites Flava: A foundational language and vision alignment model.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Flava: A foundational language and vision alignment model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.359516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.975175Z digest=sha256:f9ef1c72fc3e1cd8710bb128115b2e50cf5bc47d3863dd870f37e186e1ab848f

Observation 76d185f6-1d9b-4154-9f2f-b3e45f382573 · outbound

This paper cites Yfcc100m: The new data in multimedia research.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Yfcc100m: The new data in multimedia research

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.348027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.978562Z digest=sha256:93634e1d0f5b5ca91583a0be622463b64f5a3df540509a615793a0cd17933840

Observation 82a8c44b-5adc-4a9c-babc-da6136a43e44 · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.336292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.982108Z digest=sha256:b9cdb8e59318f9b1bd0788ba5114a2dca7e438d7e147172db12a3c7c8bbce7dd

Observation d6a37b8d-b988-4cf3-8ecc-a2c11124e0e8 · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.322934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.985873Z digest=sha256:a61b2af91205ad4d5f5b2caea7eddfb2f69aaae1f947096f494bea6685de9ad3

Observation eefbe169-f499-4469-9781-0b24b3d440ee · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.989662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.989662Z digest=sha256:b215fc7f4205084df521b19327e6e0be29a61ce8bb497dff934839a5cf150051

Observation 5d2c2618-0cf4-4fb8-a16d-838f722ac5b8 · outbound

This paper cites What you see is what you read? improving text- image alignment evaluation.Advances in Neural Informa- tion Processing Systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets What you see is what you read? improving text- image alignment evaluation.Advances in Neural Informa- tion Processing Systems, 36, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.311324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.993852Z digest=sha256:82514f08ac533e21b5a553aa39005d0eecfdb451b42814821eef4f14d3380986

Observation 3b220d1b-48cf-46f1-8116-12b54b6a27ea · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2023.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.299125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:52.997084Z digest=sha256:1c4dff00e1f4a7440f585be6d479eaf6513518cffbe365680005b7db89e7610d

Observation ed6d5e29-6cf9-4f0b-bf71-73cdb6b0e7d3 · outbound

This paper cites Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.000363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.000363Z digest=sha256:9c9727787a5d2f03d7cb012c77283caa010a5edb1ea5840dd0058c4f13dde716

Observation 9b14da76-f3fe-4ef7-9011-0ab3907d6449 · outbound

This paper cites Investigating compositional chal- lenges in vision-language models for visual grounding.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Investigating compositional chal- lenges in vision-language models for visual grounding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.287164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:53.003972Z digest=sha256:3b3cb49893386f519a73f0394a8c59c6c52ffa9e8a011c4a6dda38fbe008e468

Observation 6bec6939-e9bd-4efe-a87d-d4f9fac82942 · outbound

This paper cites Sigmoid loss for language image pre-training.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Sigmoid loss for language image pre-training

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.274342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:45:53.007487Z digest=sha256:59f72894a4da6bacd81a0f8bcdacfec10d356ee4002db4c480fc9aeec610d9b2

Observation b54f18df-cf1d-496b-8186-365b28cf9c0a · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.011960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.011960Z digest=sha256:7f7b4284f5eb79b734053b25bf4d8028a0da9750f9910f067fbcecee88be52eb

Observation 6faa031f-bd95-4ea0-9866-de705e61db24 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.015838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.015838Z digest=sha256:ee278396ded46f887e447434adb7165b03d1f8be091f40e121f020a11304f75d

Pith citing papers

No inbound Pith citation observations are available.