Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:45:53.015838Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2507.04699.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:45:53.015838Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e541e223-0112-4f2b-bbc9-a9d01ba6a773 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 227d747f-b29e-408f-a3b9-4f62ce9cb1cb · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Vismin: Visual minimal-change understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3880f207-f9cf-4cb5-b486-27be3270d16f · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c97579e-5679-49a4-bffb-f94465bededd · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a65d044-fa8b-44fe-9df2-d56f09da7c57 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Clip2scene: Towards label-efficient 3d scene understanding by clip
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92458099-43b3-46b0-9043-596746f5c15e · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets A simple framework for contrastive learning of visual representations
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f519064-4f4e-4d59-9604-802108f68bf0 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1ed6828-d923-4682-b1cd-7e32a88ecf1c · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 255f7f21-4cc0-4c52-b8bb-4f8f972f554a · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 907d0124-c006-4743-8192-02908b83132e · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fe1138b-f0b8-424d-8daf-e48f4b130349 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2870a00f-fa24-45d3-8a1f-15b453fb94fe · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36:76137–76150, 2023
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53581e1d-5568-463a-a2ed-2a178875b130 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Teaching structured vision & language concepts to vision & language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3cbef5f-286e-4ca8-abc8-4dc10f2b40d7 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff1e9ff5-461c-410f-99de-bdf4a9bf2b57 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Scaling recti- fied flow transformers for high-resolution image synthesis
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc526a0e-7a79-435c-b2c1-9d0b1222b699 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Advances in deep concealed scene understanding.Visual Intelligence, 1(1):16,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7883666e-4ec6-4b91-a616-a090dfa88957 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff800481-a222-45c8-80db-7d4383db1b2e · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Visual program- ming: Compositional visual reasoning without training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 034cbaee-71a7-48be-bae9-c85b71a30a3a · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36, 2024
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37d2f436-fce2-4aec-9234-3c5a9aa599ab · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets LoRA: Low-Rank Adaptation of Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 784193e0-7fa7-4c41-970e-390885a53cf6 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Semantic to Structure: Learning Structural Representations for Infringement Detection
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56f713ab-86a7-4915-b5d1-57a18790bbda · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Secret lies in color: Enhancing ai-generated images detection with color distribution anal- ysis
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d19b3da0-5082-4a78-aa13-25f335ec430d · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Elevater: A benchmark and toolkit for evaluating language-augmented visual models.Advances in Neural Information Processing Systems, 35:9287–9301,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3398b2d4-3be4-426a-a529-90823f509958 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c50298-4162-40d0-a531-ad5605347184 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a3fd584-54a8-4495-8687-75df5a833d64 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Vila: On pre-training for vi- sual language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e46b6ba9-fbad-431f-a29a-0b7d99b39aa6 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2280f4f3-b2b0-4be1-838b-7c667339728c · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c1b60e1-3386-462d-a2d7-9aa857a4ddf5 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Synthesize diagnose and optimize: Towards fine- grained vision-language understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d93327d3-3d0c-4963-85bc-6477f84f3fa1 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66df1fac-28a3-490d-ab5e-9321e882a381 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Learning transferable visual models from natural language supervi- sion
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f425da8-9735-4cc6-9291-eed2103a7e35 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4bcd0cc-39f7-4c3a-9b0d-7c18cb48a168 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Flava: A foundational language and vision alignment model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76d185f6-1d9b-4154-9f2f-b3e45f382573 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Yfcc100m: The new data in multimedia research
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82a8c44b-5adc-4a9c-babc-da6136a43e44 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Winoground: Probing vision and language models for visio- linguistic compositionality
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6a37b8d-b988-4cf3-8ecc-a2c11124e0e8 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eefbe169-f499-4469-9781-0b24b3d440ee · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872, 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2c2618-0cf4-4fb8-a16d-838f722ac5b8 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets What you see is what you read? improving text- image alignment evaluation.Advances in Neural Informa- tion Processing Systems, 36, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b220d1b-48cf-46f1-8116-12b54b6a27ea · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2023
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed6d5e29-6cf9-4f0b-bf71-73cdb6b0e7d3 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b14da76-f3fe-4ef7-9011-0ab3907d6449 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Investigating compositional chal- lenges in vision-language models for visual grounding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6bec6939-e9bd-4efe-a87d-d4f9fac82942 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Sigmoid loss for language image pre-training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b54f18df-cf1d-496b-8186-365b28cf9c0a · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6faa031f-bd95-4ea0-9866-de705e61db24 · outbound
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.