Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:35.161833Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2506.08227.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:35.161833Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:51:55.155521Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:40:07.157025Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1315127a-eb00-4076-a473-6c4635f35f22 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Blindfold Baselines for Embodied QA
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 514ba2d7-166e-42c1-aa13-35a90f453105 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks VisMin: Visual Minimal-Change Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff43a8e0-0a48-4bd9-9b0b-7175df3f170c · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e456ef2e-6b3c-4028-9530-ef8c92b00ed7 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Evil- probe-a composite benchmark for extensive visio-linguistic probing
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f3b9892-4235-476c-92a3-98bac5295db2 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c9aacce-3096-449c-9206-cc24b38c8c1f · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a5409ad-f390-42cf-80ed-30fddc6f4908 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e06e66ca-e87f-4833-a575-c135a8440cab · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Routledge, 2016
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38918694-0f76-466d-950b-0770cb460411 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Sugarcrepe++ dataset: Vision-language model sensitivity to semantic and lexical alterations.Advances in Neural Information Processing Systems, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99754130-9e4a-4f79-ab05-62c52a5a0c9e · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Dat- acomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Sys- tems, 36:27092–27112, 2023
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a66031c-fb9e-4c71-a5f1-ff2f5e9d7c3c · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Shortcut learning in deep neural networks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2731f8b-0024-4f47-890d-68ad2b604ddc · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1136e946-f1a2-43d0-af6f-40eaf60bf8c0 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Agqa: A benchmark for compositional spatio-temporal reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 177fad96-974e-474e-9b39-9f75b9d9a461 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Probing Image-Language Transformers for Verb Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db1d596-d0b0-4a0f-b451-e56cfff7a7c1 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.NeurIPS,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5eb6ce73-4dc4-43e1-978c-e826b9d3dfdd · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Compositional Attention Networks for Machine Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 005709c4-5903-410d-8479-eb5772704c0e · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Text encoders bottleneck compositionality in contrastive vision- language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14cc2461-dbfe-4320-a637-c8ec647f4f39 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks What's "up" with vision-language models? Investigating their struggle with spatial reasoning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc2b2ef-b1fb-4778-aa38-06d7686a4dce · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks The hard positive truth about vision-language compositionality
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c5d8276-f12b-47e7-8b59-3f0301333ac8 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Clip behaves like a bag-of-words model cross-modally but not uni-modally.arXiv preprint arXiv:2502.03566, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89abbafb-e363-47cb-9007-e6bbc3e3731a · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Building machines that learn and think like people.Behavioral and brain sciences, 40:e253,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 458b3f24-8cbf-4a93-8b7d-60019b2fd7a3 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Coco- counterfactuals: Automatically constructed counterfactual examples for image-text pairs.Advances in Neural Infor- mation Processing Systems, 2023
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93a25d74-a6cf-4136-8f00-08f039e07271 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72da09fc-fc4c-44b3-87e4-9ccfdc15f9fd · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Remov- ing distributional discrepancies in captions improves image- text alignment
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dee2741e-8289-4c75-98ec-5d5620aa12a9 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Microsoft coco: Common objects in context
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae714ec9-4090-46e2-923a-b3a46c576065 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Vera: A general- purpose plausibility estimation model for commonsense statements
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa6a9c88-529b-4294-a2d2-21b848ec309f · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Crepe: Can vision-language foundation models reason compositionally? InCVPR, 2023
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81902300-1f71-49e1-b56c-c86878d31958 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Compositional chain-of-thought prompting for large multimodal models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ad48054-2851-4434-8254-b444e1fc6a33 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75edd83d-e0b1-49ea-a639-88d5cbf2042a · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2d9873b-fb69-4996-a5ea-eeeff6c3f773 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78dec15-7bb9-42b2-a5f5-a716e191633a · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Triplet- clip: Improving compositional reasoning of clip via synthetic vision-language negatives.Advances in Neural Information Processing Systems, 37:32731–32760, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e94f47ec-3ca1-44e7-bfb7-f685dbd980f8 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Learn- ing transferable visual models from natural language super- vision
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f17bcd-ba9a-48db-97ff-4fa4316e24c9 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks cola: A bench- mark for compositional text-to-image retrieval.NeurIPS,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6622fd16-4bca-4cb5-adce-d1f7c9e7cceb · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks ColorFoil: Investigating Color Blindness in Large Vision and Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01b8eed0-e6a3-46b4-9a55-64940aa8889f · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48eca0ff-40c0-49d5-a724-6e4abdaad982 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Teaching composition- ality to cnns
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 679d7e41-8db7-4673-8be4-4eb0ea8d00b2 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Winoground: Probing vision and language models for visio- linguistic compositionality
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3698515-14c1-45aa-a30f-59ac2fb2f97a · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6d1e98b-3f4c-44db-b8a4-a016b56c09a4 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2075b482-a3c3-4074-860f-e7c0d92f2ff9 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbb94afc-f6c8-4b71-bcd1-584bfa3c92e0 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ca2d9da-5c56-47dc-badc-5060aaec599a · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Equivariant similarity for vision-language foundation models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4cce305e-31aa-4962-8c5e-eaacd241025a · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c88e7dc5-7c2d-4a58-8584-5ff64e0d218f · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2022
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfcb0d5d-3e2b-4b3d-aa1b-a8702e7af27f · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Investigating compositional chal- lenges in vision-language models for visual grounding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2f2da20-3329-4311-a1f6-e899b82cfd34 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c230ce9-de7e-47c6-9477-98231b1afb40 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42821a0d-6d10-4218-a568-1379fe734f45 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cc44a3-8be7-42e3-864c-83382a985605 · outbound
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Iterated learning improves composition- ality in large vision-language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 773f9160-465e-4d0e-bd18-743ac80e09ce · inbound
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6547264c-3fe5-413d-aa13-6fc142f74696 · inbound
Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a79a2f97-e8b3-4ba6-8924-74995bc2423d · inbound
Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.