Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:10:44.412768Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2412.09353.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:10:44.412768Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:07.218215Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T05:01:07.714487Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bbf244e4-63ba-4f69-9ac6-8f700047bdc3 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding head” word and its “dependent
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 201a70e7-41be-4705-bc70-aa5b1244e6bd · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Three teddy bears laying in a canopy bed under the covers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 825c0731-e720-4d20-8ff7-ddef98e39e14 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd71c4c-c8f6-499b-998a-a70f5e1c2a22 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Syntax-guided Localized Self-attention by Constituency Syntactic Distance
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8a77a9fd-5413-4eeb-9b7a-fc16d4850385 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding ComCLIP: Training-Free Compositional Image and Text Matching
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf57a8b0-04fe-4075-b07f-2f5ade4bdaf8 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Segment Anything
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ee15e4-391d-43a5-be99-fca1043ed5d4 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 81baf5be-a9e1-4a81-975e-c2947915c3b7 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Selective Attention Improves Transformer
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bf66ed8-85f3-4dfc-b021-690b0a3af09b · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dbc2fdf1-294a-405e-b6bc-92b1d73ae60d · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Rethinking Self-Attention: Towards Interpretability in Neural Parsing
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56508762-49ad-4f89-8690-5929d7a40e62 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f37d222-4a4f-445b-87c4-5422b71d5183 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Coarse-to-fine contrastive learning in image-text-graph space for improved vision- language compositionality
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 69530690-b4b4-4d70-b092-d26aa33fe939 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c8dc105-95b2-4711-a5f2-f7b5c1730207 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Differential Transformer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f12e40bd-4537-459c-9a49-31e8a19d0625 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding 3VL: Using Trees to Improve Vision-Language Models' Interpretability
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df23896-35c7-4108-b3f5-1628aab206ea · outbound
Causal Graphical Models for Vision-Language Compositional Understanding VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 647d2543-82b7-4c7e-be38-0fa1213a05cb · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Iterated learning im- proves compositionality in large vision-language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ab116f1c-037a-4f64-a310-70309f2bece7 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8d85ae82-75a8-48b4-bb63-74185cb427c0 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding The difference between a CGM and a Directed Graphical Model is that the former assumes that P A(Xj) are direct causes of Xj
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bd283f59-95b9-4072-a6c8-fe4f0d7071f3 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0410b5f0-87db-4ba5-8795-663842f0e09b · outbound
Causal Graphical Models for Vision-Language Compositional Understanding The same applies to those methods based on CLIP, such as NegCLIP (Yuksekgonul et al., 2023), GNM (Sahin et al., 2024), Plausible Adj
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3f9655f1-5bc5-40db-8631-9f795efe05b5 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8ed8d0d0-c64e-43ce-bdc3-c3b504322b4f · outbound
Causal Graphical Models for Vision-Language Compositional Understanding instruction
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f4b88e7f-e53c-4e5d-b067-06be1535f1e2 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 53f213fa-6fde-48aa-8f99-11d5b67165eb · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Write a description for the photo
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 281cbfd7-c3e7-4756-8559-003e45f132ec · outbound
Causal Graphical Models for Vision-Language Compositional Understanding It is composed of two main tasks: Visual Genome Relation and Visual Genome Attri- bution
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4e5ace5d-99f9-4c6a-8d29-684a4dbe8c66 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Replace”, “Swap
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 80ab0835-f486-408b-96be-b027dec86e77 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Each image is associated with two descriptions: a true and a false caption
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2e328e0b-25fa-4168-afc4-b1acb1f7562e · outbound
Causal Graphical Models for Vision-Language Compositional Understanding color-swapped
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2e75a373-0573-43f3-bb9e-f0b7355e1839 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e8844bff-36e6-4808-a45a-24dc834561c5 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Specifically, in the Trivial task, negative captions are randomly sampled from unrelated objects (of different images), offering a basic challenge for retrieval
Reference 224
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 84517f84-2ab7-41c3-87ec-ae41c01244ac · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Verbs in Action: Improving verb understanding in video-language models
Reference 1993
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c25eb1c7-2e98-4100-baeb-4f8e20af2636 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a1733a4-cc87-4e24-9482-dfc9c766bdcf · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Text-to-Image Diffusion Models are Zero-Shot Classifiers
Reference 2014
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 55857164-6169-45dc-bdde-741f83c2d8d4 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908a225e-76d2-4368-96b6-cc2f5e864581 · outbound
Causal Graphical Models for Vision-Language Compositional Understanding A generative dependency grammar
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4d7426b4-84cc-4554-8330-4ebe450de04f · outbound
Causal Graphical Models for Vision-Language Compositional Understanding What do Vision Transformers Learn? A Visual Exploration
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49858b05-ed72-4883-898b-c3696339e1ec · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Deep Biaffine Attention for Neural Dependency Parsing
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fa538cf-a18f-4f5a-8df9-c3f91b7a724b · outbound
Causal Graphical Models for Vision-Language Compositional Understanding What If We Recaption Billions of Web Images with LLaMA-3?
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9554776b-1b51-4296-adf9-b7e16ac15cdc · outbound
Causal Graphical Models for Vision-Language Compositional Understanding Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59aa08d0-dd88-4b2e-b6f4-7edee3a087db · outbound
Causal Graphical Models for Vision-Language Compositional Understanding A hierarchical quasi- recurrent approach to video captioning
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 75ad03d2-c6ca-4845-a852-39b0a661e028 · inbound
CF-VLM:CounterFactual Vision-Language Fine-tuning Causal Graphical Models for Vision-Language Compositional Understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b1e25a89-abf9-437f-b3b4-88327a90ce22 · inbound
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models Causal Graphical Models for Vision-Language Compositional Understanding
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 576ef09b-e008-4df1-987f-8bad43299b75 · inbound
Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos Causal Graphical Models for Vision-Language Compositional Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.