Pith. sign in

Paper Citation Record · LEDGER

Causal Graphical Models for Vision-Language Compositional Understanding

As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2412.09353.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09353 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:10:44.412768Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:07.218215Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:01:07.714487Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved20
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bbf244e4-63ba-4f69-9ac6-8f700047bdc3 · outbound

This paper cites head” word and its “dependent.

Causal Graphical Models for Vision-Language Compositional Understanding head” word and its “dependent

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.644892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.383990Z digest=sha256:7b92242ba5ef7d29dcbfb75a4183233125d9187ec59b3167d551aa945b61f2c4

Observation 201a70e7-41be-4705-bc70-aa5b1244e6bd · outbound

This paper cites Three teddy bears laying in a canopy bed under the covers.

Causal Graphical Models for Vision-Language Compositional Understanding Three teddy bears laying in a canopy bed under the covers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.563092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.410663Z digest=sha256:c331caf3bfcb1c2e9f45a67299ab691ff3d50eade5ce35624f6fbca77d2c98a3

Observation 825c0731-e720-4d20-8ff7-ddef98e39e14 · outbound

This paper cites ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation.

Causal Graphical Models for Vision-Language Compositional Understanding ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.316881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.316881Z digest=sha256:d2c97d0d62dfa4768e4ee0bde74e01b039ad0e6f804bf2f5578a1ec415ee7698

Observation cdd71c4c-c8f6-499b-998a-a70f5e1c2a22 · outbound

This paper cites Syntax-guided Localized Self-attention by Constituency Syntactic Distance.

Causal Graphical Models for Vision-Language Compositional Understanding Syntax-guided Localized Self-attention by Constituency Syntactic Distance

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:10:44.513979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.330776Z digest=sha256:579723d8ddd2e8feb758889fb089292e20b076c1da63422bcfe7b7ff6160f413

Observation 8a77a9fd-5413-4eeb-9b7a-fc16d4850385 · outbound

This paper cites ComCLIP: Training-Free Compositional Image and Text Matching.

Causal Graphical Models for Vision-Language Compositional Understanding ComCLIP: Training-Free Compositional Image and Text Matching

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-08-11T17:10:44.333379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.333379Z digest=sha256:409876696e9cb8b41d8a4de16032824ad0ef52ca517751f65a18b409e151499f

Observation bf57a8b0-04fe-4075-b07f-2f5ade4bdaf8 · outbound

This paper cites Segment Anything.

Causal Graphical Models for Vision-Language Compositional Understanding Segment Anything

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.335639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.335639Z digest=sha256:9cbf275d142d8b41c28bcf6c98fe9f6de0b8579efff84d4cbf6eea58c17d67e3

Observation a7ee15e4-391d-43a5-be99-fca1043ed5d4 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Causal Graphical Models for Vision-Language Compositional Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.683166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.338078Z digest=sha256:b8545de1fd2d994dc27754cce5274179f408b6ba01d6c3c9a41d00b55c956413

Observation 81baf5be-a9e1-4a81-975e-c2947915c3b7 · outbound

This paper cites Selective Attention Improves Transformer.

Causal Graphical Models for Vision-Language Compositional Understanding Selective Attention Improves Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.340058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.340058Z digest=sha256:67e1c9095d4f407c37f3bf19eee7d70fba4bad78da723d5ead1fa43af46f3618

Observation 9bf66ed8-85f3-4dfc-b021-690b0a3af09b · outbound

This paper cites Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna.

Causal Graphical Models for Vision-Language Compositional Understanding Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.677197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.344849Z digest=sha256:6cc19c865013703c76c1cd605deba3debf9ecd02b9a57200cbe9251ee40da57c

Observation dbc2fdf1-294a-405e-b6bc-92b1d73ae60d · outbound

This paper cites Rethinking Self-Attention: Towards Interpretability in Neural Parsing.

Causal Graphical Models for Vision-Language Compositional Understanding Rethinking Self-Attention: Towards Interpretability in Neural Parsing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.349412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.349412Z digest=sha256:d91f6d50907a7e622c27b9babce7e40e4c2e2e23f8c80c07f4d7769eef0f1473

Observation 56508762-49ad-4f89-8690-5929d7a40e62 · outbound

This paper cites SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels.

Causal Graphical Models for Vision-Language Compositional Understanding SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.355968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.355968Z digest=sha256:899d7c0a0a6dd47c3639529aed0dfeffa11234e4b17f54a27ab98c05c8e6db15

Observation 7f37d222-4a4f-445b-87c4-5422b71d5183 · outbound

This paper cites Coarse-to-fine contrastive learning in image-text-graph space for improved vision- language compositionality.

Causal Graphical Models for Vision-Language Compositional Understanding Coarse-to-fine contrastive learning in image-text-graph space for improved vision- language compositionality

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.671010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.359033Z digest=sha256:1f16d2411c58cd88eb070ff624702a685e8587842daeeb686d9740039361a104

Observation 69530690-b4b4-4d70-b092-d26aa33fe939 · outbound

This paper cites CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation.

Causal Graphical Models for Vision-Language Compositional Understanding CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.364395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.364395Z digest=sha256:b0d1abbe2846e859c0eacd3541b76ff6a7ac1526ab01811c30a42a9f61c75c21

Observation 6c8dc105-95b2-4711-a5f2-f7b5c1730207 · outbound

This paper cites Differential Transformer.

Causal Graphical Models for Vision-Language Compositional Understanding Differential Transformer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.366991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.366991Z digest=sha256:7ac230f26237d1fdd05742a0820bc3a08ea0765ea0a479150b824280932b5363

Observation f12e40bd-4537-459c-9a49-31e8a19d0625 · outbound

This paper cites 3VL: Using Trees to Improve Vision-Language Models' Interpretability.

Causal Graphical Models for Vision-Language Compositional Understanding 3VL: Using Trees to Improve Vision-Language Models' Interpretability

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.369610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.369610Z digest=sha256:1577bc7896dddc9048bba6c955ef808842766a4d77d1dd175417c1a566bd0bdd

Observation 2df23896-35c7-4108-b3f5-1628aab206ea · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

Causal Graphical Models for Vision-Language Compositional Understanding VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.372152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.372152Z digest=sha256:6abe28b54425254400ada13781cd98b6cc76f30513ea79fea12014e23c25c018

Observation 647d2543-82b7-4c7e-be38-0fa1213a05cb · outbound

This paper cites Iterated learning im- proves compositionality in large vision-language models.

Causal Graphical Models for Vision-Language Compositional Understanding Iterated learning im- proves compositionality in large vision-language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.664649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.374996Z digest=sha256:67435e8acde64317a42ff17d4702a7d2727b801d582e52f0c50e2a5b0b68684d

Observation ab116f1c-037a-4f64-a310-70309f2bece7 · outbound

This paper cites an unresolved cited work.

Causal Graphical Models for Vision-Language Compositional Understanding Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:44.658424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.377364Z digest=sha256:8367d3b648ec96252ea25b0ff05a22c3d55134e943febbc8bb76d8a7f53bc780

Observation 8d85ae82-75a8-48b4-bb63-74185cb427c0 · outbound

This paper cites The difference between a CGM and a Directed Graphical Model is that the former assumes that P A(Xj) are direct causes of Xj.

Causal Graphical Models for Vision-Language Compositional Understanding The difference between a CGM and a Directed Graphical Model is that the former assumes that P A(Xj) are direct causes of Xj

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.651866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.380490Z digest=sha256:b82427f1d027657520c7a28b1365d03e1683ffd64d4fc168862dc4f312b70a8a

Observation bd283f59-95b9-4072-a6c8-fe4f0d7071f3 · outbound

This paper cites an unresolved cited work.

Causal Graphical Models for Vision-Language Compositional Understanding Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:44.631397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.389014Z digest=sha256:0ba1e56dde7b74c4f5509ef32da1fd614bce6136034ad276897847225d76e964

Observation 0410b5f0-87db-4ba5-8795-663842f0e09b · outbound

This paper cites The same applies to those methods based on CLIP, such as NegCLIP (Yuksekgonul et al., 2023), GNM (Sahin et al., 2024), Plausible Adj.

Causal Graphical Models for Vision-Language Compositional Understanding The same applies to those methods based on CLIP, such as NegCLIP (Yuksekgonul et al., 2023), GNM (Sahin et al., 2024), Plausible Adj

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.624684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.391392Z digest=sha256:b74faae1a52e71af21954c752a49de05eddb82232828d48553b009f8aadf208d

Observation 3f9655f1-5bc5-40db-8631-9f795efe05b5 · outbound

This paper cites an unresolved cited work.

Causal Graphical Models for Vision-Language Compositional Understanding Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:44.616932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.393346Z digest=sha256:be800a5018d9fe28366a7bd246c3eecc2515724200206289763b210b5b817782

Observation 8ed8d0d0-c64e-43ce-bdc3-c3b504322b4f · outbound

This paper cites instruction.

Causal Graphical Models for Vision-Language Compositional Understanding instruction

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.610626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.395328Z digest=sha256:e1a3dee2c54c3e90accca6dae9b6a13621c1640f59fc20eb02588186344f8ed0

Observation f4b88e7f-e53c-4e5d-b067-06be1535f1e2 · outbound

This paper cites an unresolved cited work.

Causal Graphical Models for Vision-Language Compositional Understanding Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:44.604623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.397427Z digest=sha256:1f5f03fcb583b9ecf73766e3ff9b3cb96013ae6f5763088bc3be0eaf731bc2ee

Observation 53f213fa-6fde-48aa-8f99-11d5b67165eb · outbound

This paper cites Write a description for the photo.

Causal Graphical Models for Vision-Language Compositional Understanding Write a description for the photo

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.598741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.399533Z digest=sha256:ba6f1dabe8c91f603097f87a9f2f1a0d043a0372016440140fe5790daac873b7

Observation 281cbfd7-c3e7-4756-8559-003e45f132ec · outbound

This paper cites It is composed of two main tasks: Visual Genome Relation and Visual Genome Attri- bution.

Causal Graphical Models for Vision-Language Compositional Understanding It is composed of two main tasks: Visual Genome Relation and Visual Genome Attri- bution

Reference 37

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T17:10:44.591996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.402581Z digest=sha256:2d9f550944048ba450704dd82795602c66dfa610bd3faf81797793403890511f

Observation 4e5ace5d-99f9-4c6a-8d29-684a4dbe8c66 · outbound

This paper cites Replace”, “Swap.

Causal Graphical Models for Vision-Language Compositional Understanding Replace”, “Swap

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.584674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.404587Z digest=sha256:5523397961803679f82de62269f7ffffbe1e72ff12c89ff2817a9e9ea0178737

Observation 80ab0835-f486-408b-96be-b027dec86e77 · outbound

This paper cites Each image is associated with two descriptions: a true and a false caption.

Causal Graphical Models for Vision-Language Compositional Understanding Each image is associated with two descriptions: a true and a false caption

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.578231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.406416Z digest=sha256:4e6337ae646631a8585d279022c99662cef535deda0e89ca1d53136decd6a484

Observation 2e328e0b-25fa-4168-afc4-b1acb1f7562e · outbound

This paper cites color-swapped.

Causal Graphical Models for Vision-Language Compositional Understanding color-swapped

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.570998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.408328Z digest=sha256:6d278fd7bd7c3ade3937eb8d23ee3081a66be7d3c3e96f952d05d15a7053ed32

Observation 2e75a373-0573-43f3-bb9e-f0b7355e1839 · outbound

This paper cites an unresolved cited work.

Causal Graphical Models for Vision-Language Compositional Understanding Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:10:44.555722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.412768Z digest=sha256:2efad73d97a9cb711dc7d07759a09e4cf5a2c430f56432e889534afffe5e8af3

Observation e8844bff-36e6-4808-a45a-24dc834561c5 · outbound

This paper cites Specifically, in the Trivial task, negative captions are randomly sampled from unrelated objects (of different images), offering a basic challenge for retrieval.

Causal Graphical Models for Vision-Language Compositional Understanding Specifically, in the Trivial task, negative captions are randomly sampled from unrelated objects (of different images), offering a basic challenge for retrieval

Reference 224

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.638195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.386378Z digest=sha256:24b29038028375140631f25d58a2ad12cfdf0d5c2373aea6b17ff7a3dac040e9

Observation 84517f84-2ab7-41c3-87ec-ae41c01244ac · outbound

This paper cites Verbs in Action: Improving verb understanding in video-language models.

Causal Graphical Models for Vision-Language Compositional Understanding Verbs in Action: Improving verb understanding in video-language models

Reference 1993

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:10:44.480082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.347225Z digest=sha256:9c4e120afcb18ff5729a6f4dd3d0e3b32067f7f4b3c41228b4f66a78af4e9f58

Observation c25eb1c7-2e98-4100-baeb-4f8e20af2636 · outbound

This paper cites Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality.

Causal Graphical Models for Vision-Language Compositional Understanding Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.351701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.351701Z digest=sha256:ce70465cd39c9570fd6e9f234242412ab2399af7ca1a5ecd3cd3ab218659e521

Observation 2a1733a4-cc87-4e24-9482-dfc9c766bdcf · outbound

This paper cites Text-to-Image Diffusion Models are Zero-Shot Classifiers.

Causal Graphical Models for Vision-Language Compositional Understanding Text-to-Image Diffusion Models are Zero-Shot Classifiers

Reference 2014

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:10:44.536700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.319840Z digest=sha256:0846899d8865c307dfcddd8abf918b3a8cfeed1d11f4134838a391b85751b73b

Observation 55857164-6169-45dc-bdde-741f83c2d8d4 · outbound

This paper cites Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training.

Causal Graphical Models for Vision-Language Compositional Understanding Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.361548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.361548Z digest=sha256:a58af13a9ac68d283c6b0c28793f29954222d8b66e32865e9d6526b5ad72f9b9

Observation 908a225e-76d2-4368-96b6-cc2f5e864581 · outbound

This paper cites A generative dependency grammar.

Causal Graphical Models for Vision-Language Compositional Understanding A generative dependency grammar

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.688834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.322674Z digest=sha256:69f55cd2c1d0e942d3998d688151bbb620c265bcf1ca85c7a95caabfccbe74da

Observation 4d7426b4-84cc-4554-8330-4ebe450de04f · outbound

This paper cites What do Vision Transformers Learn? A Visual Exploration.

Causal Graphical Models for Vision-Language Compositional Understanding What do Vision Transformers Learn? A Visual Exploration

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.328119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.328119Z digest=sha256:861e97ee5800aefa051d4a4f752f9d4c3a05279e1a9b9de9bc6754b24e41c9e1

Observation 49858b05-ed72-4883-898b-c3696339e1ec · outbound

This paper cites Deep Biaffine Attention for Neural Dependency Parsing.

Causal Graphical Models for Vision-Language Compositional Understanding Deep Biaffine Attention for Neural Dependency Parsing

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.325149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.325149Z digest=sha256:3b10b7e0f506b30b241f859f58cca502bb08025afd93e0aaed32b09ab615abe5

Observation 4fa538cf-a18f-4f5a-8df9-c3f91b7a724b · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Causal Graphical Models for Vision-Language Compositional Understanding What If We Recaption Billions of Web Images with LLaMA-3?

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.342559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.342559Z digest=sha256:47474179c54027dc6f3149686554785a17e1f8544403361efa53414a492c3df7

Observation 9554776b-1b51-4296-adf9-b7e16ac15cdc · outbound

This paper cites Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP.

Causal Graphical Models for Vision-Language Compositional Understanding Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.311602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.311602Z digest=sha256:2559c17ebf01caf8007ececb134ca1066d46044885a92c960fe94d75c97d6f41

Observation 59aa08d0-dd88-4b2e-b6f4-7edee3a087db · outbound

This paper cites A hierarchical quasi- recurrent approach to video captioning.

Causal Graphical Models for Vision-Language Compositional Understanding A hierarchical quasi- recurrent approach to video captioning

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:10:44.694677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:10:44.314276Z digest=sha256:86eac9f65c876afcd5fc1c1faeeec346f01afed341ab4b7f644b66b9ca47219b

Pith citing papers

Observation 75ad03d2-c6ca-4845-a852-39b0a661e028 · inbound

CF-VLM:CounterFactual Vision-Language Fine-tuning cites this paper.

CF-VLM:CounterFactual Vision-Language Fine-tuning Causal Graphical Models for Vision-Language Compositional Understanding

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.719377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.218215Z digest=sha256:4d7eb7679e67250347b9dd282a0dc1ff5a604c429087d92ef3c3fcc9e3fb6bd2

Observation b1e25a89-abf9-437f-b3b4-88327a90ce22 · inbound

TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models cites this paper.

TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models Causal Graphical Models for Vision-Language Compositional Understanding

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-04T13:52:15.916851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:52:15.916851Z digest=sha256:85ce61ce7882475c754225b92110f83dd87f96b7ff8e17b24a511ee86a6ca571

Observation 576ef09b-e008-4df1-987f-8bad43299b75 · inbound

Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos cites this paper.

Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos Causal Graphical Models for Vision-Language Compositional Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T09:11:59.440912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:11:59.440912Z digest=sha256:e04325738a21756fd9b051823ec4b63b935256d38010ab30c1cbe516cb7703ea