Pith. sign in

Paper Citation Record · LEDGER

Multi-Sourced Compositional Generalization in Visual Question Answering

As of 9 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2505.23045.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23045 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:50.572037Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b18421b8-bae5-4053-9afc-8cb321e7af05 · outbound

This paper cites Robust visual rea- soning via language guided neural module networks.

Multi-Sourced Compositional Generalization in Visual Question Answering Robust visual rea- soning via language guided neural module networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:55.251531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:47.268997Z digest=sha256:abb55e36551533720329367524a189a705ad44f5a9acdab3606af5da6b6b0388

Observation c0bd8e5b-352e-47f5-a67e-68d015b84a34 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multi-Sourced Compositional Generalization in Visual Question Answering Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.776858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.776858Z digest=sha256:0c3fefcd8cedfca262999659198f5b3649a7bc49f50f945ca3a623b02af4e1e4

Observation b4a17059-f47e-4e7b-9d6a-569fd1894cda · outbound

This paper cites The paradox of the compositionality of natural language: A neural machine translation case study.

Multi-Sourced Compositional Generalization in Visual Question Answering The paradox of the compositionality of natural language: A neural machine translation case study

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:54.236317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:48.284091Z digest=sha256:0b8835944e561d578b21b2cb30ffd9d0def8a2a72ad4dd0150c88c4891df4191

Observation 8940b1cd-16e3-4832-ba82-3ad6cfdb8cb6 · outbound

This paper cites Compositional attention networks for machine reasoning.

Multi-Sourced Compositional Generalization in Visual Question Answering Compositional attention networks for machine reasoning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:53.562432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:48.798804Z digest=sha256:643db95a98f207f98b40f4c2e223b5d0f33209dbbe6f70581acc810c66ec8704

Observation d3fdc2e3-c1cf-452e-aaf6-3180e7314be8 · outbound

This paper cites Gqa: A new dataset for real-world vi- sual reasoning and compositional question answering.

Multi-Sourced Compositional Generalization in Visual Question Answering Gqa: A new dataset for real-world vi- sual reasoning and compositional question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:53.412975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:48.912036Z digest=sha256:c7122689d653569ce6e01c3a025ff016be8cb3004bd4bfee1f9d5ebcb0b22e62

Observation c4eeb001-8e61-4eb1-b42a-54a8daf4f156 · outbound

This paper cites Overcoming language pri- ors in vqa via decomposed linguistic representations.

Multi-Sourced Compositional Generalization in Visual Question Answering Overcoming language pri- ors in vqa via decomposed linguistic representations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:53.224012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:49.026408Z digest=sha256:e3e45c3e3a7006cf8309658307ef94d1600a1af94441908908c87e446bda41c8

Observation 429d72f8-49df-476c-b1c1-620e3924c0a1 · outbound

This paper cites On compositional generalization of neural ma- chine translation.

Multi-Sourced Compositional Generalization in Visual Question Answering On compositional generalization of neural ma- chine translation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.930710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:49.312110Z digest=sha256:bc95ab3153ab199b0d35ff69fefd4f36b663b47c8e49d6b225a94120e481f056

Observation e0913540-94e4-4ec0-8459-a5a7560ddbf9 · outbound

This paper cites Compositional temporal ground- ing with structured variational cross-graph correspondence learning.

Multi-Sourced Compositional Generalization in Visual Question Answering Compositional temporal ground- ing with structured variational cross-graph correspondence learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.702855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:49.459425Z digest=sha256:fe7a7e12ff10a173e332cde2588920447ecc61b3b370cc3f57b5b9531e93c20c

Observation eb087c05-4a06-4060-9124-558dca9fa6e8 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Multi-Sourced Compositional Generalization in Visual Question Answering BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:49.583802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:49.583802Z digest=sha256:66d26aab3e127d75b7106b80dce2f8f240024b08b190a9af75fc987e2581d2a8

Observation 609bb895-7b91-4415-bb51-01fc2ece2c4d · outbound

This paper cites Improved baselines with visual instruction tuning,.

Multi-Sourced Compositional Generalization in Visual Question Answering Improved baselines with visual instruction tuning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.492842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:49.700193Z digest=sha256:be3d5c33d74745ad8114fce9490baab30edf96f48c3015a5c1e2e5973fd124db

Observation 1be7d0da-f113-40fc-8fcd-6961a2a70cac · outbound

This paper cites Overcoming language priors in visual question answering with cumulative learning strategy.Neurocom- puting, 608:128419,.

Multi-Sourced Compositional Generalization in Visual Question Answering Overcoming language priors in visual question answering with cumulative learning strategy.Neurocom- puting, 608:128419,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.339922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:49.804989Z digest=sha256:6a13e720727eab551bb44aa04d7fd9ba8c17f13d0ae2c0dcd582e2ba6bb1980c

Observation 8451fc69-76bf-430f-9820-732a71aa2046 · outbound

This paper cites Learning graph embeddings for compositional zero-shot learning.

Multi-Sourced Compositional Generalization in Visual Question Answering Learning graph embeddings for compositional zero-shot learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.201847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:49.904424Z digest=sha256:54f5e1dea53d7ca2823e8a90538744fac249e025fdaf08fa65446e1e52603524

Observation 5013e0f5-cae0-438c-aad6-ab0b066194e8 · outbound

This paper cites Coarse- to-fine reasoning for visual question answering.

Multi-Sourced Compositional Generalization in Visual Question Answering Coarse- to-fine reasoning for visual question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.989430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:49.983404Z digest=sha256:ae7c9b08799ce94a56d4556849f97f5f9b0d82ee5adfe438c11634354f1fdd36

Observation 2f69b773-4bc7-4254-aada-d63274af7adf · outbound

This paper cites Counter- factual vqa: A cause-effect look at language bias.

Multi-Sourced Compositional Generalization in Visual Question Answering Counter- factual vqa: A cause-effect look at language bias

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.807706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:50.085036Z digest=sha256:dd7c6a0a6dd1737da6378fee31a2c39f84e07dd43e851c43a724324fb6d7ec3c

Observation 3ff5b3af-7a32-4a3e-a752-e55294efe0bf · outbound

This paper cites Combine to describe: Evaluating compositional generalization in image captioning.

Multi-Sourced Compositional Generalization in Visual Question Answering Combine to describe: Evaluating compositional generalization in image captioning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.535095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:50.199695Z digest=sha256:9d1f80e187ac82aa9d0cd828f281137d973785f45cc69cea678c20d670c7c966

Observation 5f597684-cc11-4a2a-966a-ae8ae354ffb1 · outbound

This paper cites Overcom- ing language priors for visual question answering based on knowledge distillation.

Multi-Sourced Compositional Generalization in Visual Question Answering Overcom- ing language priors for visual question answering based on knowledge distillation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.371537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:50.292955Z digest=sha256:721525bdcf628a959dc4b1f9604a4a91d26e5f484e171ffc94c866629516a5d3

Observation 5e669762-a664-4ace-b641-ec6ef45f7009 · outbound

This paper cites Faster r-cnn: Towards real-time ob- ject detection with region proposal networks.IEEE Trans- actions on Pattern Analysis and Machine Intelligence (T- PAMI), 39(6):1137–1149,.

Multi-Sourced Compositional Generalization in Visual Question Answering Faster r-cnn: Towards real-time ob- ject detection with region proposal networks.IEEE Trans- actions on Pattern Analysis and Machine Intelligence (T- PAMI), 39(6):1137–1149,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.205858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:50.359521Z digest=sha256:6b3c41656bd847c1c4100cfe6feb50444ca6c012e72b95cd0ce3b612673ecef5

Observation 4504748f-f741-42fa-bbde-df652f9351cd · outbound

This paper cites Transformer module networks for systematic generaliza- tion in visual question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence,.

Multi-Sourced Compositional Generalization in Visual Question Answering Transformer module networks for systematic generaliza- tion in visual question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:50.800820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:50.467823Z digest=sha256:767775e38ba24a24ebe39aeb50b221cfcc0dd6b1a22177cdbee329b2f2889ac4

Observation 693b9520-dbfa-4d01-87b8-5a79b1334760 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Multi-Sourced Compositional Generalization in Visual Question Answering MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:50.572037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:50.572037Z digest=sha256:115ab5d68e885c68d3e79cc91d010659d1d0ec139674dfa424f9f501ff94b403

Observation e97128ac-6ca4-48bd-a616-18c832186732 · outbound

This paper cites Composi- tional generalization for multi-label text classification: A data-augmentation approach.

Multi-Sourced Compositional Generalization in Visual Question Answering Composi- tional generalization for multi-label text classification: A data-augmentation approach

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:54.603085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:48.014768Z digest=sha256:fefcf11b7efabbdbfe07ba5f9ae9194ba256b1c5665291074cbbf1bb25b72d7f

Observation aae0f6d1-b5d3-4cf8-a230-5e53e985613f · outbound

This paper cites Systematic Generalization: What Is Required and Can It Be Learned?.

Multi-Sourced Compositional Generalization in Visual Question Answering Systematic Generalization: What Is Required and Can It Be Learned?

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.666813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.666813Z digest=sha256:e0e8d439d313ccbfa47c9565807f80a07760e86069cc259b60339e9c3c521403

Observation 20bd31ab-51f2-4100-aaf3-e99e238a3cfc · outbound

This paper cites Plug-and- play vqa: Zero-shot vqa by conjoining large pretrained models with zero training.

Multi-Sourced Compositional Generalization in Visual Question Answering Plug-and- play vqa: Zero-shot vqa by conjoining large pretrained models with zero training

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:50.956447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:50.395299Z digest=sha256:93fa34614a335234d7d582be9307864ac4b976b821934ed8a6c1cb1634af0a06

Observation 57b6876b-fe2b-4f18-b8f6-9f2eb38fc97d · outbound

This paper cites Language-conditioned graph networks for relational reasoning.

Multi-Sourced Compositional Generalization in Visual Question Answering Language-conditioned graph networks for relational reasoning

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:53.872942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:48.556333Z digest=sha256:2a80a837c18b85edb5105816eebb5a14eb87596dcba5c47d1c000ef8f48d72f3

Observation 9fad416c-f1fa-41a0-94ad-5b92de48b6c4 · outbound

This paper cites Vqa: Visual question answering.

Multi-Sourced Compositional Generalization in Visual Question Answering Vqa: Visual question answering

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:54.987320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:47.507150Z digest=sha256:4fcca86f5653ceb9c565ab6a9d46475fcf49f642b7a9c2e4375a44a57af6a14e

Observation 75cbd23e-776f-4e65-8421-0d99ae9fbce9 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Multi-Sourced Compositional Generalization in Visual Question Answering LoRA: Low-rank adaptation of large language models

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:53.700888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:48.647648Z digest=sha256:64172ee9404f93508587bec1a4c8d2c0fca038b742b6486e963b3e7e54ebfe59

Observation 48fe8c11-38a7-435c-83d9-612be10ca19a · outbound

This paper cites Retrieval-augmented primitive representa- tions for compositional zero-shot learning.

Multi-Sourced Compositional Generalization in Visual Question Answering Retrieval-augmented primitive representa- tions for compositional zero-shot learning

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:53.082479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:49.168836Z digest=sha256:2c6becf8a225da9a8cd6b61f4340f6b577b5d55f4558df6bc5985eb66cd3fa6d

Observation 4d615ff1-d4ad-49bf-80a8-cb9ff6f181f5 · outbound

This paper cites Bottom-up and top-down attention for im- age captioning and visual question answering.

Multi-Sourced Compositional Generalization in Visual Question Answering Bottom-up and top-down attention for im- age captioning and visual question answering

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:55.146044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:47.360429Z digest=sha256:73951d40670e4e135278ea46ec4c98610d980f379b1c0e5d0bb58ddedd0824ed

Observation 76936dd3-aa10-40ab-86d3-787576274996 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Multi-Sourced Compositional Generalization in Visual Question Answering Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:54.024402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:48.421701Z digest=sha256:a5e308a1abc7d8ae1df0b528f1ba0422b66eb29d8e82c261c3060349d95a372f

Observation f68d00ab-f7aa-4635-8185-ce5b4f722839 · outbound

This paper cites ” O’Reilly Media, Inc.”,.

Multi-Sourced Compositional Generalization in Visual Question Answering ” O’Reilly Media, Inc.”,

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:54.785548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:47.877093Z digest=sha256:331ce85e2d819eac0af0c38740ec81768e2ebc8b76c57311381214e04e2362f6

Observation 29fe08f9-f844-4184-b745-8224b9f0eac0 · outbound

This paper cites Meta module network for compositional visual reasoning.

Multi-Sourced Compositional Generalization in Visual Question Answering Meta module network for compositional visual reasoning

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:54.388196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:59:48.112006Z digest=sha256:175db1af2ac3486320d1a350cfcb7178638f74694adf9892321c181b241c7b1e

Pith citing papers

No inbound Pith citation observations are available.