Pith. sign in

Paper Citation Record · LEDGER

Visually Interpretable Subtask Reasoning for Visual Question Answering

As of 16 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2505.08084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08084 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:10:20.537125Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8d57195-3bbb-4101-9a88-404d38b7388c · outbound

This paper cites GPT-4 Technical Report.

Visually Interpretable Subtask Reasoning for Visual Question Answering GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.239076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.239076Z digest=sha256:6e57faea140fed1d439799da8f2cb77d0e80207c94d8d80df725bb61775c3c45

Observation b18430cf-a088-42ea-bee7-70789ea7f3a9 · outbound

This paper cites Qwen Technical Report.

Visually Interpretable Subtask Reasoning for Visual Question Answering Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.244796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.244796Z digest=sha256:9ccc15e152cf85e21c06c801b020521c61bde10c02199ab31433623a57672d12

Observation 51d7f70d-9620-4b87-a3bf-1b386f67adb4 · outbound

This paper cites Language Models are Few-Shot Learners.

Visually Interpretable Subtask Reasoning for Visual Question Answering Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.249973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.249973Z digest=sha256:c282243af4d33e85bd58810fb5df05409ea86ade366b84505cfde60f6289b71c

Observation 095b047a-0e4d-4f62-933f-5456fd8b3a0d · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Visually Interpretable Subtask Reasoning for Visual Question Answering Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.255873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.255873Z digest=sha256:e62d984c7269ee42e3302adcf797be1c39197bfed3ca6a112ab96fb36cbbd70d

Observation 130b10ec-0287-4829-85c5-c75def6851eb · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Visually Interpretable Subtask Reasoning for Visual Question Answering Gonzalez, Ion Stoica, and Eric P

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.260923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.260923Z digest=sha256:07698679efe5c5486fe3fdcbf4dde49ed3d63025f70d132bbd5f088a4e4956b6

Observation 12394637-78d5-44ce-a0c1-3a8e0f1b3c67 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Visually Interpretable Subtask Reasoning for Visual Question Answering Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.500334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.265480Z digest=sha256:3c8d51b107047cbb1e3c868af6f3b1969da44c228983337eac5ba434856f4e97

Observation 52258c8f-1b57-4fe1-a1d4-17e319040724 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.270682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.270682Z digest=sha256:a1b25645d87e33df0c45ca88e7b09844871683ee25fc98f2f87dac7458e5e8b5

Observation c8acdc8e-a4f6-4348-861c-9478d3f69d88 · outbound

This paper cites Cric: A vqa dataset for compositional reasoning on vision and commonsense.

Visually Interpretable Subtask Reasoning for Visual Question Answering Cric: A vqa dataset for compositional reasoning on vision and commonsense

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.484735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.275366Z digest=sha256:14c69313202f69288ccd14a4d3cc1d16d64d6ccae4e1f7294edfc5233613204c

Observation 022bf1e4-0d4c-41d4-8e49-8103c27707cf · outbound

This paper cites OmniFusion Technical Report.

Visually Interpretable Subtask Reasoning for Visual Question Answering OmniFusion Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.279770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.279770Z digest=sha256:e6def2770073772cd46fe805bef59483565b1c3e63d1ea1895a4da1b2b746dd1

Observation d6fe1352-bf80-47a9-ad02-1d5ff373d2f8 · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

Visually Interpretable Subtask Reasoning for Visual Question Answering Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.469059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.284387Z digest=sha256:badaa3460875fa90cbc7951b5d5ab44b44a3cbefd9ae4a89cd23ea205616bf44

Observation d6104dc0-7c3e-4659-81d1-acfa488dd8be · outbound

This paper cites Visual pro- gramming: Compositional visual reasoning without training.

Visually Interpretable Subtask Reasoning for Visual Question Answering Visual pro- gramming: Compositional visual reasoning without training

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.454740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.289765Z digest=sha256:6f86e3ec36657ccec339bcc41933e4deada6845db02056d8c7cc1bc68a09b4fd

Observation 415be4a6-2aea-444b-ba2e-fa79b0e12c3a · outbound

This paper cites Visual program distillation: Distilling tools and programmatic reasoning into vision-language models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Visual program distillation: Distilling tools and programmatic reasoning into vision-language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.440595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.294585Z digest=sha256:65c2b73a1db70d3cdbb054daa4dfd19aa6e7a6803a3b3f0120e6fe2d00ce8184

Observation 1d3a3e97-46a7-4e47-a07e-74cf9850c2dd · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Visually Interpretable Subtask Reasoning for Visual Question Answering Vtimellm: Empower llm to grasp video moments

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.424498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.299056Z digest=sha256:2cdd2db2f1f35ca8dc03bb5eb227d04d476af2f78d5ab515b3fe1461f48a9bfb

Observation 11d53f17-102f-4614-b928-e931eeb79990 · outbound

This paper cites Hudson and Christopher D.

Visually Interpretable Subtask Reasoning for Visual Question Answering Hudson and Christopher D

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.409593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.303384Z digest=sha256:11ee60e2deb3a18121eec80d9c35234fa62895097f6ddc902441e3adb47e20cb

Observation 1ec608ae-d00b-491e-acc8-c7e38c86d317 · outbound

This paper cites Unveiling the Invisible: Captioning Videos with Metaphors.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unveiling the Invisible: Captioning Videos with Metaphors

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.307370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.307370Z digest=sha256:b3a2269db2cf1bf938565168f297bc5a4ba9c656cc6ab64bb8137851aea73250

Observation 9f8812c3-d6bd-4c66-a40b-97dbea6cdb58 · outbound

This paper cites Hydra: A hyper agent for dynamic compositional visual reasoning.

Visually Interpretable Subtask Reasoning for Visual Question Answering Hydra: A hyper agent for dynamic compositional visual reasoning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.394767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.312068Z digest=sha256:9988e45371f139ecc3cddf4088eb1102da195d8b01059bfc4d51ac67a4106740

Observation 03a10c82-7b4b-4064-938e-72f2afd0b73f · outbound

This paper cites Exploring question decomposition for zero-shot vqa.

Visually Interpretable Subtask Reasoning for Visual Question Answering Exploring question decomposition for zero-shot vqa

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.380548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.316613Z digest=sha256:c3ad31323ebed5baed5fe1020fd1fe2a46bb83378ab0d3dbfb0b49f96288e5a7

Observation e2bf1655-072d-4f17-a59b-36cf535a8729 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering A Survey on Benchmarks of Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.321108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.321108Z digest=sha256:6ec11f0255e599a825f4bb44569ad809560f9bbcc8541ae3ba009a158687f537

Observation 2dbf9bf1-7862-4921-b6e7-f7b041da87a2 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.364463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.326138Z digest=sha256:4ffb17ba2ac4e49d21d27319970990458f835815dfc01b06f038cb43c02e7173

Observation 82eaea41-d40d-432f-92cf-32bd7cc8cba7 · outbound

This paper cites Grounded language-image pre-training.

Visually Interpretable Subtask Reasoning for Visual Question Answering Grounded language-image pre-training

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.349464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.330701Z digest=sha256:f02a476108a0e6792393050b685895eac3a0a7745473016ffc0fd0a174b47ec5

Observation 8422c373-7620-4f7d-b1d2-d2e6a6fa6351 · outbound

This paper cites Improved baselines with visual instruction tuning.

Visually Interpretable Subtask Reasoning for Visual Question Answering Improved baselines with visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.334420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.335097Z digest=sha256:31b58307c616388ba28a27f6d71cf642bf8b997b2f6ef1ba9fbd848b01154fd6

Observation 2a78b72e-4834-4bfe-8a47-5fdb40126675 · outbound

This paper cites Visual instruction tuning.

Visually Interpretable Subtask Reasoning for Visual Question Answering Visual instruction tuning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.319490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.339444Z digest=sha256:ef3413108810671e7b724bbf5c0bb4286b68110a782cbfba86a04bd451c75cbb

Observation 714d25c7-6672-4128-9412-8bbeed9f8543 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering NVILA: Efficient Frontier Visual Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.343778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.343778Z digest=sha256:f9084eb10b50b9fc01d6d708fcaa393a519e8eb6773a48ff865abcf7f73ed818

Observation a2da71b1-6def-4915-b75b-e768b7b39b48 · outbound

This paper cites SAE-V: Interpreting Multimodal Models for Enhanced Alignment.

Visually Interpretable Subtask Reasoning for Visual Question Answering SAE-V: Interpreting Multimodal Models for Enhanced Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.348244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.348244Z digest=sha256:8666afd24a9c0e36bfa9df5779197af52ed396fb1ede8de160feb5aed73feaf3

Observation 3314594d-8828-4ee9-a95a-72dd0bd8c313 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Groma: Localized visual tokenization for grounding multimodal large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.304205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.353038Z digest=sha256:9cf967789c5e30980ac85c4fd5f0e4c0c92a8762889973a5eba4c7d4aea51c75

Observation 8dd6e76c-f00d-4fce-a56f-341a5f0b2c00 · outbound

This paper cites Task navigator: Decomposing complex tasks for multimodal large language models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Task navigator: Decomposing complex tasks for multimodal large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.289200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.358522Z digest=sha256:95254267a6aadca5d83ec1695cc4daf88922293f128eb53baad382e6717f433a

Observation 098a1cc4-0926-49c3-a0b9-961a7620bcbc · outbound

This paper cites Language models are unsu- pervised multitask learners.

Visually Interpretable Subtask Reasoning for Visual Question Answering Language models are unsu- pervised multitask learners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.364298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.364298Z digest=sha256:86ca6deed16c721fbcc8f38b9f01da55a6008d4f539aa6c47fa96e6c0ac1e002

Observation 56242bd3-f485-4a5a-b022-16709cf86950 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Visually Interpretable Subtask Reasoning for Visual Question Answering Learning transferable visual models from natural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.369214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.369214Z digest=sha256:c0805f72a30afca5beb942becc1cc445138510eb5e8c3908591a520dadd03db8

Observation 78eaef3a-5e03-43eb-9044-0c9d17e84c99 · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

Visually Interpretable Subtask Reasoning for Visual Question Answering Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.254395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.374649Z digest=sha256:7380d1a28319427a72bdd5dd9fbef292154c66e19987dc3bc431fcb8b2f22e5f

Observation 7fb21a7c-3602-48e5-b007-df7500a0dc38 · outbound

This paper cites Reid, and Silvio Savarese.

Visually Interpretable Subtask Reasoning for Visual Question Answering Reid, and Silvio Savarese

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.239321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.379570Z digest=sha256:666dcfbff244dfe2b3b0b1ce34e9d5ed2ac307873377620d69696ffd4dd13c51

Observation 83604e42-d8a2-422a-843b-a44ed6d31051 · outbound

This paper cites Towards vqa models that can read.

Visually Interpretable Subtask Reasoning for Visual Question Answering Towards vqa models that can read

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.224403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.384030Z digest=sha256:fc0152c66c7f0c9b20c1ebaf2012c3d97be8a8fe2b19221f124ec6c937598c4c

Observation a56892f6-6ff5-441c-ba6e-6727218c0fc3 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Visually Interpretable Subtask Reasoning for Visual Question Answering Vipergpt: Visual inference via python execution for reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.209077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.387978Z digest=sha256:703235023efb4488d6a71944f8f6c86fe877b126c93e28a360d4102cf3576d82

Observation 88fa69f3-f3f3-42c7-a819-337d262f4558 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.392320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.392320Z digest=sha256:01ee03e6511fdf08337be0689e2866c9c382b39b8c59a36fbc37df29c7ebb1a9

Observation 99dc1080-943a-4b33-bf29-d3a1c23c8a73 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Visually Interpretable Subtask Reasoning for Visual Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.397283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.397283Z digest=sha256:f325926c35e30c165b4ae8509b957c491c3bed70ed77078bccdb925cec022099

Observation 80127df8-7561-464f-9fd8-d4de5a3cc5f5 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Visually Interpretable Subtask Reasoning for Visual Question Answering Emu3: Next-Token Prediction is All You Need

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.402206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.402206Z digest=sha256:f6a1fb76ebc17b646bd677e57239990878cb3a8636d2a776a820534788387068

Observation cc611ff7-469e-4ed2-a66f-4f723daf1d09 · outbound

This paper cites Sigmoid loss for language image pre-training.

Visually Interpretable Subtask Reasoning for Visual Question Answering Sigmoid loss for language image pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.193319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.406738Z digest=sha256:19e959b892bf724c0b48238438ac3caccc6a29d8bcae1c1d117060412e80c409

Observation 9e7636c1-ccd7-4748-b7f2-14d408a1f048 · outbound

This paper cites Visual Question Decomposition on Multimodal Large Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Visual Question Decomposition on Multimodal Large Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:10:20.627569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.410966Z digest=sha256:12dbdd32fa1898dd4442abf6a28b191e32622f5a41a6ff0110cb4e76252dd715

Observation 3f517add-8723-4f26-893e-fb573d440949 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.415537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.415537Z digest=sha256:ad4eebf311f15436976c24484f8d1a5c873a742d3b4e68e3e31245f6e28b4c2d

Observation c653e165-58e0-41d2-abd5-5d228e10c501 · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Visually Interpretable Subtask Reasoning for Visual Question Answering MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.420225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.420225Z digest=sha256:2c9911046d6e9380bd486dd86149add7db78d22523e9f3bd9100ca737949287d

Observation fd230b6e-d12d-4a36-835b-306f63f975c9 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.425142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.425142Z digest=sha256:bfca527abfdc6723c32fa564bd57e0fbf3a6aaf73e57f2a056f2b930420e163b

Observation 15e4033c-4b1e-480a-a8ee-0f54177fe556 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.179212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.430005Z digest=sha256:d180bc13aad1d56752747353499477f58a8ea8093719d005d3dc809019ce66c9

Observation a1bc3843-4473-4ec6-810a-652359377a16 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.163826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.434618Z digest=sha256:6c0212e97c71972666f12704532efb425a8e8e832e1024a8fd70af14f70f5bf1

Observation 24154824-e96c-4800-aee5-c2399b8ee17f · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.148370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.440312Z digest=sha256:859c81ee7a4ba610e1f128f52d257aa70202ca9f95017d0d5191acd2ab1f9712

Observation 55b571ae-f6ec-4e07-a902-966a13d8760e · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.133967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.445059Z digest=sha256:e4d7ded6c8f01900c1f9be9c37e879d08f79e3c5059a6fd7889c2780dfb4720c

Observation c0081a71-aa4e-47f9-8540-025afa851375 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.119511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.449493Z digest=sha256:4b46ec02ec5fbe6c2347b93f2cf126ce1e20597b509ec8693aa630a3dbb5a53b

Observation aa8aa685-70b3-4c5d-84ed-cf5f06946d2b · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.104652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.455126Z digest=sha256:c21ca6ba0c4c539b9392e71aaf64cdac00130417e295e0014c46a3ac3c05cb57

Observation bcc2dedb-6211-439c-8a98-893a60aa2805 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.090074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.459528Z digest=sha256:7ae5553a4ea35e448b70b857c1cce1f9f1b5cfc178ecab935cf1d3ef7528776c

Observation c3a13de9-2ac2-4ed5-8023-a54cf1a0985b · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.074152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.464202Z digest=sha256:b8f545e3293e2ad53879d5425a16d9b478bfd94b4f0f84662d1133d0907916d6

Observation 21dace39-e704-4965-8530-0ff3222ee89e · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.058533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.468971Z digest=sha256:9fdf2716f4b655153d4949d7dfd19c6a8d311289872c85b6450b74b71ab63806

Observation ee11851a-e4a0-4335-8be7-2f146bdc28e9 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.043097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.474476Z digest=sha256:bb4d832586dd2afa9e2b1c881f364f0b7e1f07ddde2f2aba2f7a2e663d11348f

Observation 431f83bc-9237-4d31-8344-0b339c2e1b6a · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.027802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.479284Z digest=sha256:e770beacb9cb1ffe865f07a549f77a4ee7d47d97e5e16fa9694fc29837f5331d

Observation 77845619-6325-4233-80eb-14ee0e682b38 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.013149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.483962Z digest=sha256:c0767231559171985a76f3d146352454d85d53f44df57d0c396fe3e9220a4f2e

Observation 9da2bdb1-3efa-4246-a605-3cb69671eb11 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.998129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.488975Z digest=sha256:92f108a63d1b6bfe7154c20043a519dc3d3b1e0791f067448d1374bf35504702

Observation f641b501-ac5a-4a7b-a387-838ba6ba578c · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.982509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.493679Z digest=sha256:dbaabb63d2abb0d3b1b0a8d99497fef91f46adab199d2880da8ac247d5420180

Observation ff43abc4-5a44-4201-9788-d387e55320e8 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.966695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.498344Z digest=sha256:1cd98b940828a654889f2f9f5921d06e48f0043eade1ed273afd75e3a258a04f

Observation caeaaf21-d7ea-46eb-aa83-801862f25198 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.951309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.503066Z digest=sha256:69123729020311a14c504da1731c6bd46aafff339c57dedd8c5af2194c9d73c4

Observation 0e6ec4c6-662e-488e-ab21-9f12c52a8dfa · outbound

This paper cites operation.

Visually Interpretable Subtask Reasoning for Visual Question Answering operation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:20.936940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.507604Z digest=sha256:fdc6afdb921168f4bf4d26c0d859582fbfb29a84c51006c858dc22e922133b5f

Observation e13435c8-bd9f-4127-aee1-3b75ae3e0f11 · outbound

This paper cites operation.

Visually Interpretable Subtask Reasoning for Visual Question Answering operation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:20.921903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.513204Z digest=sha256:f2cb60c7c34b9a27696b6ca57bb7dcbce33b233c1f5636909312d094645a38bc

Observation 18fb048a-64b3-4e47-903f-07ec1da242fe · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.905909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.518928Z digest=sha256:02fa94a9605a91690c86ec53948a53af22179681c609830fd468a843cf3bd4c4

Observation 56da4b2f-b488-4a35-8b9e-3d49453cd90f · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.891315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.523526Z digest=sha256:7472f529d905533af05421cfeffd37368766a951e9ac2385254931d00ec6e36c

Observation 4bd68034-c971-4b4a-addc-0163be50cd2c · outbound

This paper cites Noticeably, For Operation ’choose rel’ in the last step, keeping identical with Final Answer.

Visually Interpretable Subtask Reasoning for Visual Question Answering Noticeably, For Operation ’choose rel’ in the last step, keeping identical with Final Answer

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:20.875913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.528061Z digest=sha256:90386fa1ea859368d49f9e935b4ff0d29d0a7f2312d8402bd7dbafec5980b3d3

Observation 1da92d97-04e5-4781-9466-ba116211781f · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.859793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.532673Z digest=sha256:999d309aa3399ab51eb4e1fb2f300bc86276929db6823d1f0bf910770c45bba7

Observation 6cc01cd0-d368-43e8-a3b8-45b4553e1f34 · outbound

This paper cites attribute value.

Visually Interpretable Subtask Reasoning for Visual Question Answering attribute value

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:20.844059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T22:10:20.537125Z digest=sha256:f33892c2df3d8984a7352d66482e03c6cde930fefe19c51833a4d469e09cb6f0

Pith citing papers

No inbound Pith citation observations are available.