Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:10:20.537125Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2505.08084.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:10:20.537125Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c8d57195-3bbb-4101-9a88-404d38b7388c · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b18430cf-a088-42ea-bee7-70789ea7f3a9 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Qwen Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51d7f70d-9620-4b87-a3bf-1b386f67adb4 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Language Models are Few-Shot Learners
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095b047a-0e4d-4f62-933f-5456fd8b3a0d · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130b10ec-0287-4829-85c5-c75def6851eb · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Gonzalez, Ion Stoica, and Eric P
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12394637-78d5-44ce-a0c1-3a8e0f1b3c67 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 52258c8f-1b57-4fe1-a1d4-17e319040724 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8acdc8e-a4f6-4348-861c-9478d3f69d88 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Cric: A vqa dataset for compositional reasoning on vision and commonsense
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 022bf1e4-0d4c-41d4-8e49-8103c27707cf · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering OmniFusion Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6fe1352-bf80-47a9-ad02-1d5ff373d2f8 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d6104dc0-7c3e-4659-81d1-acfa488dd8be · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Visual pro- gramming: Compositional visual reasoning without training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 415be4a6-2aea-444b-ba2e-fa79b0e12c3a · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Visual program distillation: Distilling tools and programmatic reasoning into vision-language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1d3a3e97-46a7-4e47-a07e-74cf9850c2dd · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Vtimellm: Empower llm to grasp video moments
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 11d53f17-102f-4614-b928-e931eeb79990 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Hudson and Christopher D
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1ec608ae-d00b-491e-acc8-c7e38c86d317 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unveiling the Invisible: Captioning Videos with Metaphors
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f8812c3-d6bd-4c66-a40b-97dbea6cdb58 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Hydra: A hyper agent for dynamic compositional visual reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 03a10c82-7b4b-4064-938e-72f2afd0b73f · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Exploring question decomposition for zero-shot vqa
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e2bf1655-072d-4f17-a59b-36cf535a8729 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering A Survey on Benchmarks of Multimodal Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dbf9bf1-7862-4921-b6e7-f7b041da87a2 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 82eaea41-d40d-432f-92cf-32bd7cc8cba7 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Grounded language-image pre-training
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8422c373-7620-4f7d-b1d2-d2e6a6fa6351 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Improved baselines with visual instruction tuning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2a78b72e-4834-4bfe-8a47-5fdb40126675 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Visual instruction tuning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 714d25c7-6672-4128-9412-8bbeed9f8543 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering NVILA: Efficient Frontier Visual Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2da71b1-6def-4915-b75b-e768b7b39b48 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering SAE-V: Interpreting Multimodal Models for Enhanced Alignment
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3314594d-8828-4ee9-a95a-72dd0bd8c313 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Groma: Localized visual tokenization for grounding multimodal large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8dd6e76c-f00d-4fce-a56f-341a5f0b2c00 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Task navigator: Decomposing complex tasks for multimodal large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 098a1cc4-0926-49c3-a0b9-961a7620bcbc · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Language models are unsu- pervised multitask learners
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56242bd3-f485-4a5a-b022-16709cf86950 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Learning transferable visual models from natural language supervision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78eaef3a-5e03-43eb-9044-0c9d17e84c99 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7fb21a7c-3602-48e5-b007-df7500a0dc38 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Reid, and Silvio Savarese
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 83604e42-d8a2-422a-843b-a44ed6d31051 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Towards vqa models that can read
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a56892f6-6ff5-441c-ba6e-6727218c0fc3 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Vipergpt: Visual inference via python execution for reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 88fa69f3-f3f3-42c7-a819-337d262f4558 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering LLaMA: Open and Efficient Foundation Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99dc1080-943a-4b33-bf29-d3a1c23c8a73 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80127df8-7561-464f-9fd8-d4de5a3cc5f5 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Emu3: Next-Token Prediction is All You Need
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc611ff7-469e-4ed2-a66f-4f723daf1d09 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Sigmoid loss for language image pre-training
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9e7636c1-ccd7-4748-b7f2-14d408a1f048 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Visual Question Decomposition on Multimodal Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3f517add-8723-4f26-893e-fb573d440949 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c653e165-58e0-41d2-abd5-5d228e10c501 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd230b6e-d12d-4a36-835b-306f63f975c9 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e4033c-4b1e-480a-a8ee-0f54177fe556 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a1bc3843-4473-4ec6-810a-652359377a16 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 24154824-e96c-4800-aee5-c2399b8ee17f · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 55b571ae-f6ec-4e07-a902-966a13d8760e · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c0081a71-aa4e-47f9-8540-025afa851375 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aa8aa685-70b3-4c5d-84ed-cf5f06946d2b · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bcc2dedb-6211-439c-8a98-893a60aa2805 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c3a13de9-2ac2-4ed5-8023-a54cf1a0985b · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 21dace39-e704-4965-8530-0ff3222ee89e · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ee11851a-e4a0-4335-8be7-2f146bdc28e9 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 431f83bc-9237-4d31-8344-0b339c2e1b6a · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 77845619-6325-4233-80eb-14ee0e682b38 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9da2bdb1-3efa-4246-a605-3cb69671eb11 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f641b501-ac5a-4a7b-a387-838ba6ba578c · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ff43abc4-5a44-4201-9788-d387e55320e8 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation caeaaf21-d7ea-46eb-aa83-801862f25198 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0e6ec4c6-662e-488e-ab21-9f12c52a8dfa · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering operation
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e13435c8-bd9f-4127-aee1-3b75ae3e0f11 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering operation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 18fb048a-64b3-4e47-903f-07ec1da242fe · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 56da4b2f-b488-4a35-8b9e-3d49453cd90f · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4bd68034-c971-4b4a-addc-0163be50cd2c · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Noticeably, For Operation ’choose rel’ in the last step, keeping identical with Final Answer
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1da92d97-04e5-4781-9466-ba116211781f · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6cc01cd0-d368-43e8-a3b8-45b4553e1f34 · outbound
Visually Interpretable Subtask Reasoning for Visual Question Answering attribute value
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
No inbound Pith citation observations are available.