Pith. sign in

Paper Citation Record · LEDGER

Visually Interpretable Subtask Reasoning for Visual Question Answering

As of 16 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2505.08084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08084 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:10:20.537125Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8d57195-3bbb-4101-9a88-404d38b7388c · outbound

This paper cites GPT-4 Technical Report.

Visually Interpretable Subtask Reasoning for Visual Question Answering GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.239076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.239076Z digest=sha256:6e57faea140fed1d439799da8f2cb77d0e80207c94d8d80df725bb61775c3c45

Observation b18430cf-a088-42ea-bee7-70789ea7f3a9 · outbound

This paper cites Qwen Technical Report.

Visually Interpretable Subtask Reasoning for Visual Question Answering Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.244796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.244796Z digest=sha256:9ccc15e152cf85e21c06c801b020521c61bde10c02199ab31433623a57672d12

Observation 51d7f70d-9620-4b87-a3bf-1b386f67adb4 · outbound

This paper cites Language Models are Few-Shot Learners.

Visually Interpretable Subtask Reasoning for Visual Question Answering Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.249973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.249973Z digest=sha256:c282243af4d33e85bd58810fb5df05409ea86ade366b84505cfde60f6289b71c

Observation 095b047a-0e4d-4f62-933f-5456fd8b3a0d · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Visually Interpretable Subtask Reasoning for Visual Question Answering Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.255873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.255873Z digest=sha256:e62d984c7269ee42e3302adcf797be1c39197bfed3ca6a112ab96fb36cbbd70d

Observation 130b10ec-0287-4829-85c5-c75def6851eb · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Visually Interpretable Subtask Reasoning for Visual Question Answering Gonzalez, Ion Stoica, and Eric P

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.260923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.260923Z digest=sha256:07698679efe5c5486fe3fdcbf4dde49ed3d63025f70d132bbd5f088a4e4956b6

Observation 12394637-78d5-44ce-a0c1-3a8e0f1b3c67 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Visually Interpretable Subtask Reasoning for Visual Question Answering Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.500334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.265480Z digest=sha256:8c37cfc140737a7371deab2febece8c8d3dfd7338d5ab5eb953121f3ceb4e1f4

Observation 52258c8f-1b57-4fe1-a1d4-17e319040724 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.270682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.270682Z digest=sha256:a1b25645d87e33df0c45ca88e7b09844871683ee25fc98f2f87dac7458e5e8b5

Observation c8acdc8e-a4f6-4348-861c-9478d3f69d88 · outbound

This paper cites Cric: A vqa dataset for compositional reasoning on vision and commonsense.

Visually Interpretable Subtask Reasoning for Visual Question Answering Cric: A vqa dataset for compositional reasoning on vision and commonsense

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.484735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.275366Z digest=sha256:dd05bc5195c649b66deb484d3654a2fec3b8a9539cde8a35b84218a88f0a473e

Observation 022bf1e4-0d4c-41d4-8e49-8103c27707cf · outbound

This paper cites OmniFusion Technical Report.

Visually Interpretable Subtask Reasoning for Visual Question Answering OmniFusion Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.279770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.279770Z digest=sha256:e6def2770073772cd46fe805bef59483565b1c3e63d1ea1895a4da1b2b746dd1

Observation d6fe1352-bf80-47a9-ad02-1d5ff373d2f8 · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

Visually Interpretable Subtask Reasoning for Visual Question Answering Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.469059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.284387Z digest=sha256:00b931fee397a42c8fe4da0d92f7e3bdb721a62f6d82d5962f2dc037fc44cb6d

Observation d6104dc0-7c3e-4659-81d1-acfa488dd8be · outbound

This paper cites Visual pro- gramming: Compositional visual reasoning without training.

Visually Interpretable Subtask Reasoning for Visual Question Answering Visual pro- gramming: Compositional visual reasoning without training

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.454740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.289765Z digest=sha256:4c12dc913728c69cbf6d799b862244f47e0ba26b6fc41c836366e2c1d14ba068

Observation 415be4a6-2aea-444b-ba2e-fa79b0e12c3a · outbound

This paper cites Visual program distillation: Distilling tools and programmatic reasoning into vision-language models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Visual program distillation: Distilling tools and programmatic reasoning into vision-language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.440595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.294585Z digest=sha256:fdc30882d31e2b0f250060afd8e002d1a4105bfcfdb2119ec6827c4f9709a79f

Observation 1d3a3e97-46a7-4e47-a07e-74cf9850c2dd · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Visually Interpretable Subtask Reasoning for Visual Question Answering Vtimellm: Empower llm to grasp video moments

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.424498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.299056Z digest=sha256:3f506f799c9cde2043b344fbda1729d3a2925c0b23f821c5534b8f1de28e2bc1

Observation 11d53f17-102f-4614-b928-e931eeb79990 · outbound

This paper cites Hudson and Christopher D.

Visually Interpretable Subtask Reasoning for Visual Question Answering Hudson and Christopher D

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.409593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.303384Z digest=sha256:5eef0cdfc435a3ff34b695cbcf9477d76d4913869fdb15e275fb40a909875662

Observation 1ec608ae-d00b-491e-acc8-c7e38c86d317 · outbound

This paper cites Unveiling the Invisible: Captioning Videos with Metaphors.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unveiling the Invisible: Captioning Videos with Metaphors

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.307370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.307370Z digest=sha256:b3a2269db2cf1bf938565168f297bc5a4ba9c656cc6ab64bb8137851aea73250

Observation 9f8812c3-d6bd-4c66-a40b-97dbea6cdb58 · outbound

This paper cites Hydra: A hyper agent for dynamic compositional visual reasoning.

Visually Interpretable Subtask Reasoning for Visual Question Answering Hydra: A hyper agent for dynamic compositional visual reasoning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.394767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.312068Z digest=sha256:d37c6cd9e6d335668e8bb1a3185d40699e9d2486fb30a468178ef4d78e4c8874

Observation 03a10c82-7b4b-4064-938e-72f2afd0b73f · outbound

This paper cites Exploring question decomposition for zero-shot vqa.

Visually Interpretable Subtask Reasoning for Visual Question Answering Exploring question decomposition for zero-shot vqa

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.380548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.316613Z digest=sha256:1bbd02bd8b7323d2646314c9340831b40623ff5454c20864f74f062e5e901a08

Observation e2bf1655-072d-4f17-a59b-36cf535a8729 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering A Survey on Benchmarks of Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.321108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.321108Z digest=sha256:6ec11f0255e599a825f4bb44569ad809560f9bbcc8541ae3ba009a158687f537

Observation 2dbf9bf1-7862-4921-b6e7-f7b041da87a2 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.364463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.326138Z digest=sha256:857bc6331688fbd962ddd8b62e7c70c7ce5ca278ce990aad9e4975c0e0afd17e

Observation 82eaea41-d40d-432f-92cf-32bd7cc8cba7 · outbound

This paper cites Grounded language-image pre-training.

Visually Interpretable Subtask Reasoning for Visual Question Answering Grounded language-image pre-training

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.349464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.330701Z digest=sha256:5dbc5dfb27ae7ba6fbe29ff0cc7cd1e99f4b20c4942b1daa175a28d6194a53d1

Observation 8422c373-7620-4f7d-b1d2-d2e6a6fa6351 · outbound

This paper cites Improved baselines with visual instruction tuning.

Visually Interpretable Subtask Reasoning for Visual Question Answering Improved baselines with visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.334420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.335097Z digest=sha256:3f7f643da964ef8d05ed92e7cdb920461d7c5d20b599803b1352224c590c9340

Observation 2a78b72e-4834-4bfe-8a47-5fdb40126675 · outbound

This paper cites Visual instruction tuning.

Visually Interpretable Subtask Reasoning for Visual Question Answering Visual instruction tuning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.319490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.339444Z digest=sha256:01b923a8919aeffbbf09a4d394a96be33176f033ebea8b0cf6f8b63d82814959

Observation 714d25c7-6672-4128-9412-8bbeed9f8543 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering NVILA: Efficient Frontier Visual Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.343778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.343778Z digest=sha256:f9084eb10b50b9fc01d6d708fcaa393a519e8eb6773a48ff865abcf7f73ed818

Observation a2da71b1-6def-4915-b75b-e768b7b39b48 · outbound

This paper cites SAE-V: Interpreting Multimodal Models for Enhanced Alignment.

Visually Interpretable Subtask Reasoning for Visual Question Answering SAE-V: Interpreting Multimodal Models for Enhanced Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.348244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.348244Z digest=sha256:8666afd24a9c0e36bfa9df5779197af52ed396fb1ede8de160feb5aed73feaf3

Observation 3314594d-8828-4ee9-a95a-72dd0bd8c313 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Groma: Localized visual tokenization for grounding multimodal large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.304205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.353038Z digest=sha256:0a7b70b08f93e451a72ed5490fc7aabd78c43d8cf190e7975603935f4055df96

Observation 8dd6e76c-f00d-4fce-a56f-341a5f0b2c00 · outbound

This paper cites Task navigator: Decomposing complex tasks for multimodal large language models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Task navigator: Decomposing complex tasks for multimodal large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.289200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.358522Z digest=sha256:aef11862a7a4ca428f97a3d0fb3a75eda3e46d55266f7c7a6e559bb757cdb79c

Observation 098a1cc4-0926-49c3-a0b9-961a7620bcbc · outbound

This paper cites Language models are unsu- pervised multitask learners.

Visually Interpretable Subtask Reasoning for Visual Question Answering Language models are unsu- pervised multitask learners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.364298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.364298Z digest=sha256:86ca6deed16c721fbcc8f38b9f01da55a6008d4f539aa6c47fa96e6c0ac1e002

Observation 56242bd3-f485-4a5a-b022-16709cf86950 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Visually Interpretable Subtask Reasoning for Visual Question Answering Learning transferable visual models from natural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.369214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.369214Z digest=sha256:c0805f72a30afca5beb942becc1cc445138510eb5e8c3908591a520dadd03db8

Observation 78eaef3a-5e03-43eb-9044-0c9d17e84c99 · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

Visually Interpretable Subtask Reasoning for Visual Question Answering Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.254395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.374649Z digest=sha256:9917da60019f3172ea30a625e288010d865a3b984d3799e4a309cc7e77a772b6

Observation 7fb21a7c-3602-48e5-b007-df7500a0dc38 · outbound

This paper cites Reid, and Silvio Savarese.

Visually Interpretable Subtask Reasoning for Visual Question Answering Reid, and Silvio Savarese

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.239321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.379570Z digest=sha256:3e7e4807d9cbc24efdd323c402fd54281b05145b9337b09fc6930b9c8ad087ce

Observation 83604e42-d8a2-422a-843b-a44ed6d31051 · outbound

This paper cites Towards vqa models that can read.

Visually Interpretable Subtask Reasoning for Visual Question Answering Towards vqa models that can read

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.224403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.384030Z digest=sha256:4414ac78fc5801d8c0b686f10bb682ed7a87d61e7df922f7e8ca74f256171a02

Observation a56892f6-6ff5-441c-ba6e-6727218c0fc3 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Visually Interpretable Subtask Reasoning for Visual Question Answering Vipergpt: Visual inference via python execution for reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.209077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.387978Z digest=sha256:e44553752dfeafd6f980e826d6e48881f31842e38ef4ec2d1972df2fff0125bc

Observation 88fa69f3-f3f3-42c7-a819-337d262f4558 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.392320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.392320Z digest=sha256:01ee03e6511fdf08337be0689e2866c9c382b39b8c59a36fbc37df29c7ebb1a9

Observation 99dc1080-943a-4b33-bf29-d3a1c23c8a73 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Visually Interpretable Subtask Reasoning for Visual Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.397283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.397283Z digest=sha256:f325926c35e30c165b4ae8509b957c491c3bed70ed77078bccdb925cec022099

Observation 80127df8-7561-464f-9fd8-d4de5a3cc5f5 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Visually Interpretable Subtask Reasoning for Visual Question Answering Emu3: Next-Token Prediction is All You Need

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.402206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.402206Z digest=sha256:f6a1fb76ebc17b646bd677e57239990878cb3a8636d2a776a820534788387068

Observation cc611ff7-469e-4ed2-a66f-4f723daf1d09 · outbound

This paper cites Sigmoid loss for language image pre-training.

Visually Interpretable Subtask Reasoning for Visual Question Answering Sigmoid loss for language image pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:21.193319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.406738Z digest=sha256:83c0576ef19e45381f750abcd89962588d85dbf273733b79991e89af23494bba

Observation 9e7636c1-ccd7-4748-b7f2-14d408a1f048 · outbound

This paper cites Visual Question Decomposition on Multimodal Large Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Visual Question Decomposition on Multimodal Large Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:10:20.627569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.410966Z digest=sha256:774ca1e0a5bc3d6e149ad339dee72a4d58c7e200ad7ff4f98a2a092b88673fff

Observation 3f517add-8723-4f26-893e-fb573d440949 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.415537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.415537Z digest=sha256:ad4eebf311f15436976c24484f8d1a5c873a742d3b4e68e3e31245f6e28b4c2d

Observation c653e165-58e0-41d2-abd5-5d228e10c501 · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Visually Interpretable Subtask Reasoning for Visual Question Answering MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.420225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.420225Z digest=sha256:2c9911046d6e9380bd486dd86149add7db78d22523e9f3bd9100ca737949287d

Observation fd230b6e-d12d-4a36-835b-306f63f975c9 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Visually Interpretable Subtask Reasoning for Visual Question Answering MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:10:20.425142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:10:20.425142Z digest=sha256:bfca527abfdc6723c32fa564bd57e0fbf3a6aaf73e57f2a056f2b930420e163b

Observation 15e4033c-4b1e-480a-a8ee-0f54177fe556 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.179212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.430005Z digest=sha256:c522daf432222d3e715becb27955cf7202ccad3095c87a987e80c51552cb1e2e

Observation a1bc3843-4473-4ec6-810a-652359377a16 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.163826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.434618Z digest=sha256:ffe2c16c54e96a431ed9ff7d6c734ecb2c7132b386e1fa6664126cc0fb58433f

Observation 24154824-e96c-4800-aee5-c2399b8ee17f · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.148370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.440312Z digest=sha256:94c125cc1fe38250d3b7462b4e06998dad23e7ee9ef6318a7e415c05a4911e20

Observation 55b571ae-f6ec-4e07-a902-966a13d8760e · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.133967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.445059Z digest=sha256:40a88f2a2bce8f1977547db602b03ccf4b48c176fcf6fe667016b17786e634a1

Observation c0081a71-aa4e-47f9-8540-025afa851375 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.119511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.449493Z digest=sha256:0c978ad99972733c0b9a5a4be6ca766f19e7fc25cba517bb4c587efa7ead8dbf

Observation aa8aa685-70b3-4c5d-84ed-cf5f06946d2b · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.104652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.455126Z digest=sha256:2221e015f88e4f2d7b79fd905cd662797fa4043a2cb4f974a7890df8b5ef570f

Observation bcc2dedb-6211-439c-8a98-893a60aa2805 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.090074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.459528Z digest=sha256:521474998d1c9c939708d5fd1c616c374b2e9550a0b65aa10ac3e1328c81af61

Observation c3a13de9-2ac2-4ed5-8023-a54cf1a0985b · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.074152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.464202Z digest=sha256:4f8e81924750b51cca3b74076c5fbec49044b1ff3e69c98623b004b4f8426571

Observation 21dace39-e704-4965-8530-0ff3222ee89e · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.058533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.468971Z digest=sha256:1cad90ebf974659dd2e192839fac8f4749429833a2e274606baba096d56566a2

Observation ee11851a-e4a0-4335-8be7-2f146bdc28e9 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.043097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.474476Z digest=sha256:7ddd49aa5f69e2e5eac81770ee298a8e49d416ecd95be612397f1f85dd423ba2

Observation 431f83bc-9237-4d31-8344-0b339c2e1b6a · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.027802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.479284Z digest=sha256:0c01b922694ada3b4ef67ee1b238616b35ca67e89dc4d97885ad9280395dc270

Observation 77845619-6325-4233-80eb-14ee0e682b38 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:21.013149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.483962Z digest=sha256:31480a871145d978f2f92da9a9f6b38fa7d14bca87eebfdfadae5147d60fba4a

Observation 9da2bdb1-3efa-4246-a605-3cb69671eb11 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.998129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.488975Z digest=sha256:3985e5236a5846b6a847f66e56419b5767c0c17ef3776c83f70f38d7a8389c1c

Observation f641b501-ac5a-4a7b-a387-838ba6ba578c · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.982509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.493679Z digest=sha256:2bca663ecb44c21e0843ecabedf49c4363ca603c3e0d31b285971884f933cbdc

Observation ff43abc4-5a44-4201-9788-d387e55320e8 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.966695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.498344Z digest=sha256:6b3fc575d05f40491b72c1bd2f8ffbffe66bff9a397667d5375cb011efa26865

Observation caeaaf21-d7ea-46eb-aa83-801862f25198 · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.951309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.503066Z digest=sha256:52d703d928271cd4eb7895e16c54724c21b37144342cb7a41849f0d9bef169b5

Observation 0e6ec4c6-662e-488e-ab21-9f12c52a8dfa · outbound

This paper cites operation.

Visually Interpretable Subtask Reasoning for Visual Question Answering operation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:20.936940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.507604Z digest=sha256:9019e2f2934bdf2a1317135c79046f8d0657a1d83a4b9fc99c8976b65273d065

Observation e13435c8-bd9f-4127-aee1-3b75ae3e0f11 · outbound

This paper cites operation.

Visually Interpretable Subtask Reasoning for Visual Question Answering operation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:20.921903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.513204Z digest=sha256:dc46ff5dfeebb7c1015c6221024ca98bf62612f89b9fc3ced07d3de6869025be

Observation 18fb048a-64b3-4e47-903f-07ec1da242fe · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.905909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.518928Z digest=sha256:62d1052e4b8b2596c116795cb32fc648ec13669bd7f4445c0b4d5ac5fc0fc593

Observation 56da4b2f-b488-4a35-8b9e-3d49453cd90f · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.891315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.523526Z digest=sha256:a59e98d4761f0923281ca010b8b75fe32a545f4a4f2780cb8767d918301ab142

Observation 4bd68034-c971-4b4a-addc-0163be50cd2c · outbound

This paper cites Noticeably, For Operation ’choose rel’ in the last step, keeping identical with Final Answer.

Visually Interpretable Subtask Reasoning for Visual Question Answering Noticeably, For Operation ’choose rel’ in the last step, keeping identical with Final Answer

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:20.875913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.528061Z digest=sha256:483896edde0a14ca147af065d366c0c79f5e14047a5844a35fb7808ce7d00868

Observation 1da92d97-04e5-4781-9466-ba116211781f · outbound

This paper cites an unresolved cited work.

Visually Interpretable Subtask Reasoning for Visual Question Answering Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:10:20.859793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.532673Z digest=sha256:52b38aeb66a2a0414b9a01b15b20b1fa4898dca00885a98dba578c3778194c0b

Observation 6cc01cd0-d368-43e8-a3b8-45b4553e1f34 · outbound

This paper cites attribute value.

Visually Interpretable Subtask Reasoning for Visual Question Answering attribute value

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:10:20.844059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:10:20.537125Z digest=sha256:4a30f9c1d35c7b37a26a296a1baa985be0a360e4776cb3298a5682be18c57f60

Pith citing papers

No inbound Pith citation observations are available.